
Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering
Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.
Research, case studies, and engineering deep-dives from the Neo team.

Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.

Neo profiled NVIDIA Parakeet TDT 0.6B v3 on CPU, ran keep/discard ladders, and froze a static-QDQ production pack that cut primary RTF by ~2.07× on EPYC (~1.42× on Apple Silicon). Runtime-only knobs never cleared 5%.

Kyutai's Pocket TTS joins the CPU TTS benchmark: 6 configs, 180 timed runs, and 36 WAV samples across RTF, latency, throughput, and UTMOS MOS, plus zero-shot voice cloning from 5 seconds of audio. Built end-to-end with Neo.

Neo evaluated Qwythos-9B at Q4_K_M and Q8_0 on GSM8K, IFEval, and HumanEval from a single prompt. GSM8K hit 84%, IFEval 66%, HumanEval 0% — and Q4 is nearly as good as Q8 for math.

Claude Sonnet vs Ornith:35b on CodeArena — an AI coding benchmark where both passed 7/24 tests. Process scores diverged sharply: self-finalization, tool mix, and $0 local cost vs $2.73 API fees.

NEO built and ran the Ornith Evaluation Framework autonomously on Ornith-1.0-35B: 100/100 terminal safety, Level 6/15 skill ceiling. What the model scored and how the harness was verified.

GLM 5.2 built TrackLab end to end — a browser CV studio with detection, tracking, and line counting — entirely through NEO BYOK. Same agent workflow, different model.

What NEO's BYOK found comparing GLM 5.2 vs Kimi K2.6 in VS Code: Kimi won capability, GLM craftsmanship. Both failed Article 75 differently. Specs and setup guide.

A CPU-only TTS benchmark of Kokoro 82M, Supertonic 3, and Inflect-Nano-v1 across RTF, latency, throughput, and UTMOS MOS: 5 configs, 150 timed runs, and 30 WAV samples. Built end-to-end with Neo.

874 real-world agent failure records from a 7-phase governed pipeline: one prompt, HuggingFace ingestion, adversarial verification, and full provenance without synthetic padding.

A 750-example agent eval benchmark where Neo MCP turns one-shot generation into governed synthetic data operations: remediation loops, replayable verification, and audit infrastructure built for multi-release eval programs.

500 synthetic incident replays for SRE training: Neo MCP adds exact severity quotas, 11-phase timelines, generation-time dedup, and nested-field validation that turns generation output into an audit-ready dataset.