Blog

212 articles

Research, case studies, and engineering deep-dives from the Neo team.

How NEO Improved the Pi Coding Agent from 82% to 95%
AI Agents & Automation

How NEO Improved the Pi Coding Agent from 82% to 95%

NEO ran an autonomous research loop on the pi coding agent and raised QuixBugs from 49/60 to 57/60 on development, then 19/20 on a sealed holdout. Same model. Five prompt guidelines. One file changed.

August 29, 2026·12 min
Read
Qwen3.8 vs Qwen3.6 vs Muse Glimmer: Why Two Graders Picked Different Winners
LLM Evaluation & Benchmarking

Qwen3.8 vs Qwen3.6 vs Muse Glimmer: Why Two Graders Picked Different Winners

On 38 comparable answers, Muse Glimmer 30B wins deterministic checks while a blind LLM judge prefers Qwen3.8-27B. Built and run end to end with Neo — the winner depends on how you grade.

August 17, 2026·10 min
Read
Evaluating Voice Cloning Models on CPU: A Practical Benchmark of Pocket TTS, Kokoro, Audio8, and XTTS-v2
LLM Evaluation & Benchmarking

Evaluating Voice Cloning Models on CPU: A Practical Benchmark of Pocket TTS, Kokoro, Audio8, and XTTS-v2

Audio8 now clones on CPU: identity near-tie with XTTS (SIM 0.5454 vs 0.5569). Among cloners Audio8 is the practical pick (UTMOS 3.3657, RTF 1.4214); Kokoro still wins fixed-voice quality and speed. VCTK sweep built with Neo.

August 15, 2026·16 min
Read
Grok 4.6 vs Kimi K3 vs Opus 5 on Kokoro-82M TTS: Strong Research, No Production-Ready Speedup
LLM Evaluation & Benchmarking

Grok 4.6 vs Kimi K3 vs Opus 5 on Kokoro-82M TTS: Strong Research, No Production-Ready Speedup

Grok 4.6, Kimi K3, and Opus 5 each tried to speed up Kokoro-82M TTS on CPU. Study scores: Opus 83, Kimi 60, Grok 51. Grok's best verified gain was about 6–8% in one setup, but audio checks failed, so nothing shipped.

August 13, 2026·8 min
Read
Making OmniVoice 6.2× Faster Than Its 32-Step Default Without Retraining
Model Optimization & Inference

Making OmniVoice 6.2× Faster Than Its 32-Step Default Without Retraining

OmniVoice sounded good but felt slow. Without training a new model, we made it about 6× faster than the default path, and still about 3× faster than the official fast setting, then kept the quality checks honest.

August 12, 2026·15 min
Read
Three ASR models tied on the benchmark. One was almost three times worse on real speech.
LLM Evaluation & Benchmarking

Three ASR models tied on the benchmark. One was almost three times worse on real speech.

Five ASR models on 296 real clips across 19 conditions. Clean-speech WER ties at 4.4%, then accent and noise separate the field, with committed transcripts, bootstrap CIs, and a CI regression gate.

August 11, 2026·16 min
Read
Kimi K3 vs Opus 5: Which Model Built the Better Kokoro CPU Optimizer?
LLM Evaluation & Benchmarking

Kimi K3 vs Opus 5: Which Model Built the Better Kokoro CPU Optimizer?

Using Neo BYOK, both models optimized Kokoro-82M on CPU. Opus built the stronger study, but its claimed 18.66% speedup became 9.81% slower on replay.

August 1, 2026·8 min
Read
Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering

Three AutoML artifacts, all green suites (45/45, 83/83, 37/37). Execution found silent trust failures. Scores: Kimi 68, Fable 61, GLM 54.

July 24, 2026·10 min
Read
Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering

Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.

July 20, 2026·8 min
Read
Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC
Model Optimization & Inference

Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC

Neo profiled NVIDIA Parakeet TDT 0.6B v3 on CPU, ran keep/discard ladders, and froze a static-QDQ production pack that cut primary RTF by ~2.07× on EPYC (~1.42× on Apple Silicon). Runtime-only knobs never cleared 5%.

July 14, 2026·14 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark

Kyutai's Pocket TTS joins the CPU TTS benchmark: 6 configs, 180 timed runs, and 36 WAV samples across RTF, latency, throughput, and UTMOS MOS, plus zero-shot voice cloning from 5 seconds of audio. Built end-to-end with Neo.

July 6, 2026·16 min
Read
Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval
LLM Evaluation & Benchmarking

Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval

Neo evaluated Qwythos-9B at Q4_K_M and Q8_0 on GSM8K, IFEval, and HumanEval from a single prompt. GSM8K hit 84%, IFEval 66%, HumanEval 0% — and Q4 is nearly as good as Q8 for math.

July 3, 2026·10 min
Read