Blog

208 articles

Research, case studies, and engineering deep-dives from the Neo team.

Evaluating Voice Cloning Models on CPU: A Practical Benchmark of Pocket TTS, Kokoro, Audio8, and XTTS-v2
LLM Evaluation & Benchmarking

Evaluating Voice Cloning Models on CPU: A Practical Benchmark of Pocket TTS, Kokoro, Audio8, and XTTS-v2

A CPU-only VCTK sweep of Pocket TTS, Kokoro, Audio8, and XTTS-v2 across speaker similarity, UTMOS, Whisper WER, and RTF — with a true zero-shot Audio8 vs XTTS head-to-head. Built end-to-end with Neo.

August 7, 2026·12 min
Read
Kimi K3 vs Opus 5: Which Model Built the Better Kokoro CPU Optimizer?
LLM Evaluation & Benchmarking

Kimi K3 vs Opus 5: Which Model Built the Better Kokoro CPU Optimizer?

Using Neo BYOK, both models optimized Kokoro-82M on CPU. Opus built the stronger study, but its claimed 18.66% speedup became 9.81% slower on replay.

August 1, 2026·8 min
Read
Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2 vs Fable 5: Benchmarking AI-Generated ML Engineering

Three AutoML artifacts, all green suites (45/45, 83/83, 37/37). Execution found silent trust failures. Scores: Kimi 68, Fable 61, GLM 54.

July 24, 2026·10 min
Read
Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering

Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.

July 20, 2026·8 min
Read
Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC
Model Optimization & Inference

Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC

Neo profiled NVIDIA Parakeet TDT 0.6B v3 on CPU, ran keep/discard ladders, and froze a static-QDQ production pack that cut primary RTF by ~2.07× on EPYC (~1.42× on Apple Silicon). Runtime-only knobs never cleared 5%.

July 14, 2026·14 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark

Kyutai's Pocket TTS joins the CPU TTS benchmark: 6 configs, 180 timed runs, and 36 WAV samples across RTF, latency, throughput, and UTMOS MOS, plus zero-shot voice cloning from 5 seconds of audio. Built end-to-end with Neo.

July 6, 2026·16 min
Read
Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval
LLM Evaluation & Benchmarking

Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval

Neo evaluated Qwythos-9B at Q4_K_M and Q8_0 on GSM8K, IFEval, and HumanEval from a single prompt. GSM8K hit 84%, IFEval 66%, HumanEval 0% — and Q4 is nearly as good as Q8 for math.

July 3, 2026·10 min
Read
Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.
LLM Evaluation & Benchmarking

Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.

Claude Sonnet vs Ornith:35b on CodeArena — an AI coding benchmark where both passed 7/24 tests. Process scores diverged sharply: self-finalization, tool mix, and $0 local cost vs $2.73 API fees.

July 2, 2026·14 min
Read
NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously
LLM Evaluation & Benchmarking

NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously

NEO built and ran the Ornith Evaluation Framework autonomously on Ornith-1.0-35B: 100/100 terminal safety, Level 6/15 skill ceiling. What the model scored and how the harness was verified.

June 27, 2026·9 min
Read
GLM 5.2 Built TrackLab: Browser Computer Vision with NEO BYOK
LLM Evaluation & Benchmarking

GLM 5.2 Built TrackLab: Browser Computer Vision with NEO BYOK

GLM 5.2 built TrackLab end to end — a browser CV studio with detection, tracking, and line counting — entirely through NEO BYOK. Same agent workflow, different model.

June 26, 2026·10 min
Read
GLM 5.2 vs Kimi K2.6: Same Agent Workflow, Same Citation Bug, Opposite Failures
LLM Evaluation & Benchmarking

GLM 5.2 vs Kimi K2.6: Same Agent Workflow, Same Citation Bug, Opposite Failures

What NEO's BYOK found comparing GLM 5.2 vs Kimi K2.6 in VS Code: Kimi won capability, GLM craftsmanship. Both failed Article 75 differently. Specs and setup guide.

June 23, 2026·8 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1: A Real CPU TTS Benchmark

A CPU-only TTS benchmark of Kokoro 82M, Supertonic 3, and Inflect-Nano-v1 across RTF, latency, throughput, and UTMOS MOS: 5 configs, 150 timed runs, and 30 WAV samples. Built end-to-end with Neo.

June 22, 2026·14 min
Read