Blog

205 articles

Research, case studies, and engineering deep-dives from the Neo team.

Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering
LLM Evaluation & Benchmarking

Kimi K3 vs GLM 5.2: Benchmarking Frontier Models on Production ML Engineering

Both shipped green suites; execution inverted the ranking. Kimi K3 68/100 vs GLM 5.2 54/100 on production ML frameworks. Built with NEO BYOK.

July 20, 2026·8 min
Read
Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC
Model Optimization & Inference

Making Parakeet Faster on CPU: Static QDQ Cut Primary RTF ~2× on EPYC

Neo profiled NVIDIA Parakeet TDT 0.6B v3 on CPU, ran keep/discard ladders, and froze a static-QDQ production pack that cut primary RTF by ~2.07× on EPYC (~1.42× on Apple Silicon). Runtime-only knobs never cleared 5%.

July 14, 2026·14 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 vs Pocket TTS: A Real CPU TTS Benchmark

Kyutai's Pocket TTS joins the CPU TTS benchmark: 6 configs, 180 timed runs, and 36 WAV samples across RTF, latency, throughput, and UTMOS MOS, plus zero-shot voice cloning from 5 seconds of audio. Built end-to-end with Neo.

July 6, 2026·16 min
Read
Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval
LLM Evaluation & Benchmarking

Qwythos-9B Evaluation: Benchmarking a 9B Reasoning Model on GSM8K, IFEval, and HumanEval

Neo evaluated Qwythos-9B at Q4_K_M and Q8_0 on GSM8K, IFEval, and HumanEval from a single prompt. GSM8K hit 84%, IFEval 66%, HumanEval 0% — and Q4 is nearly as good as Q8 for math.

July 3, 2026·10 min
Read
Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.
LLM Evaluation & Benchmarking

Claude and Ornith Tied on Tests. Their Behavior Couldn't Be More Different.

Claude Sonnet vs Ornith:35b on CodeArena — an AI coding benchmark where both passed 7/24 tests. Process scores diverged sharply: self-finalization, tool mix, and $0 local cost vs $2.73 API fees.

July 2, 2026·14 min
Read
NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously
LLM Evaluation & Benchmarking

NEO Evaluated Ornith-1.0-35B: Terminal Safety and Coding Skill Ceiling, Built Autonomously

NEO built and ran the Ornith Evaluation Framework autonomously on Ornith-1.0-35B: 100/100 terminal safety, Level 6/15 skill ceiling. What the model scored and how the harness was verified.

June 27, 2026·9 min
Read
GLM 5.2 Built TrackLab: Browser Computer Vision with NEO BYOK
LLM Evaluation & Benchmarking

GLM 5.2 Built TrackLab: Browser Computer Vision with NEO BYOK

GLM 5.2 built TrackLab end to end — a browser CV studio with detection, tracking, and line counting — entirely through NEO BYOK. Same agent workflow, different model.

June 26, 2026·10 min
Read
GLM 5.2 vs Kimi K2.6: Same Agent Workflow, Same Citation Bug, Opposite Failures
LLM Evaluation & Benchmarking

GLM 5.2 vs Kimi K2.6: Same Agent Workflow, Same Citation Bug, Opposite Failures

What NEO's BYOK found comparing GLM 5.2 vs Kimi K2.6 in VS Code: Kimi won capability, GLM craftsmanship. Both failed Article 75 differently. Specs and setup guide.

June 23, 2026·8 min
Read
Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1: A Real CPU TTS Benchmark
LLM Evaluation & Benchmarking

Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1: A Real CPU TTS Benchmark

A CPU-only TTS benchmark of Kokoro 82M, Supertonic 3, and Inflect-Nano-v1 across RTF, latency, throughput, and UTMOS MOS: 5 configs, 150 timed runs, and 30 WAV samples. Built end-to-end with Neo.

June 22, 2026·14 min
Read
From Synthetic Data Generation to Dataset Engineering: What Changed When We Added Neo MCP
LLM Evaluation & Benchmarking

From Synthetic Data Generation to Dataset Engineering: What Changed When We Added Neo MCP

874 real-world agent failure records from a 7-phase governed pipeline: one prompt, HuggingFace ingestion, adversarial verification, and full provenance without synthetic padding.

June 17, 2026·14 min
Read
Managed Synthetic Data: What Happens When Dataset Pipelines Start Governing Themselves
LLM Evaluation & Benchmarking

Managed Synthetic Data: What Happens When Dataset Pipelines Start Governing Themselves

A 750-example agent eval benchmark where Neo MCP turns one-shot generation into governed synthetic data operations: remediation loops, replayable verification, and audit infrastructure built for multi-release eval programs.

June 17, 2026·16 min
Read
Synthetic Data Generation Is Easy. Dataset Engineering Is Hard.
LLM Evaluation & Benchmarking

Synthetic Data Generation Is Easy. Dataset Engineering Is Hard.

500 synthetic incident replays for SRE training: Neo MCP adds exact severity quotas, 11-phase timelines, generation-time dedup, and nested-field validation that turns generation output into an audit-ready dataset.

June 17, 2026·12 min
Read