Inflect-Micro-v2: complete voice in 9.36M parameters

https://news.ycombinator.com/rss Hits: 16
Summary

Inflect-Micro-v2 Complete local text-to-waveform speech synthesis under 10M parameters. Fixed-voice English TTS with deterministic seeds, long-text handling, and CPU or CUDA inference. A note from Owen I built and funded Inflect v2 independently. If this release finds a real audience, I would like to continue the project with a broader v3, which might include things like more langauges, voices, and stability improvements. If the model is useful to you, leaving a like on Hugging Face genuinely helps more people discover it. 9,356,513 deployable parameters · 37.53 MB FP32 · 24 kHz mono output Inflect v2 uses one public API across two sizes: Micro prioritizes quality below 10M parameters; Nano prioritizes footprint below 4M. Explore this model card Listen These are held-out text generations, not reconstructions of training audio. Each transcript is shown exactly as passed to the public frontend. Test Exact transcript Generated audio Conversational It wasn't until later that I realized what had actually happened. Punctuation First, close the window; second, turn off the lamp; finally, lock the door. Numbers The package weighs twelve point six kilograms and arrived on July twenty-first. Names and places Gwendolyn photographed the eucalyptus trees outside Ljubljana. Technical The system runs on three core components that all have to stay in sync. Evaluation No single metric captures TTS quality. Inflect v2 reports human preference, predicted naturalness, multi-ASR intelligibility, complete footprint, and runtime separately rather than compressing them into one unverifiable score. Community preference ↑ UTMOS22 ↑ Two-ASR semantic WER ↓ Complete FP32 weights ↓ 4-thread CPU throughput ↑ 66.2% 4.395 3.99% 37.53 MB 6.28× real-time The headline row always refers to Inflect-Micro-v2. Detailed competitor results and protocol boundaries are kept visible below. Comparison set. Results include KittenTTS Nano, Piper Low, and Supertonic 3, established compact or local TTS baselines wi...

First seen: 2026-07-26 02:51

Last seen: 2026-07-26 18:02