Foundation Model vs LLM Selection in Bark and VALL-E Speech Pipelines

Audio codec choice matters more than whether you pick Bark or VALL-E.

Contributing Editor · · 11 min read
Cover illustration for “Foundation Model vs LLM Selection in Bark and VALL-E Speech Pipelines”
Voice AI Research · September 30, 2026 · 11 min read · 2,526 words

Old-school TTS treated speech as a signal regression problem: predict a spectrogram, hand it to a vocoder, hope the seams don't show. It worked, technically, the way a fax machine technically sends a document. What it didn't do was sound like a person who meant what they said, because nothing in that pipeline modeled meaning at all, just curves.

VALL-E broke that by reframing text-to-speech as a language modeling problem over discrete codec codes pulled from a neural audio codec. Instead of regressing a waveform, the model predicts tokens, the same basic move that made GPT-style text generation work, just aimed at sound instead of sentences. That's the whole trick, and it's a bigger deal than it sounds: it makes the language model a first-class citizen inside the speech pipeline rather than a text preprocessor sitting upstream of the "real" audio work.

Bark takes the opposite bet. It's a generative audio foundation model in the AudioLM lineage: the language model living inside it is native to audio, not imported from the text-processing world next door. Neither approach is wrong. But they answer a different question, and every engineer picking a foundation model needs to know which question they're actually asking: is this an audio-native generative system, or a speech-shaped wrapper bolted onto a text LLM? Everything downstream, latency budgets, cloning fidelity, hardware bills, deployment shape, traces back to that one fork.

Bark's internal token hierarchy

Bark runs on a three-stage hierarchy: semantic tokens, then coarse acoustic tokens, then fine acoustic tokens, borrowing the same structure Google built for AudioLM. Each stage refines the last, semantic tokens capture the "what," coarse tokens rough in the "how it sounds," fine tokens polish the texture. It's an assembly line, not a single leap from text to waveform.

That structure buys Bark something the codec-LLM lineage doesn't get for free: it can generate laughter, sighing, crying, background noise and simple music, alongside speech, all from one model pass. For creative audio work and games, nothing in the open-source field currently matches it. Most TTS comparisons get reduced to "which one sounds more human," but Bark is answering a different question entirely: which one sounds most alive.

The trade-off appears in the numbers. Generating one second of audio has an RTF of 0.85, among the slowest in the open-source field, before any network overhead. It's a batch tool, built for pre-rendered creative work.

The architecture leaves a door open that Bark itself doesn't walk through: plugging an external LLM upstream to sharpen semantic token quality is optional and architecturally available. That open door is exactly the hallway VALL-E and its descendants built a house around.

VALL-E's codec language model approach and its zero-shot cloning breakthrough

VALL-E's mechanism is almost embarrassingly simple to state: take discrete codec codes from an off-the-shelf neural codec, originally EnCodec, treat them as a vocabulary, and have a language model predict them conditioned on text plus a short acoustic prompt. Three seconds of enrolled recording is enough to clone a voice with no fine-tuning required, per Microsoft Research's project documentation Microsoft / VALL-E-X. The model doesn't learn the speaker, it reads them, in-context, the way a good voice actor picks up an accent from hearing it once. Emotion and acoustic environment ride along in that same three-second sample, preserved without any additional training step Microsoft / VALL-E-X.

The lineage kept moving. VALL-E 2 hit human parity in zero-shot TTS, the first system to do so, and introduced Repetition Aware Sampling specifically to stop the model from stuttering on repeated tokens during decoding. VALL-E R tackled a related failure mode, the tendency to repeat or skip tokens outright, by enforcing monotonic alignment between text and audio. Neither fix is cosmetic. Both attack the exact failure mode that makes codec-language models occasionally sound drunk.

The architectural payoff is the real story, though. Because the language model is the generative engine itself and not a helper sitting to the side, upgrading that LLM directly upgrades the whole speech system. Bark's audio-native design doesn't expose that lever in the same way. Swapping Bark's internals leaves no clean seam. Swapping the backbone LLM in a VALL-E-style system potentially makes the entire pipeline smarter overnight.

None of this is free. VALL-E gives up Bark's native non-speech audio generation, its paralinguistic range, its audio-domain flexibility, all of it. It's a voice model. It is not, and doesn't pretend to be, a sound-effects model.

The codec layer as a shared bottleneck and an independent design axis

Both paradigms lean on a neural codec to turn continuous sound into discrete tokens, VALL-E originally via EnCodec, Bark via its own parallel hierarchical structure. That shared dependency ties both paradigms' output quality to the same codec's limitations, capping performance regardless of which camp's marketing claims apply. The 2026 MOSS-Audio-Tokenizer paper argues codec scaling is the next real bottleneck for audio foundation models, not LLM selection. In other words, developers have been arguing about which brain to use, when the ears might be the limiting factor.

Frame rate is the concrete lever here. DualCodec runs at 12.5 Hz, the choice behind the TontaubeV1 preprint from September 2026, while VibeVoice's tokenizer runs even sparser, at 7.5 Hz BentoML / Best Open-Source TTS Models 2026 TontaubeV1 / ArXiv. Lower frame rates mean shorter token sequences for the same audio. That means less compute and more context room to work with BentoML / Best Open-Source TTS Models 2026 TontaubeV1 / ArXiv. At 12.5 Hz, a full minute of audio compresses to 750 tokens, tight enough that a model can hold onto a meaningful stretch of both text and audio history without blowing its context window TontaubeV1 / ArXiv.

That's not a footnote. Frame rate decides whether streaming is even feasible, how sharp the prosody can get, and whether residual codebook prediction runs fast or crawls. Anyone comparing Bark-style and VALL-E-style systems on LLM choice alone is skipping half the homework. The codec is a design axis in its own right, sitting underneath the "real" model.

Diagram: Codec Frame Rate: The Hidden Design Axis. Visualizes: Visualize how codec frame rate determines token economy, with concrete numbers from the article.

Language-model-backbone systems like Spark-TTS and the VALL-E paradigm's trajectory

Spark-TTS, published in 2025, is a prominent published example of explicit LLM-backbone TTS: it integrates Qwen2.5 as the language model, paired with a single-stream neural audio codec called BiCodec. The result reads like an efficiency argument dressed up as a research paper. That is not a rounding error; it is a statement about where the value actually lives, in codec design and data curation, not raw parameter count.

Controllability comes along for the ride. Spark-TTS uses chain-of-thought generation to steer both coarse attributes, gender, speaking style, and fine ones, exact pitch, exact rate, which sidesteps the need for a separate flow-matching model entirely. The VoxBox dataset, a 100,000-hour curated dataset with comprehensive attribute annotations, was introduced by the Spark-TTS team to enable this level of control Hyperstack / Popular Open-Source Text-to-Speech Models Spark-TTS / ArXiv Spark-TTS / Medium Ringlyn / AI Voice Synthesis 2026 BentoML / Best Open-Source TTS Models 2026. Good control needs good labels, not just a bigger model.

Other variants are pulling the same thread in different directions. P VALL-E, developed across 2025 and 2026, swaps Transformer layers for Performer layers through a layer-wise strategy, matches the original VALL-E's accuracy, runs about 20% faster, shrinks the parameter count for constrained hardware, and adds a language embedding for accent control. MoE-TTS takes a more conservative path, freezing the pre-trained text LLM entirely and bolting on speech-modality expert weights, preserving whatever text-world knowledge the base model already had, no retraining required.

The common thread across all of them is the structural advantage the VALL-E paradigm was built on. Every one of these systems inherits improvements to its backbone LLM automatically. Bark's audio-native design has no equivalent shortcut. Sharpening its semantic understanding means retraining the whole model, top to bottom. That asymmetry is quiet, but over a few product cycles, it compounds.

Hybrid two-stage architectures between the two paradigms

A third camp splits the difference. An autoregressive language model predicts semantic tokens, then a separate flow-matching model turns those tokens into acoustic features, two specialized stages instead of one model doing everything. FireRedTTS runs this pattern, structurally similar to Seed-TTS. So does the CosyVoice family. Spark-TTS is explicitly positioned against this camp: single codebook, pure language model, no flow-matching stage at all, described in its own literature as the tightest LLM integration published to date.

The trade is this: one process handles drafting, the other handles editing. It's the audio equivalent of separating drafting from editing.

Bark, oddly, sits closer to the hybrid camp than to VALL-E on this particular axis. Its fine acoustic token stages serve a similar refinement role to flow matching in the hybrid systems. Choosing a hybrid architecture means accepting that extra pipeline complexity in exchange for quality levers that can be tuned independently at the semantic and acoustic stages, which is either a feature or a maintenance headache depending on how many engineers are on call that week.

Latency profiles across the streaming and real-time spectrum

The industry-wide number is stark: end-to-end TTS latency fell from around 800 milliseconds to under 200 milliseconds across 2024 and 2025, a shift that amounts to a different product category rather than incremental progress Ringlyn / AI Voice Synthesis 2026 TontaubeV1 / ArXiv. That is not incremental progress; it is a different product category. Under 800ms total end-to-end is considered acceptable for voice agents to feel like natural interaction, and TTS streaming is essential for meeting that threshold The Neural Base / Voice Cloning Ethics.

Bark simply isn't built for that world. No streaming, an RTF of 0.85, a real VRAM appetite, none of it plays nicely with sub-800ms agent pipelines short of serious engineering workarounds The Neural Base / Voice Cloning Ethics. It's an offline instrument, suited to pre-rendered creative work The Neural Base / Voice Cloning Ethics.

The VALL-E lineage is where the streaming numbers actually live. TontaubeV1, a September 2026 preprint, achieves approximately 200ms to first audio on a single RTX 5090 via hierarchical codec modeling and bounded context. Its real-time factor is 0.08 for one input and drops to an aggregate 0.02 across eight concurrent streams.

Microsoft's VibeVoice-Realtime-0.5B lands audible speech in about 300 milliseconds with streaming text input, single-speaker, tuned deliberately for speed over polish Spark-TTS / ArXiv. The rule falls out cleanly: if the requirement is streaming with sub-200ms first audio, the LLM-codec paradigm is the only real option on the table, and Bark-style systems require accepting batch mode as the ceiling. TontaubeV1's total parameter count is 2.9B across four predictors (a 1.7B semantic predictor and three 0.6B acoustic refiners), demonstrating that LLM-backbone systems can hit sub-200ms streaming from a single consumer GPU.

Diagram: Latency Collapse: From 800ms to Under 200ms. Visualizes: Visualize the dramatic compression in end-to-end TTS latency across 2024–2025, anchored by specific named systems and their real-world numbers.

Voice cloning fidelity and expressiveness across the two paradigms

VALL-E's original party trick, cloning a voice from three seconds of audio while carrying over emotion and acoustic environment, still holds as the paradigm's signature move, done entirely through in-context learning with zero fine-tuning Microsoft / VALL-E-X. Fish Audio's S2 Pro, sitting in the same LLM-codec lineage, extends that into zero-shot cloning across more than 80 languages with genuine cross-lingual transfer, plus inline emotion control written as plain-language descriptions embedded at specific word positions. The benchmark numbers back it up: an 81.88% win rate on EmergentTTS-Eval, a 0.515 posterior mean on the Audio Turing Test, and word error rates of 0.54% in Chinese and 0.99% in English on Seed-TTS Eval Spark-TTS / ArXiv.

Bark answers a different brief. Laughter, sighs, hesitations, background sound, all deliverable through inline tags, a form of expressiveness the VALL-E lineage doesn't replicate natively. Comparing the two head-to-head on "expressiveness" without specifying which kind of expressiveness is a category error.

The frontier both paradigms are racing toward now is instruction-following, not raw cloning accuracy. VoiceSculptor, published on arXiv in January 2026, hit state-of-the-art on InstructTTSEval-Zh by pairing instruction-based voice design with high-fidelity cloning, using retrieval-augmented iterative refinement to make attribute-level edits, pitch, rate, age, emotion, style. EmergentTTS-Eval, presented at NeurIPS 2025, is the current reference point for judging models on prosody and linguistic nuance using a model-as-judge setup, and it is the framework to apply when actually choosing between these architectures. MINT-Bench, arriving in 2026, adds a multilingual axis that word error rate alone was always going to miss.

Open-source licensing and hardware realities that constrain the architectural choice

Licensing narrows the field faster than architecture does. CodeSOTA's survey lists the commercially clear field as Kokoro (Apache 2.0, MOS 4.2, 82 million parameters), Fish Speech (CC-BY-NC-SA-4.0 weights, meaning non-commercial use unless a separate license is arranged, MOS 4.1, 80-plus languages), Dia, Parler-TTS, Bark, and Piper, all sitting under Apache 2.0 or MIT terms.

XTTS v2 is the cautionary tale in the group. Its CPML license blocks commercial use outright, and there's no path to a separate commercial agreement because Coqui shut down in January 2024. Whatever XTTS v2 can do technically, 17 languages, six-second reference cloning, is beside the point if the license is a dead end.

Kokoro deserves a second look for a different reason: at 82 million parameters, built on StyleTTS 2, it posts the highest MOS score in the entire CodeSOTA comparison, 4.2, despite being the smallest model on the list. It doesn't do arbitrary voice cloning, and that is the trade, but the RTF of 0.03 on GPU says a lot about what's possible when a model isn't trying to be everything at once. Piper goes further toward the edge: VITS/VITS2 exported to ONNX, an RTF of 0.008, running on CPU under 100 MB of RAM. It belongs to neither the Bark nor the VALL-E camp, and for latency-constrained hardware, it's the one to know about.

Microsoft's VibeVoice, released in April 2026, is at the research end of the spectrum, shipped with audible disclaimers and watermarking, limited to English and Chinese, with its 1.5B variant supporting a 64K token context for up to 90 minutes of continuous multi-speaker audio across four distinct voices. Bark carries the MIT license, fully commercial with no separate agreement, and offers a GPU speed-up of up to 2× on GPUs and up to 10× on CPUs versus prior versions, per the Hyperstack Cloud guide. The hardware asymmetry between Bark's 6 GB VRAM requirement and Kokoro's CPU-capable 82M parameter footprint illustrates that paradigm choice and hardware budget are inseparable decisions.

The commercial API landscape as a reflection of these architectural bets at production scale

The API and SDK layer is currently the fastest-growing corner of the AI voice market, and it's where these architectural bets get tested against paying customers rather than benchmark leaderboards. Fish Audio's S2 Pro hosted API achieves approximately 100ms time-to-first-audio using a SGLang-based streaming engine Hyperstack / Popular Open-Source Text-to-Speech Models Spark-TTS / ArXiv. Cheaper inference, faster time-to-first-audio, and a licensing story clean enough to build a product roadmap around, all flowing from the same underlying decision made back at the "codec vocabulary versus audio-native tokens" fork. The architecture chosen at the whiteboard stage is the same one showing up on the invoice. Fish Audio's S2 Pro hosted API costs approximately $15 per 1M characters at approximately 100ms TTFA, compared to approximately $165 per 1M for a leading proprietary alternative, per BentoML.

Sources

  1. 6 Popular Open-Source Text-to-Speech Models in 2026
  2. TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
  3. The Best Open-Source Text-to-Speech Models in 2026
  4. P VALL-E: An efficient multilingual speech synthesis system based on the performer architecture - Shih-Hsiung Lee, Yu-Hsiang Chang, Chu-Sing Yang, 2026
  5. VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
  6. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
  7. VoiceSculptor: Your Voice, Designed By You
  8. MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models

More in Voice AI Research