Pika Labs, best known for AI video generation, is moving into audio. On August 14 the company launched Pika Audio Models, a family of four "frontier foundation sound models" spanning the generative audio spectrum: Soundtrack, Music, SFX and Speech.

Pika Soundtrack turns a video into a synchronized native soundtrack — motion-aware sound effects, music, ambience and voiceover that follow what happens on screen. Pika Music turns prompts, lyrics, voice references and reference tracks into complete songs of up to six minutes. Pika SFX generates clean, prompt-faithful sound effects from natural-language direction, and Pika Speech is an expressive text-to-speech model with preset voices or a cloned voice trained on a few seconds of reference audio.

The headline claim is price. Pika says the models are "the lowest-priced audio models on the market — up to 20x cheaper than other audio models," thanks to highly efficient training and inference techniques, few-step generation and a fast inference stack. Concretely: Soundtrack is 2x more cost-efficient than Hunyuan Foley, Speech is 9x more cost-efficient than ElevenLabs v3 (and 4.5x vs. Cartesia and ElevenLabs Turbo), and Music is up to 10x more cost-efficient than comparable music models.

The company also published speed and quality figures: Soundtrack covered 529 seconds of video at 0.617 seconds of wall time per generated second, with the strongest semantic alignment and lowest audiovisual desynchronization in its benchmarks; Music generated a 90-second song in 6.21 seconds on average, about 14.5x faster than playback; SFX completed an end-to-end generation in 0.847 seconds on average; and Speech runs at a real-time factor of 0.02 — a minute of speech in about a second.

The models are available now through the Pika API Club, the company's developer platform. Pricing comparisons are current as of August 14, 2026, based on providers' official list prices.