The business of supplying the data that AI models are trained on just got a new valuation high-water mark. According to TechCrunch, Snorkel AI has tripled its valuation to $3.5 billion, a jump the company ties to booming demand for AI training data.
The timing is not accidental. Attention has shifted from pre-training to everything that happens after it: fine-tuning, reinforcement learning from human and executable feedback, evaluation datasets, and the tooling that systematically finds where a model fails. All of that is data-hungry and cannot simply be scraped from the public web, because it is about proprietary, domain-specific examples — legal filings, clinical notes, invoices, sensor logs, transcripts of solved customer cases.
For enterprises the uncomfortable implication is that the gap between a model that demos well and one that reliably does a job is rarely about architecture and almost always about data quality. Snorkel's raise reflects a bet that a meaningful share of that work becomes an infrastructure product rather than a bespoke project repeated inside every company.
The sector is competitive and unusually sensitive. Anonymisation, provenance and licensing questions now appear in procurement reviews, and buyers are wary of vendors that retain or reuse their data. At the same time, the volume of demand pushes suppliers to replace manual labelling with programmatic generation verified by humans, which shifts both throughput and margin.
What to watch: whether this valuation is supported by revenue or by expectations, which customer segments actually drive the growth, and whether the category consolidates into a durable tooling layer instead of a low-margin services business. Details of the round, its investors and its terms were not included in the reporting available at the time of writing.




