Liquid AI has released two open-weight models that do not write anything. The d1 family — d1-3B and an experimental d1-omni-600M — answers structured questions about a given input in a single forward pass, returning calibrated, typed answers instead of a stream of generated tokens.

The pitch is edge computing: instead of a chat model, you hand d1 a "state" (text, JSON, an image, or a mix) and a set of named questions with constraints — a yes/no, a choice between labelled options, a score on a scale — and get back one decision. Because nothing is generated, there are no output tokens to pay for and no decoding loop to wait on.

The published numbers aim squarely at that use case. d1-3B scores 48.57 on the Decision Index v0.2.1, ahead of every 4B and 9B model tested and of Decider 35B-A3B (47.11). Across seven public datasets covering reading comprehension, toxicity detection, intent classification, medical QA and cross-lingual understanding, it averages 82.9 — the best in the comparison table. The 600M model, built on a bidirectional 350M encoder, averages 78.4 with a quarter of the parameters of the Decider 2B it beats, and adds audio input alongside text and images.

Speed is the headline: in work with NVIDIA, d1-3B answered a question in 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin 64 GB and 50 ms on an Orin Nano; on an Apple M5 Pro the figure was 30 ms. On an RTX 4090 it takes 8 ms, and on an AMD MI325X 9 ms. Three questions cost about 1.3x the time of one. Both models are on Hugging Face with native llama.cpp support across Apple, AMD, Qualcomm and NVIDIA hardware.

Two caveats deserve stating plainly. Liquid reports no vision or audio benchmarks for this release — it says the Decision Index's vision split is private and that audio decision benchmarks are an open problem — and d1-omni-600M is an early research release with no published speed figures. The company is also explicit that these are decision models, not general assistants: ask one to write you a paragraph and you have the wrong tool.

That narrowness is the point of the release. If a classifier, router or triage step is what a pipeline actually needs, a model that emits one typed answer in one pass is cheaper and simpler than a generative model asked nicely not to ramble.