Wispr has raised $280mn at a $2bn valuation. Menlo Ventures led the round, having led the last one too. Total funding now stands at $361mn.

The company makes Wispr Flow, a dictation app that turns speech into cleaned-up text in whatever application your cursor happens to be in. Reuters reports the previous round valued it at $700mn, in November. That is close to a tripling in nine months.

The pitch is that the model layer is finished

Menlo published its reasoning the same day, and it is unusually blunt.

“The frontier labs have produced superhuman AI intelligences. They have not produced delightful and effective human interfaces,” wrote partner Matt Kraning. He argues the bottleneck has moved from the model to the interface.

He put the commercial version of that to Fortune more directly.

“It isn’t a dictation market. Dictation is how you get in the door,” Kraning said. “The labs have mostly solved intelligence. Nobody has solved how a normal person tells it what they want.”

Menlo also calls this one of the largest AI investments it has made.

It is not the only firm making that argument this week. In Aachen, a startup called amber raised €7mn on the near-identical claim that the competition for models is more or less over, and that what matters now is the context you feed them. One round is forty times the other. The thesis is the same.

The error rate in the announcement

Wispr used the funding day to preview Canto, its first proprietary speech recognition model. The detail inside that announcement is worth reading slowly.

“In the hardest conditions, with background noise, wind, heavy accents or music, error rates fall from more than 30% of words to somewhere between 5% and 10%,” chief executive Tanay Kothari wrote.

Read that backwards. In difficult conditions today, before Canto ships, the product gets more than three words in ten wrong.

Users noticed. TechCrunch reported that several complained about a dip in output quality over recent weeks. Canto arrives as an answer to that as much as a milestone.

The sceptical read, from someone who does this for a living

Wispr says revenue has grown more than 150% in each of the last four quarters. Menlo puts it at over 30 times year on year.

Michael Ashley Schulman, a partner at Cerity Partners, gave Reuters the caution that belongs next to those numbers.

“Triple digit quarterly growth off a small base is the easiest number in venture capital to generate for a few quarters and the hardest one to keep producing once the early adopters stop being the entire customer base,” he said.

That is the whole question in one sentence, and Wispr is not the only company it applies to. Higgsfield raised $400mn yesterday on a run rate that went from $20mn to $700mn in a year. Both rounds price a trajectory rather than a business.

Founders Tanay Kothari and Sahaj Garg met in a Stanford freshman dorm and started the company in 2021. They spent years on wearables and on silent speech, an interface for controlling a computer without making a sound. Neither worked. Flow arrived about two years ago.

How many customers, exactly

Here the published figures do not agree, and the gap is large.

Fortune says 100,000 businesses. Menlo says tens of thousands of paying businesses. Reuters says more than 10,000 enterprises. Those may be three different definitions rather than three different counts, but nobody has reconciled them publicly.

What is consistent: millions of consumer users, 162 countries, and more than 100 languages.

The company’s own site names Microsoft, Amazon, Notion, Klarna, Groupon, Rivian, Vercel, and Mercury among the employers where staff use it.

The competition is everyone

Apple, Google, Microsoft, Anthropic, and OpenAI are all working on voice input. Below them sit a crowd of smaller dictation apps including Willow, Monologue, Aqua, and Superwhisper, plus free tools aimed at the same users.

Menlo answers that objection head on, which is a sign the firm expects it. Free dictation has existed for a decade and almost nobody uses it, Kraning argues, because a feature can transcribe but acting on intent takes a company.

Wispr charges nothing for 2,000 words a week, then sells Pro, team, and enterprise tiers. It certifies to SOC 2 Type II, ISO 27001, and HIPAA, and says Privacy Mode keeps dictation out of its training data.

Investors in the Wispr Series B include Notable Capital, NEA, Neo Ventures, 8VC, Acrew, Forerunner, Goodwater, Peak XV, Together Fund, and PLUS Capital. Fortune reports that the athletes Joe Burrow, Shaun White, Klay Thompson, and Paul George also took part.

What it wants to be instead

Dictation is not the endpoint anyone involved describes. In July the company opened a research lab under Ariya Rastrow, who worked on Amazon Alexa in its early years.

Rastrow’s diagnosis, as Menlo relays it, is that the previous generation of voice assistants failed on intelligence, and that the speech-to-text-then-model pipeline everyone still uses is now the wrong architecture. The lab is meant to replace it with something Menlo calls voice-to-outcome.

Wispr has also shipped a meeting notetaker, which puts it against Granola, Fireflies, and Read AI.

Co-founder Sahaj Garg framed the trust problem the product creates for itself. “Privacy and security aren’t features for us: They’re the foundation for earning ambient access to your life,” he told Fortune.

What would settle it

Voice funding has been busy all year. Bland raised $50mn in June, and Fish Audio took $52mn in July.

The furthest along is ElevenLabs, at roughly $600mn in revenue, which is the number Wispr is being priced against whether anyone says so or not.

Two things decide whether this round looks cheap or silly in a year. The first is whether Canto actually closes the error gap the company just published, because a 30% miss rate in noisy conditions is not a product people build habits around. The second is Schulman’s point: whether the growth holds once the customers are ordinary rather than enthusiastic.

Get the TNW newsletter

Get the most important tech news in your inbox each week.