Why We Backed Wispr: The Interface That Finally Listens
By Aliisa Rosenthal, Lauren Kolodny, Tom Porter, and Kunaal Patel
Most of us think faster than we can type. A thought arrives whole, in paragraphs, and what makes it onto the keyboard is fragments. We’ve quietly paid that tax for decades, because the alternative never worked: voice broke the moment you had an accent, a noisy room, or a sentence that did not sound like it was read off a script.
Which is why our office looks a little different lately. Walk through on any given day and you’ll catch at least one of us quietly whispering into a device. Not on a call, not a voice memo. Drafting an email, a Slack, a doc. It looks a little ridiculous. It’s also the first time we’ve watched software push people back toward talking instead of typing.
The speech recognition industry spent a decade optimizing for the wrong number. Word accuracy: 85%, 88%, 90%%, each incremental point celebrated as progress. But nobody dictates a message and checks it against a ground-truth transcript. They check whether they can hit send. A system that transcribes someone’s words perfectly, false starts and mid-sentence corrections and all, still hands back text that needs cleaning up. The number that actually matters is how many messages need editing afterward, and getting it down requires context most transcription models never see: who you are writing to, what thread you are replying in, how you write when you are formal versus quick, which words you always fix. Voice becomes the default input not when the model gets a little better, but when the editing disappears.
That’s why we backed Wispr: it’s the first company to treat the editing step, not the transcription step, as the problem worth solving.
The team: they have been circling this problem for a long time
Tanay Kothari built his first voice assistant at nine years old, inspired by Iron Man and convinced that talking to computers should not be science fiction. That project, Evi, scaled to millions of users in its first year, before Siri or Alexa existed. He kept going anyway: a Stanford AI research background, a first company (FeatherX, acquired by Cerebra), and eventually a decision to start over from scratch. He built Wispr with Sahaj Garg, his roommate from the first day of Stanford undergrad, who spent his own academic years publishing early research in diffusion-based generative modeling under Stefano Ermon. Before writing a line of code together, the two spent two months talking through their values. That is not typical founder behavior, and it shows in how deliberately they have built since.
Since then, Wispr has recruited a leadership team that reads like a list of the people building the most important speech and multimodal systems anywhere: Chief Scientist Ariya Rastrow, a founding member of the Alexa team who most recently led multimodal foundation model work at Meta; VP of Engineering Cliff Chang from Asana; and a Chief Revenue Officer who built out enterprise motion at Watershed and Procore. That is a rare density of senior operators for a four-year-old company, and exactly the kind of team we look for when betting on a category this early.
The product: the numbers reflect a habit, not a novelty
Wispr’s product, Flow, works everywhere we already type, across essentially every application and website, with no integrations to configure. What makes it stick isn’t the transcription. It’s everything that happens after: Flow reads the thread you’re replying to, understands what’s already on the screen, adapts to how you write, and improves every time you correct it. More often than not, the output goes out without a cleanup pass, and that trust is what turns Flow from an occasional shortcut into a daily habit.
That habit shows up in the numbers. Flow is now used by millions of people around the world, including employees at the majority of Fortune 500 companies and more than 10,000 enterprises total. Meanwhile, revenue has grown more than 150% in each of the last four quarters. The growth has an unusually organic shape, and we watched it happen inside our own walls: several of us were paying for Flow out of pocket before we even thought about the company as an investment opportunity. One person skeptically tried it, and before you know it, half the team was Wispr-ing. Before long, typing felt inefficient by comparison. Anyone who has used the product knows that after a week of using Flow, your fingers start to feel like unnecessary middlemen.
That same everyday use is what compounds into a moat. Every opted-in dictation shows Wispr something a clean recording never could: what a real person was trying to say, in the app and thread where they said it, using the nomenclature of their profession, and how they fixed the output the model guessed incorrectly. Clean speech recorded in a quiet room is easy to buy. Real-world speech paired with context and user corrections is not. The signal isn’t the recording, it’s the correction, made in context, during real work.
The timing: the models finally cleared the bar
Voice has been the obvious interface for a long time and a bad product for just as long. What changed is that the models finally cleared the bar where the editing step gets short enough to trust, and the first company across that line gets to build the data advantage that keeps it there. Wispr has been attacking this from both ends: it started in 2021 on hardware, building toward silent speech, and moved to software when it became clear the problem could be solved there first. That is not a team chasing a trend. It is a team that has lived with this problem long enough to know where it actually breaks.
We are proud to share that Acrew has joined Wispr’s $280 million Series B, led by Menlo Ventures, alongside Notable Capital, Forerunner, Goodwater, Peak XV, and new investors Together Fund and Plus Capital, valuing the company at $2 billion.
What comes next
With this round, Wispr is extending its lead. The new Wispr Advanced Interfaces Lab, led by Rastrow, has already unveiled Wispr’s first proprietary speech model, cutting word error rates in the hardest real-world conditions, background noise, wind, and strong accents, from more than 30% down to 5-10%. That matters because every edit falls into one of two buckets: words the model misheard, and everything context should have gotten right. Flow already solves the second. The new speech model attacks the first. The team expects the combination to mean roughly a third fewer dictations needing any edit at all. As Tanay puts it, “dictation was always the starting point for something bigger.”
That something bigger is the real prize. Wispr has already built the engine that knows who you are writing to and what is on your screen, so now it can begin to do much more. Its interface that understands what you mean and what you want to do creates the interaction layer between people and technology. Dictation was the entry point, and for the first time the underlying technology is finally good enough to build the rest.
Every earlier attempt at voice made you talk like a machine. Wispr built the first one that listens like a person.
To Tanay, Sahaj, and the entire Wispr team: welcome to the crew.






