In The Flow

Your mom was right.

You’re a lazy communicator.

Oftentimes you’ll say something and mean another.

When you have those moments of “I get what you mean,” that’s where the magic is.

Or in some cases, where the magic is not, like with voice agents like Alexa back in 2014 up to, well… two years ago.

Because of the lack of context, understanding, etc., your interactions were limited to “what’s the weather outside” and “what’s my next meeting,” which worked. What didn’t work was anything with any level of nuance or anything outside of the template, e.g., “Move my 3pm to tomorrow and tell everyone I’m running behind on the deck”: two actions plus a vague reason, and no template exists.

Wispr Flow is one of my most used apps. Today it’s primarily used for speech to text, which is awesome.

But in the vein of this new wave of AI, you get really interesting questions about what it looks like to work with intelligent systems that have the ability to get to those “I get what you mean” moments, and to deliver on the magic there.

There’s an interesting blog post about how they’re thinking about that from their CSO: Introducing Wispr Advanced Interfaces Lab. The multimodal hook is what I had fun thinking about after reading this post. Imagine I’m walking down Broadway near 31st and I say, “Where’s a good coffee spot?” Location preferences can infer proximity and recommend something close by, but what they can’t do is notice the line out the door of the place I’m facing, the laptop bag on my shoulder, and the rain starting, and then answer the question I actually meant.

Voice carries the intent, and vision and context fill in the rest of the meaning.

Leave a comment