Lukáš Chudý
I looked at a hundred startups trying to improve how humans interact with AI. How will we communicate with agents once they break free from the text chat?
-JRgQ.jpg)
Issue #120
My main thesis is simple: as AI agents become more capable, interacting with them will start to feel less like using software and more like communicating with another person.
We rarely rely on words alone. We combine speech with eye contact, gestures and touch. Much of that communication happens almost automatically, and we only give something our full attention when a decision needs to be made or something unexpected happens.
Startups building voice rings, silent-speech interfaces, smart glasses and haptic devices are each creating one piece of this future interface. The really interesting part begins when someone connects them.
First, AI needs to leave the chat window
Voice is the obvious first step. When you are driving, walking or cooking, saying what you want is much easier than pulling out your phone and typing it. Voice is also useful because new tasks often need context: what you are doing, why you are doing it and what matters to you.
Sandbar moves the microphone from your phone to your finger. Hold the ring, speak a thought, and a short vibration confirms that it has been captured. No unlocking your phone, opening an app or shifting your attention to a screen.
But voice also has an obvious problem: everyday life.
I do not want to dictate a private message on the subway, and an open-plan office where 20 people are constantly talking to their agents would be unbearable. If AI is going to stay with us throughout the day, it needs a private communication channel too.
What if we could speak without making a sound?
Silent speech is one of the most interesting parts of this entire space to me. It could preserve the natural expressiveness of language while allowing us to give complex instructions around other people without saying everything aloud or typing it into a phone.
Augmental is approaching this through two devices. MouthPad sits against the roof of the mouth and lets users control a computer with their tongue. VOX picks up speech through the body, already enabling extremely quiet dictation. The company plans to combine the two signals, learning the relationship between sound and tongue movement while someone speaks, with the long-term goal of recognizing words even without sound.
MIT’s AlterEgo explores another approach. It detects subtle neuromuscular signals generated when a person intentionally verbalizes words without speaking them aloud. The research system sends responses back through bone conduction, making the entire interaction feel almost like talking to yourself. Importantly, it detects deliberately formulated words rather than arbitrary thoughts.
Imagine leaving a meeting and silently reviewing the outcome with your AI as you walk. You ask about a connection you noticed, have it prepare a follow-up, and adjust the rest of your afternoon. Your phone stays in your pocket and nobody around you hears anything.
If the agent sees what we see, we need far fewer words
Even with a private voice channel, we will not want to describe everything explicitly. If I am sitting at a table, look at a glass and say, “Pass me that, please,” another person immediately understands what I mean. Glasses with eye tracking could let an agent interpret instructions in the same way.
A small movement could then be enough to confirm an action. The Meta Neural Band reads electrical signals from muscles in the wrist and translates subtle gestures into digital commands. Haptics could work in the opposite direction, telling us that something is done or that the agent is waiting for input.
Over time, this could create a simple communication grammar:
Voice, text or silent speech: a new task, context, or change of plan.
Gaze: what the instruction refers to.
Subtle gesture: yes, no, stop.
Haptics: done, waiting, warning.
Screen: evidence, comparisons, and confirmation of important actions.
That is why I think voice usage will grow significantly at first, but over time we may need it less. Agents will receive more context through other channels and eventually learn our personal shortcuts, gestures and habits. They will begin to understand a language that developed specifically between us.
The more interesting question: when should the agent say nothing?
Once an agent can receive instructions and respond privately, the next question is not how it should communicate. It is when it should interrupt us at all.
Most of the signals around us never reach our conscious attention. We notice changes, uncertainty and things that matter. Everything else fades into the background. A personal agent could work in much the same way. It could learn what is normal for us and recognize when something genuinely requires our judgment.
Imagine it handling a delayed delivery or a meeting that gets moved, as long as everything stays within the limits you have already set. If the outcome is still fine, there is no reason to send another notification. But if that meeting change suddenly puts your airport departure at risk, then it should interrupt you. One of the most valuable capabilities an agent can have will be knowing what to handle quietly.
For this to work, the agent will need memory that follows us across devices. A ring, glasses, headphones, car and home will each provide a different part of the picture. That creates an opening for one primary agent that carries your goals, preferences and boundaries across all of them, something like a personal Apple ID, but with a much deeper understanding of you.
Who gets to hold the full picture?
This is where the next platform war begins. Google has a strong position in our digital lives. Apple controls the devices around us. Meta knows our relationships and communication. Microsoft sits inside our work. Amazon knows our shopping habits and increasingly our homes. Each of them already holds a different part of the context an agent would need.
The advantage will go to the system that can see enough of our lives to distinguish routine from exception and most importantly, that we trust enough to act on that judgment.
But that convenience comes with a cost. I can reject a recommendation
but I cannot reject information the agent decided not to show me.Whoever decides what deserves our attention gains enormous influence over us. That means we will need ways to look back and understand what the agent handled, what it filtered out, why it made those decisions and whose interests shaped them.
What might an ordinary day look like in 2036?
In the morning, my agent does not brief me because nothing important has changed. On the subway, I silently tell it about a change of plan and talk through a new idea. At the office, I look at a contract, ask it to explain one clause, and approve the preparation of a reply with a gesture. Before anything is sent, I review it on a screen.
Throughout the day, the agent quietly handles routine tasks. A small haptic signal tells me when something is done. If it has a question, it waits for a natural break in my attention instead of interrupting me immediately. At night, it consolidates what it learned during the day, forgets details that no longer matter and prepares one question for the morning because it has noticed a conflict between two of my goals.
My role gradually shifts from telling the agent exactly what to do toward setting goals, defining boundaries and dealing with exceptions. Words will still matter. I just will not need them to control everything.
Who would you trust with an agent like this: Google, Apple, one of the AI labs, or someone entirely new?
Share
(Press Release) Two leading Czech search engines from Miton's portfolio, fashion-focused Glami and furniture-focused Biano, are entering a new chapter. As of March 2026, both companies are led by a single CEO, Peter Hupka, who previously headed Biano. Both companies hold normalized data on millions of products data that large language models lack and aim to build on this foundation to develop a new generation of online product discovery tools.
The second half of the year, like the whole year (and the two previous ones as well), was marked by AI. This year will be no different. There will be a lot of AI in this summary too. But we’ll also add a bit of crypto and psychedelics.
(press release) A startup founded by Johanna von der Leyen and Marek Miltner at Stanford is changing the way companies and public institutions work with geospatial data. Instead of needing to hire a team of GIS experts, PangeAI agents allow making complex analyses and decisions involving physical infrastructure as easy as typing a prompt. The goal is to make the physical world as searchable and understandable as the digital one.
