top of page

Designing for Porosity: The Missing Layer in Voice AI (Part 1)

Voice AI keeps getting smarter. It keeps sounding wrong. The problem isn’t the script — it’s the invisible layer we haven’t designed for yet.


There is a particular kind of frustration that comes from sitting in a room full of smart people trying to design voice AI by feel.


The engineers are capable. The dialog flows technically work. And still, something is off — the system sounds cold, or slightly robotic, or wrong in a way no one in the room can precisely name. So the conversation circles. People adjust the script. Someone suggests a different word choice. The problem doesn’t move.


I came to this problem through my work as Head of Design at Byte by Yum! — the AI platform behind some of the largest restaurant brands in the world. Voice AI isn’t hypothetical in my world. It’s on the table right now, in real products, with real stakes.


And I think it’s about to be on everyone’s table.


Design leaders are already deep in the AI conversation — copilots, generative interfaces, AI-assisted workflows. That’s where most of the energy is. But voice AI is the one that will sneak up on us. Because it doesn’t look like a design problem until it already is one. There’s no screen to critique, no layout to redline. By the time it feels wrong to a user, the system is already in production — and no one can quite explain why.


The issue is never the words. It never is.



The Problem Isn’t the Script


For decades, design has codified standards for visible interfaces. Grids, color systems, typography, accessibility rules. These frameworks exist because practitioners recognized that shared principles — not individual intuition — were what separated craft from guesswork.


Voice AI has no such foundation.


Most implementations are built around dialog flows and intent classification. That is necessary. It is not sufficient. To understand why, try a simple demonstration. One sentence.


Five words.


“I didn’t take her wallet.”


Say it out loud. Now say it again, stressing a different word each time.


Stress I — and you’re denying personal involvement. Someone else did it. Stress didn’t — a flat denial of the act itself. Stress take — maybe you borrowed it. Stress her — it was the ownership that mattered, not the act. Stress wallet — you took something, just not that.


Same five words. Five completely different meanings.


This is the hidden physics of conversation — the way tone, lilt, and timing change intent before a single syntactic element shifts. And this is exactly where most voice AI fails. It hears the words. It misses the meaning.


Most teams respond to this problem by rewriting the script. More words, better words, cleaner words. But the script was never the problem. The problem is that we’ve been designing for language when we should be designing for something that exists beneath it.


What Porosity Is


The word I’ve landed on for that something is porosity.


Porosity is the patterns of rhythm and sound that carry meaning between and beneath the words. It is the invisible layer of conversation — the part that operates before language, beyond language, and sometimes instead of language entirely.


Humans are wired to parse meaning from tone and timing automatically, faster than conscious thought, across every language and culture. We hold entire conversations without words — just sounds. A sigh signals frustration. A gasp signals surprise. A long pause signals something is wrong. None of that requires a sentence. The meaning arrives before the language does.


Think about the last time you answered the phone and knew, from a single “hello,” that something was off. No content. No context. Just a sound — and you knew.


That is porosity. And most voice AI systems are completely deaf to it.


The question for designers is not how to write better scripts. It is how to teach a system to operate in this invisible layer of sound and silence. That is a fundamentally different design problem. It requires different inputs, different metrics, different prototyping methods, and a different vocabulary for talking about what we’re building.


What We’re Actually Designing For


When I gave this talk at the Design Leadership Summit, I walked through what designing for porosity actually means in practice. It is not one thing. It is four.


Lilt. The musical quality of a phrase — the rise and fall that makes a sentence feel like a question, a reassurance, or a command even before the words land. Lilt is what makes a customer service voice feel warm versus procedural. It is entirely absent from most voice AI outputs, which tend toward a flat, declarative rhythm that feels efficient and nothing else.


Silence. The gaps between turns that human brains read as social signals. A pause of the right length signals thinking. A pause that runs too long signals hesitation, confusion, or rejection. Silence is social glue — until it becomes social distance. In voice AI, silence is almost always treated as a technical problem to minimize rather than a design variable to calibrate.


Stress. The weight placed on individual words within a phrase. As the wallet sentence shows, stress is not decoration — it is meaning. A system that produces correct words in a flat, unstressed delivery is not communicating. It is reciting.


Tone. The acoustic quality that tells someone, within seconds, whether they are trusted, rushed, heard, or dismissed. Tone operates below the level of content. It is what users are responding to when they describe a system as “cold” or “robotic” — and it is almost never explicitly designed.


Without porosity, a voice AI can’t hear what we really mean. It can only hear what we literally said. And without designing for porosity, a voice AI can’t sound like it understands — regardless of how accurate its responses are.


Why This Is a Design Leadership Problem


It would be convenient if porosity were purely a technical problem — something for NLP engineers to solve with better models. It is not.


The decisions that determine whether a voice system feels human are design decisions: what emotional signals to detect, how silence should be calibrated, what tone the system projects in different contexts, how recovery should sound when something goes wrong. These require design judgment, design standards, and design governance. They require someone in the room asking not just “does it work?” but “does it feel right — and do we know why?”


Right now, most organizations don’t have that person in that room.


Voice AI is moving fast. The teams building these systems are capable and well-intentioned. But without a design framework for the invisible layer, they are optimizing for accuracy while the user experience degrades at the level of relationship. Users don’t file bug reports about porosity. They just stop trusting the system. They call it robotic. They disengage. And by the time the data surfaces, the damage is done.


Design leaders need to get ahead of this. Not because voice AI is new — it isn’t — but because the standards for designing it don’t exist yet. We have decades of practice designing for screens. We are starting from scratch on sound.


What Comes Next


Understanding porosity changes how you see voice AI failures. The script isn’t the problem. The flows aren’t the problem. What’s missing is a set of governing principles for the invisible layer — for the physics of spoken interaction that operates whether your system accounts for it or not.


In my next article, I’ll introduce the framework I developed to address that gap: The Five Laws of Invisible Conversation. These are not tips or tactics. They are the structural conditions under which human conversation functions — grounded in linguistic, psychological, and conversational science. They are, as best I can define them, the laws that govern the invisible layer.


Once you understand porosity, you can start to name what it requires.


That’s where we’re going.


Jessa Parette is Head of Design at Byte by Yum! and a speaker on Voice AI, design systems, and the future of invisible interfaces. This piece is adapted from a talk originally presented at the Design Leadership Summit, 2025.

 
 
 

Comments


bottom of page