Something shifted in how companies work over the past two years, and the software has not caught up.
Teams came back to offices. Some returned by choice, some by mandate, and some through a messy hybrid compromise. Amazon, JP Morgan, and hundreds of smaller companies pushed return-to-office policies through 2024 and 2025. Those policies did not create the whole shift, but they are part of the backdrop. For many teams, the default meeting has moved from Zoom back into a room.
But the meeting stack was built during a different era. These are the tools companies rely on to capture, summarise, and act on what happens in meetings. They were designed around one central assumption: the meeting is a scheduled video call, every participant is logged in separately, and software can join or capture the call as part of the workflow.
That assumption made sense when it was built. It describes fewer meetings now.
The meeting stack was built around the video call
Look at the major AI meeting tools and you can see how strongly the category was shaped by the video-call era.
Otter, Fireflies, Fathom, Granola, Read.ai and similar tools have each taken a different path. Some join meetings as visible assistants. Some capture system audio without a bot. Some now offer mobile or in-person recording modes. The details vary, and the products are improving quickly.
But the centre of gravity is still clear: the modern AI notetaker grew up around the scheduled video call.
That architecture solved a real problem. When teams went fully remote in 2020, meetings became video calls by default. The audio was usually captured from individual devices. Participants were tied to accounts. Calendar links created a clean workflow for joining, recording, transcribing, and summarising.
The tools that emerged from that period are genuinely good at remote meetings.
The gap appears when the meeting moves into a room. A room has multiple people, shared microphones, side conversations, informal starts and stops, and social dynamics that make visible recording software feel heavier than it does on Zoom.
Return-to-office changed the meeting environment
The default meeting has changed.
For many teams, a typical week now includes a mix: video calls with remote colleagues, in-person meetings in conference rooms, hybrid sessions with people both in the room and on the call, and informal conversations that never get a calendar invite.
The software handles the first category well, then struggles with the rest.
In-person meetings are captured poorly or skipped entirely. Hybrid meetings create a two-tier documentation problem, where remote participants get clean transcripts and in-room participants get lumped together or missed. Informal conversations, such as the hallway follow-up, the five-minute debrief after a client visit, or the leadership discussion before the official meeting starts, disappear by default.
The stakes are higher than convenience. These are often the most consequential conversations. The decisions that shape strategy, resolve disagreements, and commit resources frequently happen outside of a formal Zoom invite.
The current meeting stack was built around the scheduled video call. Modern work has moved beyond that boundary.
The bot assumption breaks down
Underneath the technical gap is a cultural one.
The bot model works reasonably well in a fully remote context. A piece of software joins the meeting as a visible participant, alongside everyone else. Everyone is already on a screen. Another participant icon in the corner usually blends into the workflow. Consent is also easier to establish at the start of a call.
In a physical room, the same model feels different.
Some conversations are sensitive. Some are internal and exploratory. Some involve people thinking out loud before anything is formalised. Some happen in rooms where recording infrastructure changes the nature of the conversation. Setting up a device, launching a bot, or explaining that the room is being recorded can make an exploratory discussion feel formal.
You can see versions of this complaint in Reddit threads where teams try to solve the problem. People want notes without another bot joining the meeting. They want something that works offline. They want local recording that stays private until they choose to share it. They want something that does not make a sensitive internal conversation feel like a formal deposition.
Those examples point to something important. The bot-first model was designed for a specific context, and that context is no longer universal.
The room has become the blind spot
Even when teams do try to record in-person meetings, the results expose another gap: the meeting stack was not built for room audio.
The standard workaround is a speakerphone, a room mic, or a phone placed in the centre of the table. This captures the audio, but speaker attribution remains the harder problem.
On a video call, speaker identification is straightforward. The platform knows which microphone is active because each participant is connected separately. On a room recording, everyone speaks into the same device. The AI has to work out who is talking from a single mixed stream.
That is a much harder problem. Diarisation means splitting a mixed audio recording into separate speakers. It works in controlled conditions. In a real conference room with six people, overlapping conversation, and ambient noise, it degrades quickly.
The result is a transcript that may be roughly accurate about what was said, while remaining unreliable about who said it.
Hybrid meetings make this worse. Remote participants, each on their own microphone, show up as distinct speakers. The people in the room, sharing one device, often collapse into a single anonymous voice. The people who physically showed up can end up with worse documentation than the people who stayed home.
The room has become the blind spot in a stack that was supposed to capture everything.
What an office-native meeting stack needs to do differently
Fixing this requires rethinking the capture model, rather than only improving transcription.
The video-call model works because audio is already per-person before any AI gets involved. An office-native model needs to recreate that property in a room. The way to do it is distributed capture: each participant in the room contributes their own audio stream from their own device, rather than sharing one room microphone.
Per-person capture changes what the system knows. Instead of guessing who is speaking from a mixed recording, the transcript can attribute each line to the person whose device captured it. Speaker labels no longer depend on acoustic guesswork alone. Attribution comes from the capture model itself.
Beyond capture, an office-native stack also needs to handle the contexts where a bot can feel heavy or awkward: informal conversations, sensitive internal discussions, hybrid rooms where in-room and remote participants need equal documentation, and meetings that happen outside of a formal calendar invite.
That means working without a bot dependency. It means recording locally when privacy matters. It means being usable in a room with no built-in AV, no conference room system, and no Zoom link.
The next meeting stack will be built for the room
The tools that defined the last era of meeting software were right for their moment. They captured what needed to be captured when work was remote and meetings were video calls.
That moment has not disappeared. Remote work still exists. Video calls still happen. The tools built for them still have a place.
But work has changed shape. The office returned. Hybrid rooms became normal. Important conversations moved back into physical spaces where single-recorder workflows struggle.
The next meeting stack needs to start with the room: multiple people, uneven audio, sensitive discussions, informal conversations, and a real need to know who said what afterwards.
Roundtable is a meeting notetaker built for that environment: distributed audio capture, per-person attribution, no bot required, and designed for the conversations that actually matter, whether they happen on a call or around a table.