Technology

Vojta Ciml at Ai4 2026: How a Conference Room Broke Our Computer Vision Assumptions

5 min read

Close your eyes and try to tell where the speaker is standing, just by listening. That's how Vojta Ciml, SlidesLive Founder & CEO, opened his talk at Ai4 2026, and in a room full of loudspeakers it's much harder than it sounds. It was also the first of five assumptions that real conference rooms broke while we were building robotic cameras.

📢 Watch the full talk ›

A Production Crew in a Suitcase

SlidesLive records talks at conferences around the world: 120,000+ talks, 4,000+ events, 14+ years. That includes Ai4 itself, where our robotic cameras ran in every track.

Our mission is to record every presentation and preserve it for future generations, and the obstacle has always been cost. Video production is labor intensive and gear heavy, with plenty of strange cables and boxes. So we asked a simple question: what if everything needed to record a session fit in a single case?

That question became RoboCam, our robotic camera system built around a Canon CR-N500 PTZ camera and run by our software NATE.

Assumption 1: Sound Tells You Where the Speaker Is

On a panel with five people, tracking faces isn't enough. The camera has to know who is actually talking.

Instinct says we can hear where a voice comes from. On a stage we can't. The sound arrives from the loudspeakers, not from the speaker's mouth, so a microphone array tries to locate the voice points at the PA system instead of the person.

That's the easy case. A conference stage is not a sealed Hollywood studio. Some sit next to the expo hall, some share a wall with a live band, and what the camera microphone picks up is close to unusable.

So we stopped relying on a single signal. NATE uses audio for speaker diarization, recognizing each voice so it knows who is who even without seeing them, and computer vision for active speaker detection, watching whose lips move in sync with the audio. Both run on two streams: the main robotic camera and a static wide shot.

For each person on stage, NATE builds a memory of their face, their camera position, and their voice fingerprint. Cross-checked against each other, those sources keep the system right even when one model is wrong.

Assumption 2: One Camera Is All the Coverage You Need

One camera is enough to follow one speaker, and it's cheaper. Then an MC walks on. A second speaker joins. Or a motivational speaker leaves the stage entirely and keeps talking from the back of the room, which has happened to us. The camera loses track, and with it any sense of what's going on.

So we added a static second camera that sees the whole room. That creates a new problem: two devices with different resolutions and angles have to agree on where everything is. We align them with homography, finding matching keypoints in both frames and calculating how one view maps onto the other. It has to work with whatever setup the room has, because no two venues are the same.

And it's live. One shot, one take. A bad alignment could send the camera zooming into the drapes mid-session, so every alignment gets a sanity check. Does the transformation keep right angles right? Is the image stretched more in one direction than the other? Do the edges line up?

Even a passing check isn't the final word, because someone can bump the tripod halfway through the day. NATE detects the speakers in both cameras and compares their positions. If they don't match, the alignment is off.

Assumption 3: Tracking a Speaker Means Following Their Face

Following the face sounds simple. It helps to look at how humans do this job first.

Human operators aren't perfect. A conference day is eight hours behind a camera with a few breaks, and nobody stays fully focused that long. Speakers drift out of frame, focus slips, and in a live show that split second ends up on the recording.

So how would a well-rested, fully focused videographer decide when to move? Not by following every movement. A camera locked onto the face of someone swaying at a lectern would make viewers seasick. A good operator waits to see whether the speaker settles into a spot, then holds the shot.

NATE mimics that with a heatmap of where the speaker's head has been recently. Positions accumulate and slowly fade, which separates someone shifting around one spot from someone genuinely moving to a new one. It also has to be fast. React more than two frames late and viewers notice.

Assumption 4: The Model Recognizes Every Face Equally Well

In a long panel, speakers switch, move, and swap seats. The system has to know who is who at all times, or the camera goes looking for the wrong person.

A good face recognition model should handle that. It didn't, at least not equally for everyone. Many face models are biased toward specific ethnic groups, and some speakers were recognized less reliably than others. Speakers also turn their backs to point at their slides, and then there is no face to recognize at all.

Months of looking for a model that would always work ended in a different answer: stop relying on face recognition alone. The wide camera keeps continuous track of everyone on stage, so faces need recognizing far less often, and when they do, recognition is cross-checked with re-identification models that look at the whole person. The goal is recognition that holds up in any circumstance, from low light to different hairstyles to a speaker facing the screen.

Assumption 5: Getting the Shot Is the Whole Job

Solve everything above and the job is done. Roll out thousands of cameras, mission accomplished.

Real events had other plans. In one session the power cut out mid-talk. Batteries kept the cameras recording, but how do you track a speaker in a pitch-black room?

That's one surprise out of many: a wrong input, full storage, missing slides, background noise, dropped frames, no signal, audio clipping, an underexposed image. Any of them can quietly ruin a recording that looks fine at a glance. So NATE doesn't stop at the shot. It keeps checking video, audio, slides, and storage throughout the session and alerts the technician when something needs attention.

It's also why humans stay in the loop. Electronic boxes fail, cables snap, and even with backups everywhere something needs troubleshooting on site. What should go away is the mundane part, eight hours standing behind a camera, so people can focus on the moments when the system is lost and a human has to step in.

Introducing Autopilot

Put the five lessons together and you get Autopilot, NATE's fully autonomous mode. Both cameras know who is on stage, active speaker detection runs on both feeds, and together they tell NATE who is speaking right now. The main camera moves to that person while the technician supervises instead of steering a joystick.

Test It Where It Breaks

The advice for anyone building AI systems: test in real-world scenarios, and go looking for the worst cases on purpose.

Back to the opening test. Ears alone couldn't place the voice on stage. The audience needed their eyes too, and reliable systems work the same way. Connect the dots, combine several approaches, validate them against each other.

And for anyone hosting events or meetups: record your sessions. For thousands of years, once a talk was over, the knowledge was lost. Recording is how it stays.

From the Q&A

❔How do you tackle so many problems at once?

One step at a time. Conferences like ICLR, ICML, and Ai4 keep us close to the computer vision community. We also tested constantly, with robotic cameras running alongside human operators at real events, plus simulations.

❔Why capture the screen instead of using the uploaded slides?

The deck a speaker sends is often not the one they present, or they switch to a demo. So we capture the screen feed directly, run OCR on it, and combine it with the transcript for better AI summaries and captions.

❔What does it need from the venue?

Standard AV: an XLR audio feed, an HDMI or SDI feed for slides, power, and an ethernet drop. Instead of a videographer with a tripod, a box arrives, and setup takes about 30 minutes. No six-foot table of gear that takes half a day to set up and another half to strike.


📢 Running a multi-track event? Book a call with us and we'll show you how RoboCam would cover your rooms.

Made for Real Conference Rooms

Power cuts, panel swaps, wandering speakers. Leave your email and see how NATE handles them.

Photo of Gareth Schumacher

Gareth Schumacher

Business Development Manager · SlidesLive

Made for Real Conference Rooms

Power cuts, panel swaps, wandering speakers. Leave your email and see how NATE handles them.

Photo of Gareth Schumacher

Gareth Schumacher

Business Development Manager · SlidesLive

Made for Real Conference Rooms

Power cuts, panel swaps, wandering speakers. Leave your email and see how NATE handles them.