Summary: Most telehealth platforms are built around a single assumption: one patient, one provider, one telehealth video call. That holds for a routine check-in, but breaks down the moment care involves a caregiver, an interpreter, a specialist, a care team, or an emergency response — all of which need more than two people on the call. This guide covers five points in the care journey where a standard two-party video call stops being enough, and what that means for the platforms built to support it.
Table of Contents
Ask most people what a telehealth video call looks like, and they’ll describe two windows on a screen: a patient, and a doctor. That picture is accurate for a huge share of virtual care — a med refill check-in, a routine follow-up, a quick symptom review. It’s also incomplete in a way that matters more than most telehealth builds account for.
Healthcare is rarely a conversation between exactly two people. A patient’s care decisions often involve a spouse or adult child. A significant share of the population needs a professional interpreter to participate meaningfully in their own appointment. Primary care physicians routinely need a specialist’s input in the middle of a visit, not after a multi-week referral cycle. Complex cases — an unusual cancer diagnosis, a rare condition — get better outcomes when several clinicians look at them together, in real time. And in an emergency, care is inherently a team activity: a remote specialist, on-site staff, and sometimes a coordinating physician all need to be present in the same moment.
None of that fits into a video architecture designed for one-to-one calls. Bolting a third or fourth participant onto a system built for two is where a lot of telehealth platforms discover, usually mid-deployment, that “add a participant” is a much bigger technical and UX decision than it looks. This guide covers five points in a care journey where single-party video stops being sufficient, grounded in the research behind each — and what that means for the infrastructure decision underneath.
Key Takeaways:
For a large share of telehealth visits, “patient” and “the person who needs to be on the call” aren’t the same thing. An adult child helping an aging parent navigate a diagnosis. A spouse who remembers the medication list better than the patient does. A parent present for a pediatric visit, alongside a caregiver who handles day-to-day care. In all of these cases, a two-person video call forces an awkward workaround — a phone held up to a laptop screen, a caregiver crowded into frame, or the visit happening without the person who most needs to hear it.
The Department of Veterans Affairs has already built this as a first-class feature rather than a workaround. Its Caregiver Connect program lets a patient bring up to five family members or caregivers into a video telehealth appointment — either scheduled in advance or added mid-visit, when a provider realizes partway through that a caregiver’s input would help. That second capability matters more than it might sound: care conversations don’t always reveal their complexity before they start.
The data backs up why this matters. A national survey of VA occupational therapy practitioners found that increased access to video telehealth was the single most-cited benefit of caregiver participation, reported by over 90% of respondents, alongside improved collaboration with family reported by a similar share. A separate national study of informal caregivers went further, concluding that telehealth platforms designed specifically to support multi-participant visits — alongside better scheduling and coordinated after-visit follow-up — are best positioned to reduce caregiver burden and improve continuity of care.
For development teams, the implication isn’t just “allow a third video tile” to support a multi-party call. It’s that identity, consent, and access all need to account for a non-patient participant from the architecture stage — who can join, whether that requires the patient’s consent on record, and how that participant’s presence is reflected in the visit’s documentation. Retrofitting that after a platform is built around a strict one-to-one model is a meaningfully harder problem than designing for it from the start.
An estimated 29.6 million people in the U.S. have limited English proficiency, per Census Bureau American Community Survey data — and a separate analysis of national survey data puts the Deaf and Hard of Hearing population at roughly 11 million, split between about 10 million hard of hearing and close to 1 million who are functionally deaf. Together, that’s a substantial population for whom a standard two-person video visit isn’t fully usable without a third participant on the call.
The gap this creates is measurable. A study of over 955,000 primary care telemedicine visits at Kaiser Permanente Northern California found that patients with documented limited English proficiency used video visits at a lower rate than other patients — 34.5% versus 39.8% — even after adjusting for technology access and other factors. Notably, that gap closed almost entirely among patients who had used video before: with prior experience, LEP and non-LEP patients chose video at statistically indistinguishable rates. The likely story isn’t that video doesn’t work for LEP patients — it’s that the first visit, without an interpreter built into the experience, creates a barrier that follow-up visits don’t.
This is a compliance question as much as a UX one. Section 1557 of the Affordable Care Act requires covered healthcare entities to provide qualified medical interpreters for patients with limited English proficiency, and Title VI of the Civil Rights Act extends similar obligations to any organization receiving federal funding through Medicaid or Medicare. A telehealth video call that can’t seamlessly add a third participant for interpretation isn’t just creating friction — for many organizations, it’s creating a compliance gap.
The practical shape of this in production is video remote interpreting: a qualified interpreter joins the same call as the patient and provider, in real time, rather than the visit being conducted over a translated relay or postponed until an in-person interpreter is available. That requires the same participant-management capability as the caregiver case above, with one addition — the interpreter is typically a third-party professional joining from an entirely separate organization, which raises its own questions around identity verification and session access that a purely internal multi-participant model doesn’t have to solve.
A primary care visit often surfaces something that isn’t quite the primary care physician’s call to make alone — a rash that might be more than a rash, a cardiac symptom worth a second read, a case that would normally trigger a referral and a multi-week wait. The traditional path is to refer out, and the patient waits. The alternative — pulling a specialist into the conversation while the patient is still on the call — is a fundamentally different care model, and one the evidence increasingly supports.
Most of the research here comes from eConsults: asynchronous, secure messaging exchanges where a primary care physician sends a case to a specialist and gets guidance back, often without the patient needing a separate appointment. More recent evidence shows how often that can change the referral pathway. A study of eConsults for medically underserved primary care patients found that specialists considered a face-to-face referral unnecessary in 66% of cases, while primary care providers documented acting on at least one specialist recommendation within 90 days in 77% of cases. A study of Ontario’s eConsult service in federal correctional facilities found an even greater effect: 81% of 906 cases were resolved without an in-person specialist appointment, with specialists responding in a median of 0.9 days. The exact impact varies by setting and specialty, but the broader pattern is consistent: specialist input can often be delivered without putting the patient through a separate referral and appointment process.
That evidence base is for asynchronous, text-based consultation — not a specialist joining a live video call. But the underlying principle is the same one that makes eConsults work: specialist input doesn’t always require a full standalone appointment; sometimes what is needed is the specialist’s judgment applied to the case in front of the primary care physician. Live video is the natural real-time extension of that same idea — instead of sending the case and waiting a day for a written response, the specialist joins the call, sees what the primary care physician sees, and the patient gets an answer before they leave the visit rather than after a referral cycle.
This plays out in two distinct ways. The first is unplanned: a decision made by the provider in the moment, based on what the visit reveals. That case needs a platform that supports adding a participant to an already-live session with minimal friction — pulling a specialist from a directory, sending an instant invite, getting them connected mid-call.
The second is pre-planned: a rural primary care visit scheduled in advance to include a specialist from another department, a different hospital, or a different city entirely — the same access problem eConsults solve, but addressed synchronously rather than through a written exchange. This is one of the ways telehealth can extend specialist access in rural areas, where distance and workforce shortages can otherwise turn a referral into a significant barrier to care. A patient in a small town shouldn’t need to travel to see a specialist who’s willing to join the call from three states away; the visit just needs to be built to expect a third participant from the start.
Both cases require the same underlying capability — flexible, on-demand or scheduled participant addition — but they place different demands on the platform: one needs speed and low friction in the moment, the other needs the scheduling and directory infrastructure to coordinate across organizations in advance.
Some cases aren’t a two-person decision, or even a two-plus-specialist decision — they need several clinicians looking at the same patient at the same time. Oncology is the clearest example: a new cancer diagnosis routinely benefits from a tumor board where oncology, radiology, pathology, and surgery all weigh in together, but tumor boards have historically required every specialist to be in the same room, which limits how often smaller or rural facilities can offer them at all.
Virtual tumor boards remove that constraint, and more recent evidence shows how much participation can expand when clinicians no longer need to be in the same place. A study comparing 12 months of in-person tumor boards with 12 months of virtual meetings at an NCI-designated Comprehensive Cancer Center found that physician attendance increased by 46%, while the number of patient cases discussed rose by 20%. More recently, a study of the VA National TeleOncology service reported 113 virtual tumor board sessions involving 233 patient cases from 51 VA facilities across 33 states and Puerto Rico, demonstrating how the model can extend multidisciplinary oncology expertise across geographically dispersed healthcare facilities.
This is the most demanding of the five use cases from an infrastructure standpoint, because it isn’t three participants, it’s potentially six or eight — an oncologist, a radiologist, a pathologist, a surgeon, a nurse coordinator, and the patient, all on the same call, often with someone screen-sharing imaging or a diagnostic report at the same time. Session stability, participant management, and screen-sharing all get meaningfully harder as the participant count grows, and it’s precisely the case that most exposes a video architecture that was only ever tested with two or three people on a call.
Emergency care is multi-party by nature in a way none of the previous four cases quite are — it’s rarely a single provider making a single decision. A rural emergency department connecting with a remote neurologist for a suspected stroke, an ICU consulting a specialist while on-site nursing staff manage the patient directly, a transport coordinator joining a call to arrange a transfer while the receiving physician is still assessing the case — these scenarios routinely need more than two participants on the line simultaneously, and they need it fast, with no time to troubleshoot a platform that wasn’t built for it.
The clinical case for telehealth in emergency settings, including telestroke and Tele-ICU programs, is covered in depth in The Impact of Telehealth on Emergency Healthcare. The point worth making here is narrower: emergency escalation is the use case where multi-participant capability stops being a convenience feature and becomes a patient-safety requirement. A call that only supports two people is a call that cannot, by design, support the way emergency medicine actually gets practiced — and unlike the other four cases, there’s rarely time to work around that limitation in the moment.
Two-person video and multi-participant video are not the same engineering problem scaled up — they’re different problems. A platform built around a single patient and a single provider can make simplifying assumptions that break the moment a third participant enters: who has permission to join a session, whether a participant can be added after a call has already started, how many simultaneous video feeds the infrastructure can sustain without degrading quality, and what happens to screen-sharing, recording, and session logging when there are six participants instead of two instead of one shared view.
None of the five scenarios above are solved by the same configuration. A caregiver joining a routine visit and an eight-person tumor board with simultaneous screen-share are both “multi-participant video,” but they sit at opposite ends of what the underlying session needs to support — participant count, join/leave flexibility, and bandwidth demands all scale differently depending on which of these cases a platform is actually built for. And to be clear on a basic point that’s easy to assume rather than state outright: telehealth was never inherently limited to a single patient and a single provider — that’s simply the narrowest and most common deployment of it, not a ceiling built into the concept of video-based care.
For development teams, there are two broad ways to support these use cases. One is to build multi-participant telehealth video calls directly into an existing healthcare application using a video SDK, which provides the underlying real-time communication infrastructure while giving the development team control over the user experience and surrounding workflows. The alternative is a pre-built, customizable white-label consultation platform, where multi-party video and features such as virtual waiting rooms, chat, and patient workflows are already part of the product.
Whichever route a team takes, it is important to evaluate how the technology handles participant management at scale, not simply whether it technically supports “more than two people” on a call. For teams taking the build-it-yourself route, the video layer is only one part of the communication stack. Developers also need to consider how video, chat, authentication, notifications, and other communication features work together within the application. Our guide to Telehealth Chat APIs and Video SDKs looks more closely at those infrastructure choices.
Multi-participant video also needs to be considered as part of the application’s broader scaling strategy. As usage grows, development teams need to plan not only for larger calls, but for increasing numbers of users, sessions, messages, integrations, and concurrent workloads across the platform. See How to Build a Telemedicine App That Scales for a broader look at those architecture decisions.
The standard two-person telehealth video call is the easiest version of telehealth to build, which is exactly why so many platforms are architected around it and nothing else. But the moments that matter most in care — a family member helping a patient understand a diagnosis, an interpreter making a visit actually accessible, a specialist weighing in before a case gets worse, a care team reviewing a complex diagnosis together, an emergency unfolding in real time — are precisely the moments where that assumption runs out. The evidence across these use cases points in the same direction: multi-participant video can improve access, collaboration, and continuity of care — not simply make virtual visits more convenient.
For teams that want to build this themselves, QuickBlox’s Video SDK provides the infrastructure to add multi-party video directly into your own application. For teams that would rather not build the consultation layer from scratch, Q-Consultation is a ready-to-go, white-label video consultation product built specifically for healthcare — multi-party video, chat, virtual waiting rooms, and the surrounding workflow already in place. Get in touch to talk through which approach fits your platform’s care model.
If you’re evaluating how to support video and real-time communication within a healthcare application, these guides explore some of the technical, security, and infrastructure decisions in more detail: