WebRTC vs HLS for live classes: latency, scale and cost

Sub-second WebRTC or CDN-scale HLS? Latency budgets for live classes, SFUs, Low-Latency HLS, RTMP vs SRT ingest, the cost of a 5,000-student class and a hybrid design.

8 min read
On this page 8 sections
  1. Latency budgets for live classes
  2. WebRTC and SFUs
  3. HLS and LL-HLS
  4. Ingest: RTMP vs SRT, and WHIP
  5. Cost at scale: a 5,000-student class
  6. A hybrid architecture for large live classes
  7. Key takeaways
  8. Frequently asked questions

WebRTC delivers live video with well under a second of delay, which makes real conversation possible, but every viewer holds a live connection to a media server you run and pay for. HLS adds delay, from a few seconds with Low-Latency HLS to 18 seconds or more with standard settings, but it rides on ordinary CDNs, so growing from 500 viewers to 5,000 adds CDN traffic but almost no load on your own servers. Most large live classes need both: WebRTC for the teacher and the few students on stage, HLS for everyone watching.

Latency budgets for live classes

Latency here means glass to glass: from the teacher's camera to a student's screen. It builds up at every step: capture and encoding, the upload to your servers, transcoding, packaging, the CDN and finally the player's buffer. Different moments in a class tolerate very different amounts:

Class momentWhat students doDelay that worksUsual technology
Doubt-solving with a student on stage, one-to-one mentoringTalk back and forthWell under a secondWebRTC
Live quiz or poll during a lectureAnswer against a timerA few seconds, if the timer is synchronised to the videoLow-Latency HLS, or WebRTC for small groups
Questions in chatType; the teacher reads them aloudA few seconds is workableLow-Latency HLS
Lecture broadcast to thousandsWatch and listen15 to 30 seconds is usually acceptableStandard HLS

Delay also causes a subtler problem: a poll that opens on the teacher's screen can close before students 20 seconds behind have even seen the question. Live HLS playlists carry EXT-X-PROGRAM-DATE-TIME timestamps, which Apple requires for live streams, so the app can open and close each poll at the moment the student's video reaches it.

WebRTC and SFUs

WebRTC is the W3C and IETF standard behind video calls in the browser. It was designed for direct connections between participants: ICE finds a network path, STUN helps a device discover its public address, and a TURN server relays the media when no direct path works. Encryption is mandatory: under RFC 8827, media must be secured with SRTP, keys are negotiated with DTLS, and unencrypted media is not allowed.

Direct connections stop scaling after a handful of people, because every participant would have to upload a copy of their video to everyone else. Group sessions therefore go through a Selective Forwarding Unit (SFU), a server that receives each participant's stream and forwards selected streams to the others without decoding them. With simulcast, the teacher's browser sends two or three qualities at once, and the SFU forwards whichever one each viewer's connection can take.

The cost of WebRTC sits in that forwarding. Every viewer is a stateful, encrypted connection that some SFU must hold, pace and send packets to for the whole class. Very large audiences need a fleet of SFUs, often cascaded, plus TURN capacity for students on networks that block UDP. That works well for a class of 30 in which anyone can speak, and gets expensive for a broadcast to thousands.

HLS and LL-HLS

Standard HLS packages the live stream into segments of about six seconds, and RFC 8216 tells players not to start within three target durations of the live edge. With six-second segments, that alone puts viewers at least 18 seconds behind, before encoding and CDN delays. In exchange you get everything HTTP delivery offers: CDN caching, adaptive bitrate, easy recording and replay, and playback on every device.

Low-Latency HLS keeps those benefits and cuts the delay. Instead of waiting for a full segment, the server publishes partial segments; Apple's authoring specification recommends a one-second part target. Players ask for the next playlist update before it exists and the server holds the request until it does, so there is no polling. Because players must still stay at least three part durations behind the live edge, a realistic target is a few seconds. Your CDN has to support these blocking requests and short cache lifetimes for playlists, so test it before relying on it. Low-latency DASH works along similar lines, using chunked CMAF segments; our comparison of HLS vs DASH covers the formats.

Ingest: RTMP vs SRT, and WHIP

Ingest is how the teacher's stream reaches your servers. Three options matter today:

ProtocolTransportStrengthsWatch out for
RTMPTCPSupported by almost every encoder and streaming toolA lossy or congested uplink causes stalls rather than recoverable glitches; the original specification predates HEVC and AV1, which the Enhanced RTMP extension adds
SRTUDP, with retransmission of lost packetsDesigned for unreliable networks; optional AES encryption with 128, 192 or 256-bit keys; a configurable latency buffer, 120 ms by default in live modeNeeds UDP to be open between the studio and the ingest server
WHIPWebRTC over HTTP signallingStandardised as RFC 9725 in March 2025; lets a browser or encoder publish straight into a WebRTC or streaming serviceNewer, so check support in your encoder and media server

For a teacher streaming from a studio over imperfect broadband, SRT usually copes with packet loss better than RTMP. Whatever you choose, send a backup stream to a second ingest point before any class that matters.

Cost at scale: a 5,000-student class

Consider a 90-minute live class with 5,000 students watching at an average of 1 Mbps. These figures are illustrative.

  • Bytes delivered: each student receives about 675 MB, so the class delivers roughly 3.4 TB. That number is the same whichever technology you use.

  • With HLS: the CDN serves those bytes from cache. Your servers encode one ladder and publish each segment once; the CDN's mid-tier fetches each segment a handful of times, not 5,000 times. You pay the CDN's per-GB rate, and scale is mostly the CDN's problem.

  • With WebRTC: your SFUs send 5,000 individually encrypted streams, about 5 Gbps of sustained egress, and must hold 5,000 connections for 90 minutes. You need enough SFU capacity for the peak, with headroom for a node failing mid-class, and you pay your cloud provider's egress rates for every byte.

The bytes are identical; what changes is who carries them and how much state your own servers must hold. That's why pure WebRTC suits interactive sessions of tens or low hundreds, and HTTP streaming suits audiences in the thousands.

A hybrid architecture for large live classes

  1. The teacher, and any students invited on stage, join a WebRTC room on an SFU, where they can talk with sub-second delay.

  2. A recorder or compositor subscribes to the room and produces one programme feed: the teacher's camera, screen share and whoever is on stage.

  3. That feed goes to a live transcoder, which produces an adaptive ladder, and a packager, which writes HLS or Low-Latency HLS for the CDN.

  4. Thousands of students watch through the CDN in the app. Chat, polls and raised hands travel over a separate real-time channel; our guide to scaling WebSockets covers that side.

  5. When the teacher brings a student on stage, that student's app switches from HLS playback to a WebRTC connection. Everyone else hears the exchange a few seconds later on the stream.

  6. The recording becomes an on-demand lecture as soon as the class ends.

For very large events, put the HLS side behind more than one CDN; see our guide to multi-CDN failover. For how live classes fit alongside recorded video and tests, see our guide to scaling a video learning platform.

Key takeaways

  • WebRTC gives sub-second delay for conversation, at the cost of one server-side connection per viewer.

  • Standard HLS sits at least 18 seconds behind with six-second segments; Low-Latency HLS brings that down to a few seconds while keeping CDN delivery.

  • Use SRT for studio ingest over imperfect networks, RTMP where tools demand it, and WHIP for browser publishing.

  • For big classes, combine them: WebRTC on stage, HLS for the audience, and a separate channel for chat and polls.

Upclass gives institutes live classes in their own branded app, alongside their recorded courses and tests. See what's included on our LMS for coaching institutes page.

Frequently asked questions

What is low latency HLS?

Low-Latency HLS is an extension of Apple's HTTP Live Streaming that cuts live delay to a few seconds. The server publishes small partial segments as they are encoded, players ask for playlist updates before they exist and the server answers as soon as they do, and preload hints let players request upcoming parts early. It keeps CDN delivery and adaptive bitrate, so it scales like standard HLS.

Is WebRTC peer to peer?

It can be. WebRTC was designed for direct connections between browsers, using ICE and STUN to find a network path and TURN to relay traffic when no direct path works. For anything beyond a few participants, though, media usually goes through a server such as a Selective Forwarding Unit, because sending your video separately to every other participant quickly exhausts your upload bandwidth.

Is WebRTC encrypted?

Yes, always. The WebRTC security architecture, RFC 8827, requires media to be protected with SRTP using keys negotiated over DTLS, and forbids unencrypted media; data channels must use DTLS too. Encryption is between each endpoint and the next hop, so an SFU in the middle can technically access the media unless the application adds end-to-end encryption on top.

Which is better for live classes, WebRTC or HLS?

It depends on how many people need to talk. For small interactive classes, doubt sessions and mentoring, WebRTC's sub-second delay matters most. For lectures watched by thousands, HLS or Low-Latency HLS is cheaper and sturdier because CDNs carry the load. Large platforms usually combine them: WebRTC for the teacher and students on stage, HLS for the audience.

Share this article

Looking for something else?

Talk to Us