Video transcoding pipelines: from upload to HLS
How to turn an uploaded lecture into encrypted HLS: probing, encoding a ladder with FFmpeg, queue-driven workers, CPU vs GPU vs spot capacity, retries and what one lecture really costs.
On this page 10 sections
Video transcoding converts an uploaded video into the versions students actually stream: several resolutions and bitrates, cut into short segments and listed in HLS playlists. A transcoding pipeline is the chain of steps that does this for every upload without anyone watching over it: check the file, encode a bitrate ladder, package it, encrypt it and publish it. For a coaching platform that uploads long lectures every day, the dependable way to run it is as queued jobs on workers that can fail and retry safely.
What transcoding does
A lecture arrives in whatever format the recording device chose: HEVC from an iPhone, a variable frame rate from screen-capture software, or a 4K camera file of several gigabytes. Transcoding decodes it to raw frames, optionally scales it or changes the frame rate, and encodes it again with settings every student's phone and browser can play.
A few terms get mixed up, so it helps to separate them:
| Term | What it means | Example |
|---|---|---|
| Codec | The compression format of the audio or video itself | H.264, HEVC and AV1 for video; AAC for audio |
| Container | The file format that wraps encoded streams with timing information | MP4, fragmented MP4, MPEG-TS |
| Encoding | Compressing raw video with a codec | A camera or OBS writing H.264 |
| Transcoding | Decoding and re-encoding, usually to other sizes or bitrates | One 1080p upload turned into five renditions |
| Transmuxing | Moving the same encoded streams into another container, without re-encoding | MPEG-TS segments repackaged as fragmented MP4 |
Transcoding is expensive because every frame is decoded and encoded again. Transmuxing is cheap. A good pipeline keeps the two apart, so you can repackage or re-encrypt a lecture without paying for the encode twice. For how this fits with CDNs, live classes and tests, see our guide to scaling a video learning platform.
The stages of a transcoding pipeline
- Upload straight to object storage. Use multipart or resumable uploads from the teacher's app or browser, so a 3 GB file never passes through your web servers and a dropped connection doesn't start from zero.
- Probe and validate. Run
ffprobeto read duration, codecs, resolution, frame rate, rotation and audio channels. Reject files with no video, zero duration or no audio before you spend any compute on them. - Create a job. Write a job record (upload ID, ladder version, status) and put a message on the transcoding queue.
- Encode the ladder. Produce each rendition with keyframes at fixed intervals, so every rendition can be cut at exactly the same points.
- Package. Split the renditions into segments and write the media and multivariant playlists. Apple's HLS authoring specification recommends six-second segments with a keyframe every two seconds.
- Encrypt. Encrypt the segments and register the keys with your key service. Our guide to HLS encryption explains AES-128 and SAMPLE-AES.
- Check the output. Confirm that every rendition matches the source duration to within a frame or two, that the playlists parse, and that a sample of segments plays.
- Publish. Mark the lecture ready, notify students if the batch schedule says so, and move the original to a cheaper storage class. Keep it: you will want to re-encode when your ladder or codec choices change.
FFmpeg to HLS: a three-rung example
FFmpeg can encode and package in one command. This produces 360p, 480p and 720p renditions of a lecture as HLS, with one playlist per rendition and a multivariant playlist that lists them all:
ffmpeg -i lecture.mp4 \
-filter_complex "[0:v]split=3[a][b][c];[a]scale=-2:360[v0];[b]scale=-2:480[v1];[c]scale=-2:720[v2]" \
-map "[v0]" -map "[v1]" -map "[v2]" -map 0:a:0 -map 0:a:0 -map 0:a:0 \
-c:v libx264 -profile:v high -preset medium \
-force_key_frames "expr:gte(t,n_forced*2)" -sc_threshold 0 \
-b:v:0 350k -maxrate:v:0 525k -bufsize:v:0 700k \
-b:v:1 600k -maxrate:v:1 900k -bufsize:v:1 1200k \
-b:v:2 1200k -maxrate:v:2 1800k -bufsize:v:2 2400k \
-c:a aac -b:a 96k -ac 2 \
-f hls -hls_time 6 -hls_playlist_type vod \
-hls_flags independent_segments -master_pl_name master.m3u8 \
-var_stream_map "v:0,a:0 v:1,a:1 v:2,a:2" \
-hls_segment_filename "seg_%v_%05d.ts" "index_%v.m3u8"
The flags worth understanding:
-force_key_frames "expr:gte(t,n_forced*2)"places a keyframe every two seconds whatever the source frame rate, and-sc_threshold 0stops x264 adding extra keyframes at scene changes. Together they keep segment boundaries identical across renditions, which players need in order to switch cleanly.-hls_time 6sets the target segment length. The default in FFmpeg's HLS muxer is 2 seconds, and it cuts each segment at the next keyframe after the target, which is why keyframe placement matters so much.-maxrateand-bufsizecap each rendition's peaks. Apple's specification says the peak bitrate of on-demand content should be no more than 200% of the average.-var_stream_mappairs each video rendition with an audio stream and writes one playlist per pair;-hls_playlist_type vodmarks the playlists as complete.
Add -hls_segment_type fmp4 to write fragmented MP4 segments instead of MPEG-TS; Apple's specification requires fragmented MP4 for HEVC and AV1, and DASH players can share the same files. Choosing the rungs themselves, how many and at what bitrates, is its own design problem, covered in our guide to adaptive bitrate streaming.
Queues and workers
Transcoding is slow, bursty and prone to retries, which makes it a textbook background job. The upload request should finish as soon as the file is safely in storage; everything after that runs on workers pulling from a queue. A few rules keep this healthy:
- One job per lecture per ladder version. Use the upload ID plus the ladder version as an idempotency key, so a duplicate message can't produce two sets of files or two "new lecture" notifications.
- Write to a new path and publish at the end. Encode into a job-specific folder and flip the lecture to "ready" only after every rendition, playlist and check has passed. A half-finished job then never becomes visible to students.
- Let each worker take one job at a time. A worker that reserves several two-hour encodes sits on work that idle workers could have started. In Celery, that means a prefetch multiplier of 1 and late acknowledgement on this queue.
- Give transcoding its own queue. A backlog of re-encodes must never delay OTP or password-reset messages, and tonight's lecture should jump ahead of a back-catalogue re-encode.
- Scale workers on queue depth, not CPU. The numbers that matter are the hours of video waiting and the age of the oldest job.
Our guide to message queues and background jobs covers retries, idempotency and dead-letter queues in more depth. For very long lectures, you can also split the source into chunks at keyframes, encode the chunks on several workers in parallel and join the results. It needs identical encoder settings for every chunk and care at the joins, but a two-hour encode becomes many short jobs that can each be retried on its own.
CPU, GPU or a managed service
| Option | Strengths | Trade-offs | Fits when |
|---|---|---|---|
| Software encoders on CPUs (x264, x265, SVT-AV1) | The most control over quality per bit; slower presets fit more quality into the same bitrate; runs anywhere | Slowest per video, so a large backlog needs many cores | Recorded lectures, where quality per bit, and therefore delivery cost, matters most |
| Hardware encoders on GPUs (such as NVIDIA NVENC or Intel Quick Sync) | Much faster per machine, with many streams at once; suits live transcoding | Quality at a given bitrate can differ from slow software presets; fewer tuning options | Live classes, big one-off backlogs, tight turnaround |
| Managed transcoding services from cloud providers | No encoders to operate; typically priced per minute of output | Less control; per-minute pricing adds up for a large library | Small teams, spiky volumes, a first version of the pipeline |
Whichever you pick, compare options on your own lectures rather than on film trailers. Encode a blackboard lecture, a slide lecture and a phone recording at the same bitrates, score them with an objective metric such as Netflix's open-source VMAF, and have a person check that handwriting is still readable at 360p.
Spot and preemptible capacity
Batch transcoding suits spare cloud capacity sold at a discount; AWS, for example, advertises Spot Instances at up to 90% off on-demand prices. The catch is that the provider can take the machine back, and on AWS you get a two-minute interruption notice. That is harmless when jobs are short or chunked and idempotent: the worker stops taking new work when the notice arrives, the unfinished job returns to the queue, and another worker picks it up. Keep a small on-demand pool for urgent work, such as a lecture that must be live by 7 p.m.
A worked example: what one lecture costs
Take a two-hour lecture and a five-rung ladder of 110, 350, 600, 1,200 and 2,000 kbps, plus a 96 kbps audio track stored once. All figures here are illustrative; substitute your own.
- Storage: the renditions add up to about 4.4 Mbps, so two hours comes to roughly 4 GB, on top of the original upload.
- Compute: suppose your benchmark shows one 16-core worker encoding the whole ladder at twice real time. The lecture then needs one worker-hour, which is half a worker-hour per hour of lecture. At ₹60 an hour for that worker, that is about ₹30 per lecture-hour, and less on spot capacity.
- Delivery: if 500 students each watch the full lecture at an average of 800 kbps, each one pulls about 720 MB, or 360 GB in total.
The lesson is in the ratio. You encode and store about 4 GB once, but deliver around 90 times that, and the figure grows with every enrolment and every rewatch before an exam. A ladder that trims 15% of delivered bytes at the same perceived quality is worth far more than a faster encoder, so it pays to spend compute on slower presets and per-title tuning.
Failures and retries
Most failed jobs fall into a few groups, and each needs different handling:
| Failure | Example | What to do |
|---|---|---|
| Bad input | A corrupt file, no audio track, zero duration, an unsupported codec | Fail fast at the probe step, don't retry, and tell the uploader what to fix |
| Awkward input | Variable frame rate screen recordings, phone video with rotation metadata | Normalise early (constant frame rate, rotation applied) and keep sample files of each case in your test suite |
| Transient infrastructure | A storage timeout, a worker running out of memory, a spot interruption | Retry automatically with exponential backoff, up to a limit |
| Stuck job | The encoder hangs on a damaged frame | Set a timeout proportional to the video's duration, retry once, then move the job to a dead-letter queue for a person to inspect |
Alert on the age of the oldest queued job and on failures per hour, not just on worker CPU. And keep every original: when you add HEVC or AV1 renditions, or retune the ladder, re-encoding the library is just another batch of jobs.
Key takeaways
- Transcoding turns one upload into a ladder of renditions; packaging and encryption turn those into HLS segments and playlists.
- Upload straight to storage, validate early, and run everything else as queued, idempotent jobs.
- Force keyframes at fixed intervals so every rendition is segmented at the same points.
- Choose CPU, GPU or a managed service by testing quality per bit on your own lectures, and use spot capacity for work that can retry.
- Delivery, not encoding, dominates video cost, so spend compute on a better ladder.
Institutes on Upclass get courses, live classes and tests in their own branded app, with VidSafe proprietary encryption, dynamic watermarking, and screen- and camera-recording detection. See what's included on our LMS for coaching institutes page.
Frequently asked questions
What is video transcoding?
Video transcoding is converting an already compressed video into another version: a different codec, resolution, bitrate or frame rate. The file is decoded to raw frames and encoded again with new settings. Streaming platforms transcode every upload into several renditions, for example from 240p to 1080p, so each viewer's player can pick the one their screen and connection can handle.
What is transcoding video files?
Transcoding video files means re-encoding them from one format to another, such as turning an iPhone's HEVC .mov file into an H.264 .mp4 that plays almost everywhere, or shrinking a 4K recording to 720p. Every transcode is lossy, so always work from the best original you have rather than transcoding a transcode, and keep that original in case you need to re-encode later.
What is video encoding?
Video encoding is compressing raw video frames with a codec such as H.264, HEVC or AV1 so the result is small enough to store and stream. Encoders discard detail viewers are unlikely to notice and reuse information between neighbouring frames. Transcoding includes an encoding step: it first decodes an existing compressed file, then encodes it again with the new settings.
Should you use a video transcoding API or run FFmpeg yourself?
A managed transcoding API is the quicker start: there are no servers to run, and you pay per minute of output. Running FFmpeg on your own workers takes more engineering, but gives you full control over the ladder, encoder presets and per-title tuning, and can cost less at steady, high volumes. Many teams start with a managed service and bring the high-volume path in-house later.