Heading to IBC in September? Download Streaming Media's curated guide to the MUST-SEES.
We'll see you there!

SRT over the internet vs. SRT in the cloud: where the operational differences begin

Article Featured Image

For many broadcasters, Secure Reliable Transport, or SRT, has become synonymous with IP contribution. Together with RIST, it has enabled broadcasters to replace or complement satellite capacity and dedicated fibre with secure, low-latency transport across standard IP networks.

Both protocols address the same fundamental challenge such as transporting high quality live video reliably across unpredictable, best-effort networks. They compensate for packet loss, jitter, and varying latency, making the public Internet a viable contribution path for many professional applications.

Because these protocols work so well, it is easy to assume that selecting SRT or RIST largely solves the problem. In reality, selecting the transport protocol is only the first step.

The difference between sending an SRT stream directly across the public Internet and operating the same stream through a managed cloud environment is not necessarily the transport protocol. In both cases, the video may still be carried using SRT.

The difference lies in everything surrounding it such as visibility, fault isolation, routing, scalability, media processing, interoperability, resilience, and centralised orchestration.

This becomes increasingly important as broadcasters move from occasional point to point links to operating hundreds of simultaneous feeds across multiple venues, production partners, rights holders and regions.

Transport reliability is not operational reliability

Public IP networks were not designed specifically for real time broadcast video. Packet loss, congestion, delay variation and changing latency are normal characteristics of Internet transport.

SRT operates above UDP and uses selective retransmission through Automatic Repeat reQuest (ARQ), buffering, round trip time measurement, congestion management. and AES encryption to maintain service across these conditions. RIST addresses the same overall problem through a different protocol architecture.

These mechanisms improve the probability that packets will arrive correctly and within the configured latency window. They cannot guarantee that every packet will always be recovered. If a retransmitted packet cannot arrive before the receiver’s latency buffer expires, it may still be discarded as too late.

More importantly, neither protocol is intended to operate an entire contribution network.

An SRT connection may be healthy while a processing node is overloaded. The transport may report no unrecovered packet loss while the MPEG transport stream contains timing or structural characteristics that affect the receiving decoder. The source may be working correctly while one destination is misconfigured.

SRT and RIST provides the transport foundation. It does not provide the complete operational environment.

“SRT can recover packets within a defined latency window. It does not orchestrate the wider workflow, manage multiple destinations or determine whether the media being transported is compatible with every receiver.”

What can the operator actually see?

Consider a live football feed being delivered from Madrid to London using SRT. The transmission begins to experience intermittent packet loss.

The first question is obvious: Where is the problem?

Internet-only connections

In a direct connection across the public Internet, the encoder may report retransmissions while the decoder reports packet loss or late packets. Between them sit multiple Internet Service Providers, transit carriers, autonomous systems and peering networks outside the broadcaster’s control.

SRT implementations expose valuable connection statistics, including round trip time, packet loss, retransmissions, estimated bandwidth, late packets, and buffer behaviour. These indicate the health of the overall SRT session, but they do not reveal every network segment between source and destination.

direct SRT connection

SRT can show that retransmissions are rising. It cannot identify the particular carrier, router or peering point introducing the impairment.

Cloud-managed contribution

A cloud managed contribution platform approaches the workflow differently. Instead of treating the entire path as one opaque connection, it divides the service into observable stages.

managed contribution workflow (observability)

This does not necessarily identify the individual router dropping packets. It does allow the operator to determine whether the problem is occurring between the venue and cloud ingress, within the processing environment, between cloud egress and the receiver, or at the destination itself.

Rather than asking, “Is the stream healthy?”, the operator can ask, “Which stage is unhealthy?” That difference can significantly reduce fault isolation and restoration time.

Creating a meaningful SLA boundary

Another misconception is that using SRT automatically creates an end-to-end Service Level Agreement (SLA).

An unmanaged Internet workflow traverses infrastructure owned by multiple organisations. Routing decisions, congestion, maintenance and failures across those networks remain outside any single provider’s control. No organisation can realistically guarantee the performance of an entirely unmanaged public Internet path.

A managed platform creates a clearer operational boundary. The public Internet may still be used across the first and last miles. Once the stream enters the managed environment, however, the cloud infrastructure, processing nodes, routing functions, monitoring services and distribution resources can be operated as one service.

creating a meaningful SLA boundary

This allows a provider to offer meaningful commitments covering the infrastructure and services it controls. First and last mile connections are as short as possible to ingress and egress the cloud.

The SLA is not based on controlling the whole Internet. It is based on operating, monitoring and providing resilience across the managed parts of the workflow.

Where provider backbones, private interconnections or managed network services are available, dependence on unmanaged routing can also be reduced. This does not make the entire path deterministic, but it creates a clearer technical and contractual boundary.

Routing, resilience, and service control

Direct Internet traffic follows routes selected by the Border Gateway Protocol (BGP). BGP decisions are influenced by reachability, network policies, commercial relationships and available paths. They are not based on broadcast specific requirements such as video latency, jitter or packet loss performance.

Routes may also change because of congestion, maintenance, provider failures or alterations to peering arrangements. Two transmissions between the same locations can therefore experience different performance at different times. Managed platforms cannot control public BGP routing either. What they can do is reduce the proportion of the end-to-end media path that depends on it.

A distributed gateway architecture, with media processing and routing nodes deployed across the Virtual Private Cloud (VPC), enables contribution feeds to enter the cloud through a gateway located as close as practical to the source. At the receiving end, broadcasters connect through a gateway positioned close to the destination. This minimises the unmanaged first and last mile of the journey, while keeping the core contribution and distribution path within the managed cloud environment, where performance, resilience and operational control can be maintained.

This distributed architecture changes the role of the public Internet. Rather than carrying the complete feed across one long, uncontrolled path between the venue and receiver, it is used primarily to connect each endpoint to the nearest suitable node within the VPC. The longer middle section of the workflow is then carried through the managed cloud core.

Once the stream enters the VPC, the platform has greater control over its routing and operational behaviour. Processing can be placed on the most appropriate available node. Distribution can be rerouted through another part of the VPC. Services can be migrated away from unhealthy infrastructure, and receivers can be redirected towards an alternative egress node without requiring a new contribution session from the venue.

Redundant ingress, processing or egress paths can also be activated when a node, region or route becomes unavailable. Because the source and receiver connect to a distributed service rather than depending on a single fixed end-to-end Internet path, the platform can adapt the workflow while keeping the first and last miles short.

Ingest once, distribute many times

Scalability is another area where architecture matters more than protocol selection.

Consider a venue delivering one 20 Mbps feed to ten international rights holders. In a direct point to point model, the encoder establishes ten independent SRT sessions. The venue’s transmission requirement rises from approximately 20 Mbps to 200 Mbps before retransmissions and network overhead are considered. Each connection also requires separate configuration, connection state, encryption, retransmission management, monitoring and fault handling.

With a managed platform, the venue contributes the feed once and replication takes place inside the distribution environment.

scaling architecture

The overall distribution requirement still grows as destinations are added, but that scaling moves away from the venue encoder and into infrastructure designed to support it.

The engineering team manages one contribution service rather than many separate source connections. A new rights holder can be added without creating another transmission from the venue.

For major sporting and news events, where many broadcasters may require the same feed, this has significant operational and bandwidth advantages.

The platform becomes a processing environment

Distribution is only part of the opportunity.

SRT transports the media payload reliably, but it does not correct or adapt the internal structure of the video, audio or MPEG transport stream it carries. A connection may report no unrecovered packet loss while a receiving decoder still experiences compatibility problems.

Two devices may both support MPEG-TS, H.264, H.265 and SRT but behave differently in production because of variations in:

  • PCR behaviour and timing
  • PAT and PMT repetition intervals
  • PID and service allocation
  • Constant or variable bitrate operation
  • Codec profiles, levels and resolutions
  • Audio formats and channel layouts
  • Transport stream multiplexing
  • Vendor specific decoder tolerances

Transport is therefore not always the problem, codec compatibility is. A software defined processing platform can adapt each output independently instead of forcing every receiver to accept the original contribution unchanged.

One broadcaster may require 1080i while another requires 1080p. One destination may need H.264 while another can accept HEVC. An IRD may expect specific PID structures or constant bitrate (CBR), while another receiver may need variable bitrate (VBR) transport stream.

platform becomes a processing environment

A single contribution can be transcoded, remultiplexed, filtered, remapped, or converted into another delivery protocol. Audio layouts can be adapted, selected services can be extracted, and outputs can be monitored for bitrate, transport-stream and ETR 101 290 conditions.

These functions were traditionally performed by dedicated hardware appliances distributed throughout the network. In a software defined environment, they can be instantiated only where required and applied only to the destinations that need them. The result extends beyond transport and protocol conversion to deliver true interoperability.

From alarms to operational intelligence

Traditional Internet contribution is often reactive. Packet loss occurs, an alarm is generated, an engineer investigates and corrective action begins. By then, the service may already have been affected. Managed platforms collect telemetry across the complete workflow, including transport performance, bitrate, latency, RTT, retransmissions, connection status, processing resources, transport stream behaviour and receiver availability.

These measurements can be correlated across ingress, processing, routing, egress and delivery rather than viewed as separate device statistics. This matters because not every fault begins as a complete outage. A rising retransmission rate may indicate a deteriorating path. Increasing RTT may make the configured SRT latency insufficient. A processing node may be approaching its resource limit while the output remains on air. A connected receiver may begin reporting increasing late packets.

Continuous observability allows operators to identify these conditions earlier and potentially reroute a service, move processing, activate resilience or investigate a receiver before the problem becomes visible on air.

The value lies not in generating more alarms, but in providing the operational context needed to respond effectively. Operators can immediately see which service is affected, where the issue resides within the workflow, which destinations are at risk and what corrective action can be taken.

Centralised orchestration

As contribution networks grow, so does the number of encoders, decoders, connections, processing nodes, receivers, recordings, monitoring points, destination profiles, schedules and redundancy paths. Managing each element separately may be practical for a small number of transmissions but becomes increasingly difficult at scale.

A software defined platform can define the complete service centrally. The operator selects the source and destinations, applies the required processing, defines the schedule and chooses the resilience level. The platform then creates transport paths, processing stages and receiver outputs as one workflow.

This reduces manual configuration and the opportunity for human error, particularly when changes must be made at short notice. A new rights holder can be added during an event. A destination can receive a different bitrate or multiplex structure. A receiver can move to another region, or a backup path can become primary.

Where edge devices are integrated into the same orchestration layer, codec settings, transport parameters, scheduling and monitoring can also become part of the complete service definition.

Cloud, on-premises, or hybrid?

Cloud has transformed broadcast contribution, but not every workflow belongs permanently in a public cloud.

Temporary events may require substantial processing and distribution capacity for only a few hours or days. Cloud infrastructure is well suited to these occasional use requirements because resources can be created quickly, scaled and released when no longer needed.

Permanent 24-hour services are different. A broadcaster may already own resilient data centres, private fibre and virtualisation infrastructure. Moving every permanent workload to public cloud can create continuous compute, storage and network charges for resources used throughout the year. The practical answer is often a hybrid architecture.

hybrid architecture

Always on services can remain on customer owned infrastructure, while cloud resources are introduced for major events, breaking news, temporary rights holders, disaster recovery, international delivery or additional processing. Rather than treating cloud and on premises as mutually exclusive strategies, broadcasters should consider which parts of the workflow are best suited to each environment based on their operational, technical and commercial requirements.

The greatest benefit is realised when the same orchestration, scheduling, monitoring and processing model operates consistently across cloud, on premises and edge infrastructure. The location of the processing then becomes a deployment decision rather than an operational one, allowing workloads to move between environments while maintaining a consistent management and operational experience.

It’s not just about transport

None of this diminishes the importance of SRT or RIST. Without reliable transport protocols, modern IP contribution would not exist, and they remain the foundation for moving live media securely and efficiently across unpredictable IP networks.

However, operating a modern broadcast network requires far more than simply transporting packets from one location to another. Broadcasters need end-to-end visibility, rapid fault isolation, scalable distribution, media processing, interoperability between diverse broadcast systems, centralised orchestration and resilience across multiple networks and locations.

These capabilities are not inherent features of SRT or RIST themselves. Rather, they are delivered by the operational platforms built around them, transforming reliable transport into a complete software defined media workflow.

Beyond choosing a transport protocol

None of this diminishes the importance of SRT or RIST. Without reliable transport protocols, modern Internet contribution and distribution would not exist. They provide the foundation for moving live media securely and efficiently across unpredictable IP networks.

However, as contribution networks continue to grow in size and complexity, selecting a transport protocol is only one part of the overall solution. Broadcasters increasingly require the ability to monitor services end-to-end, scale distribution efficiently, adapt content for different receivers, automate operational workflows and deploy processing wherever it makes the most technical and commercial sense.

This is where the distinction between transport and platform becomes increasingly important. The transport protocol determines how media moves across the network, while the platform determines how that media is processed, monitored, orchestrated and managed throughout its entire lifecycle. Together, these capabilities enable broadcasters to build resilient, scalable contribution and distribution networks that extend across cloud, on-premises and edge environments while maintaining a consistent operational model.

As software defined media architecture continues to evolve, the focus will increasingly shift from the transport protocol itself to the intelligence of the platform that surrounds it. Reliable transport will remain fundamental, but it is the operational platform that ultimately determines how efficiently, flexibly and reliably an IP media network can be deployed and managed.

[Editor's note: This is a contributed article from GlobalM. Streaming Media Global accepts vendor bylines based solely on their value to our readers.]

Streaming Covers
Free
for qualified subscribers
Subscribe Now Current Issue Past Issues
Related Articles

Video encoding and decoding's interoperability problem—and what we can do about it

Video encoder and decoder interoperability remains one of the most frustrating and time-consuming challenges in live video workflows, especially in IP-based, protocol-native distribution. Even when everything looks good on paper, the result on screen can be anything from smooth playback to stuttering video and missing audio. So why does this happen? And more importantly, what can we do about it?