How we reduced CPU and memory use—and what our measurements against Haivision SRT 1.5.7 show.
Many concurrent SRT connections place different demands on a system than a single stream. Alongside packet processing itself, memory allocation, frequent wake-ups, and the distribution of work across threads can consume a substantial share of resources. These are the areas we focused on in Robotweax SRT 0.2.6.
Our starting point was a series of scaling measurements in the SRT Network Lab. They led to targeted investigations and changes: less unnecessary work on the receive path, packet memory used as needed, and revised scheduling. The result is our published Performance Optimization release, version 0.2.6. The version number refers to the Robotweax implementation; the SRT protocol and compatible API remain at version 1.5.7. See the release and compatibility scope.
Less CPU Work on the Receive Path
One important finding concerned notifications triggered when an application reads received data. These reads free space in the receive window. But if every read immediately starts another channel-processing pass, high receive loads generate additional scheduling and synchronization work.
We now coalesce these notifications and use processing passes that are already pending. Receive-window progress and protocol deadlines are preserved. Idle channels can also wait for shared socket-readiness monitoring instead of being polled unnecessarily often.
A targeted Linux before-and-after measurement with 32 streams at 10 Mbit/s each showed the following effect of the receive-path optimization:
| Transmission direction | Receiver CPU before | Receiver CPU after | Change |
|---|---|---|---|
| Haivision → Robotweax | 6.364 CPU-seconds | 2.536 CPU-seconds | −60.2% |
| Robotweax → Robotweax | 6.436 CPU-seconds | 2.451 CPU-seconds | −61.9% |
The comparison used two successive Robotweax development revisions, with two runs per case in alternating order.
Use Memory as Needed While Preserving Buffer Capacity
The second major focus was packet storage. Previously, large payload areas became resident when a session was created, even if the connection used only a small part of its configured capacity.
The new design separates packet metadata from payload data. It still reserves the configured capacity, but touches payload memory pages only when needed. Freed payload slots can be reused independently of the packet ring’s position.
In a Linux component test, the additional resident memory of an empty session with default capacities fell from approximately 23.74 to 1.18 MiB—about 95%. The effect was also visible under real network traffic. In an isolated before-and-after measurement of the storage change with 16 streams, the results were:
| Robotweax → Robotweax, 16 streams | Process RSS before | Process RSS with payload pool | Change |
|---|---|---|---|
| Sender | 915.1 MiB | 53.5 MiB | −94.2% |
| Receiver | 917.8 MiB | 97.8 MiB | −89.3% |
Comparison with Haivision SRT 1.5.7
How does the optimized implementation compare with the reference implementation? The final evaluated Linux series supports a specific comparison at 32 concurrent streams and a total offered payload rate of 320 Mbit/s.
The following values were observed with Robotweax version 0.2.6. Each case used a sender and receiver running the same implementation on the same Linux ARM64 VM: four Robotweax runs and two Haivision control runs.
| Metric | Robotweax P28 | Haivision SRT 1.5.7 |
|---|---|---|
| Received rate | 319.696 Mbit/s | 319.994 Mbit/s |
| Sender CPU time | 3.607 s | 3.661 s |
| Receiver CPU time | 1.992 s | 3.980 s |
| Combined CPU time | 5.610 s | 7.641 s |
| Combined peak RAM use | 43.12 MiB | 24.34 MiB |
| Process threads, sender / receiver | 4 / 8 | 66 / 36 |
| Total process threads | 12 | 102 |
At nearly the same received rate, Robotweax used approximately 50% less receiver CPU time and 26.6% less CPU time overall in this series. The number of process threads was also substantially lower: 12 versus 102, or 88.2% fewer.
Haivision had the advantage in memory use. Despite the substantial improvement over earlier Robotweax revisions, Robotweax’s simultaneous RAM peak was 77.1% higher in this comparison.
Under What Conditions Were These Measurements Taken?
The final comparison series ran on a dedicated Ubuntu ARM64 VM with four vCPUs and approximately 6 GiB of RAM. Sender and receiver communicated over loopback. The tests used unencrypted LIVE/Message streams, 20 ms SRT latency, and TLPKTDROP disabled. Artificial loss and netem delays were off.
The 320 Mbit/s figure was the specified load, not a measured maximum transport capacity. These measurements do not support general claims about encrypted connections, WAN links, other platforms, or an arbitrary number of streams.
Further Changes for More Reliable Operation
The work went beyond CPU and memory. A send cursor avoids repeatedly scanning packets that have already been sent but not yet acknowledged. Bounded work per channel pass and rotating visits to connections distribute processing more evenly. We handle temporary UDP backpressure with bounded retries; send accounting and pacing advance only after a successful handoff. Reusable callback workers avoid repeatedly creating threads.
Measurement Also Makes the Limits Visible
Not every plausible optimization produced a measurable gain. Removing a redundant scheduler notification passed the functional checks but showed no demonstrated CPU advantage over its immediate predecessor.
Improving tolerance of packet reordering, further parallelizing work at the shared listener, and batching receive system calls are separate tasks outside this release.
Robotweax SRT 0.2.6 is therefore a first step in performance optimization, with further investigation to follow.
Robotweax SRT 0.2.6 is available: as a GitHub release with signed Windows SDKs and through our Homebrew tap with a bottle for Apple Silicon on macOS 15.

