Project

General

Profile

Actions

Feature #9111

open
LS LS

dpdk: support packet backlog to smooth out heavy bursts of traffic

Feature #9111: dpdk: support packet backlog to smooth out heavy bursts of traffic

Added by Lukas Sismis 8 days ago. Updated 8 days ago.

Status:
In Progress
Priority:
Normal
Assignee:
Target version:
Effort:
Difficulty:
Label:

Description

Especially worthwhile on NICs with low rx-descriptors (e.g. Intel X710/E810 i40e/ice),
When lot of packets arrive at once, NICs, especially ones with low RX descriptors (Intel i40e, ice), can become saturated easily. In that scenario, af-packet seems like a more performant option. However, AFP has an internal buffer, configured through ring-size directive which is filled up by the kernel/driver thread, even in workers runmode.
On the other hand, when ehavy traffic comes and DPDK workers don't RX packets quickly enough, the RX descriptors can be quickly occupied.


Files

afp-optimize-iface.sh (17.6 KB) afp-optimize-iface.sh Lukas Sismis, 09/22/2026 03:36 PM
suricata.yaml.afpdpdk.battle (95.2 KB) suricata.yaml.afpdpdk.battle Lukas Sismis, 09/22/2026 03:36 PM
test4.rules (76 Bytes) test4.rules Lukas Sismis, 09/22/2026 03:36 PM
suri_microburst_udp.py (6.26 KB) suri_microburst_udp.py Lukas Sismis, 09/22/2026 03:41 PM
trex_it1_1_cfg.yaml (862 Bytes) trex_it1_1_cfg.yaml Lukas Sismis, 09/22/2026 03:41 PM

LS Updated by Lukas Sismis 8 days ago Actions #1

Setup

The proposed improvement can be demonstrated using Trex/Suricata setup. In my instance, those are 2 machines connected together (over switch but with the right MAC setup, it should not matter).

TREX.port0 -> switch -> Suricata.mlx0

I compared both AF-Packet + DPDK and tried to simulate heavy burst of e.g. 100k packets.
AF-Packet have more favorable settings, it has 8k descriptors, in comparison to DPDK's 256 or 1024 descriptors. I've noticed this later in the process but the improvement is still very well visible when comparing DPDK with/without the backlog.

Suricata setup with configured hugepages:

./autogen.sh && ./configure --enable-dpdk --enable-hwloc CFLAGS="-O3 -g" && make -j10

# Run afp-optimize-iface.sh before af-packet runs

sudo ./src/suricata -c suricata.yaml.afpdpdk.battle -l /tmp/ -vvvv -S ./test4.rules --dpdk
sudo ./src/suricata -c suricata.yaml.afpdpdk.battle -l /tmp/ -vvvv -S ./test4.rules --af-packet

Trex setup, from the installed location, e.g. /opt/trex/v3.06:

Window#1:

# allocate hugepages appropriately
echo 1024 | sudo tee /sys/devices/system/node/node0/hugepages/hugepages-2048kB/nr_hugepages

sudo ./t-rex-64     --cfg /tmp/trex_it1_1_cfg.yaml     -i --stl     -c 10

Window#2:

./trex-console -p 5601

# Then in trex interactive console shell 
start -f stl/suri_microburst_udp.py -m 1 -d 10 -p 0   # single UDP flow, with heavy burst and a cool-off period afterwards. 

Results

AFP with the maximum number of descriptors (8192) receives all replay packets. With 1024, it loses packets ~20%.
Initially I thought it is because of the ring-size setting (AFP Suricata.yaml option) but better AFP results are likely caused thanks to higher amount of RX descriptors.

DPDK with only 256/1024 descriptors and no backlog loses packets (~20% drop rate). With the backlog of 131k elements, 1024 descriptors receives all packets while 256 descriptors only lose ~0.5% packets.

Notes

- Mind that the results can be severely different on different machines but though manipulation of the Python scapy script, you could potentially reach similar results.
- This is a reproducer that works with lower descriptor counts -- it would be better if there would be one that works on 32k descriptors (maximums of ConnectX card) but I believe it demonstrates the case.
- Suricata uses 10 cores but one is sufficient as it is a single flows and only lands on one worker.

Trex machine: Intel(R) Xeon(R) Silver 4210
Suricata machine CPU: Intel(R) Xeon(R) CPU D-1557

Actions

Also available in: PDF Atom