That’s very interesting, thank you for sharing. Looks like it could be a very useful tool for testing high performance networking.
I wonder if something similar could be done using TC BPF instead of AF_XDP? My only reservation about AF XDP is that it requires a special NIC to support it, so it may not be useful for a “regular Joe” user. I wonder if TC BPF would also work since it similarly bypasses the Kernel networking stack, I believe you can put packets directly into the NIC TX queue for transmission
You can, but the interesting thing about AF_XDP is that you've got a userland path to writing directly to the card's DMA buffers; TC BPF still allocates an skbuff for every packet you send.
Also with liburing you have zero-copy send and receive operations for normal protocols like UDP or TCP, but this requires a NIC that supports scatter/gather DMA (so that the packet headers go to/from kernel buffers, while the data goes to/from userland buffers).
AF_XDP is also available in older kernels, but with recent enough kernels (zero-copy receive is a recent addition) and with a good NIC, liburing should provide a similar performance.
I think, and I'm saying this in part to get someone to correct me, that post-XDP (so 5 years or so now) DPDK is basically obsolete. Is there a circumstance where it would make sense to start from DPDK rather than XDP?
I think access to offload engines is a big part of the appeal of dpdk still, especially for me all the GPUdirect nvidia-only packet steerer.
I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.
Some silly stuff in DPDK are very useful for testing too (in-memory devices).
Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.
At this point there's a much bigger ecosystem for DPDK than AF_XDP I think. Also more people know about DPDK than AF_XDP right now e.g. the Ostinato traffic generator's line-rate Turbo functionality uses AF_XDP but most customers assume it uses DPDK.
While I am a big proponent of XDP there are use-cases better suited to DPDK: Being able to offload flows and crypto operations to the NIC is important to a lot of use-cases. The first packet to user space is essentially the slow path even with DMA, that sets up the fast path.
I think you're correct. Not really aware of any real limitations, other than it's slightly slower than dpdk (it's not a complete bypass), but at a much easier ease of use.
Yeah fair point. I should have been more precise. AF_XDP with ZC bypasses the kernel networking stack and the NIC DMA's directly into UMEM and so in that sense it absolutely is a kernel bypass.
The difference I was getting at vs DPDK is that the kernel NIC driver/NAPI/XDP path is still involvd. With DPDK the userspace PMD is effectively driving the NIC and accessing the queues directly.
either way, it's great and everyone should use it :) that is assuming they have a use case for it. The use-cases are perhaps somewhat limited as it also bypasses the kernel tcp-ip stack, so you gotta do a lot yourself.
I wonder if something similar could be done using TC BPF instead of AF_XDP? My only reservation about AF XDP is that it requires a special NIC to support it, so it may not be useful for a “regular Joe” user. I wonder if TC BPF would also work since it similarly bypasses the Kernel networking stack, I believe you can put packets directly into the NIC TX queue for transmission
AF_XDP is also available in older kernels, but with recent enough kernels (zero-copy receive is a recent addition) and with a good NIC, liburing should provide a similar performance.
I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.
Some silly stuff in DPDK are very useful for testing too (in-memory devices).
Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.
Disclosure: Ostinato creator here.
The difference I was getting at vs DPDK is that the kernel NIC driver/NAPI/XDP path is still involvd. With DPDK the userspace PMD is effectively driving the NIC and accessing the queues directly.
either way, it's great and everyone should use it :) that is assuming they have a use case for it. The use-cases are perhaps somewhat limited as it also bypasses the kernel tcp-ip stack, so you gotta do a lot yourself.