4 comments

  • mjevans 1 day ago
  • spaceywilly 1 day ago
    That’s very interesting, thank you for sharing. Looks like it could be a very useful tool for testing high performance networking.

    I wonder if something similar could be done using TC BPF instead of AF_XDP? My only reservation about AF XDP is that it requires a special NIC to support it, so it may not be useful for a “regular Joe” user. I wonder if TC BPF would also work since it similarly bypasses the Kernel networking stack, I believe you can put packets directly into the NIC TX queue for transmission

    • tptacek 1 day ago
      You can, but the interesting thing about AF_XDP is that you've got a userland path to writing directly to the card's DMA buffers; TC BPF still allocates an skbuff for every packet you send.
      • adrian_b 1 hour ago
        Also with liburing you have zero-copy send and receive operations for normal protocols like UDP or TCP, but this requires a NIC that supports scatter/gather DMA (so that the packet headers go to/from kernel buffers, while the data goes to/from userland buffers).

        AF_XDP is also available in older kernels, but with recent enough kernels (zero-copy receive is a recent addition) and with a good NIC, liburing should provide a similar performance.

    • Palomides 1 day ago
      it seems like every NIC on the market that can do 100Gb has support in its linux kernel driver, so probably not a big deal in practice
    • bgpdude 1 day ago
      you can use generic af_xdp which sits at the TC layer. Just get a bit less performance.
      • bgpdude 1 day ago
        That's what the veth examples are using
  • bgpdude 23 hours ago
    The video demo is pretty neat as well https://www.youtube.com/watch?v=8EmxvaOR5uA
  • barryvand 23 hours ago
    ooh no more DPDK, this will make it a lot easier. I just started working with trex but it's all really complicated. Going to give this a try
    • tptacek 21 hours ago
      I think, and I'm saying this in part to get someone to correct me, that post-XDP (so 5 years or so now) DPDK is basically obsolete. Is there a circumstance where it would make sense to start from DPDK rather than XDP?
      • touisteur 9 hours ago
        I think access to offload engines is a big part of the appeal of dpdk still, especially for me all the GPUdirect nvidia-only packet steerer.

        I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.

        Some silly stuff in DPDK are very useful for testing too (in-memory devices).

        Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.

        • tptacek 4 hours ago
          I am corrected. Thanks! :)
      • pstavirs 15 hours ago
        At this point there's a much bigger ecosystem for DPDK than AF_XDP I think. Also more people know about DPDK than AF_XDP right now e.g. the Ostinato traffic generator's line-rate Turbo functionality uses AF_XDP but most customers assume it uses DPDK.

        Disclosure: Ostinato creator here.

      • trevex 13 hours ago
        While I am a big proponent of XDP there are use-cases better suited to DPDK: Being able to offload flows and crypto operations to the NIC is important to a lot of use-cases. The first packet to user space is essentially the slow path even with DMA, that sets up the fast path.
      • bgpdude 19 hours ago
        I think you're correct. Not really aware of any real limitations, other than it's slightly slower than dpdk (it's not a complete bypass), but at a much easier ease of use.
        • tptacek 18 hours ago
          AF_XDP kind of is a complete bypass, right? RX get scooped right off the DMA buffer for the card, and TX get shoved right back in.
          • bgpdude 18 hours ago
            Yeah fair point. I should have been more precise. AF_XDP with ZC bypasses the kernel networking stack and the NIC DMA's directly into UMEM and so in that sense it absolutely is a kernel bypass.

            The difference I was getting at vs DPDK is that the kernel NIC driver/NAPI/XDP path is still involvd. With DPDK the userspace PMD is effectively driving the NIC and accessing the queues directly.

            either way, it's great and everyone should use it :) that is assuming they have a use case for it. The use-cases are perhaps somewhat limited as it also bypasses the kernel tcp-ip stack, so you gotta do a lot yourself.