Rust's derive often implies inline

(yossarian.net)

104 points | by woodruffw 3 days ago

7 comments

  • Sharlin 6 hours ago
    I’m fairly convinced that Debug should never be inlined. Display probably neither, the fmt machinery is heavy enough that not inlining is probably not a bottleneck even in serialization-heavy workloads. I’ve had to #[inline(never)] some of my own Debug/Display impls, shrinking the binary by tens of kilobytes (out of a few hundred, so relatively a significant reduction).
    • aw1621107 6 hours ago
      For what it's worth, according to the PR that added the annotation [0] doing so generally resulted in decreases in compile times and binary sizes on benchmarks. Furthermore, an additional experiment that avoided emitting the inline attribute on structs with >5 fields resulted in benchmark regressions compared to always emitting the attribute [1]. I'd guess this is one of those things which may help in aggregate but hurts for specific cases.

      That being said, one of the Rust devs indicated in the corresponding lobste.rs discussion [2] that they're open to revisiting/rebalancing things if they get enough bug reports indicating something is up, so it might not hurt to tag onto the bug report the author will (hopefully) eventually submit.

      [0]: https://github.com/rust-lang/rust/pull/117727

      [1]: https://github.com/rust-lang/rust/pull/118031

      [2]: https://lobste.rs/s/dldhpw/rust_s_derive_often_implies_inlin...

      • afdbcreid 5 hours ago
        A reasonable conjecture was raised on lobsters that this is because `#[inline]` makes actual codegen (LLVM IR and down from MIR) lazy, and most `Debug` impls are never used.
        • saghm 2 hours ago
          That makes sense. If you're a library, it's just good manners to derive Debug on types that you don't have a good reason not to, but you have no way of knowing if anyone will actually use it in a given program.
        • infogulch 4 hours ago
          How much code is never used and compilation could be skipped entirely? Maybe applying a reachability pass to skip compiling unused code would be helpful.
          • saghm 2 hours ago
            I've said for a while that I think that the "real" issue with compile times is that there's a ton of dead code getting compiled across the dependency tree. I personally blame Cargo features not being ergonomic enough, because on paper they're perfectly suited to solve this problem, but in practice the amount of boilerplate needed to use them for this is infeasible. I wrote a manifesto about this a while back: https://saghm.com/cargo-features-rust-compile-times/
            • infogulch 35 minutes ago
              So the thesis is that crates massively oversubscribe the features of their dependencies, and that if features were managed more carefully and maybe dependency features exposed through each crate's own exposed feature list, that would cut down on dependency bloat and compile times a lot.

              My idea above is roughly equivalent to the "zero-config" strategy you mention.

              It's kinda crazy that if a direct dependency doesn't use some feature from a transitive dependency and the direct dependency forgets to disable it or expose a way to disable it you're still stuck with it.

          • mgsloan2 3 hours ago
            A cross-crate dead code analysis would mean that compilation of a crate now depends on information about its dependents. This would break reuse of compiled crates and cause recompiles when the analysis changes.

            Something does seem a little off about this, though. Ideally for this `Debug` case there would be an annotation that says "compile this lazily, don't inline". Maybe there doesn't even need to be a new annotation, just `#[inline] #[cold]`. Which looks pretty weird, but might work already.

            • afdbcreid 1 hour ago
              `#[inline]` in Rust is weird and not similar to `inline` in C or C++. It has two effects: it duplicates the MIR for every codegen unit that needs it instead of precompiling (which is what causes the laziness, and also why it's commonly recommended to avoid `#[inline]` on generic functions as this happens anyway for them), and also adds LLVM `inlinehint`. It is possible to avoid the second but not the first if you want to stay lazy, and that can also cause code duplication (unless you set `codegen-units = 1`) and LLVM deciding to inline things even without the hint.
            • afdbcreid 1 hour ago
              And `#[inline] #[cold]` doesn't work: it adds both `inlinehint` and the cold hint to LLVM. Which may cause, for example, the function to be inlined but the branch be marked cold for better machine code organization.
      • Sharlin 6 hours ago
        Thanks, interesting!
  • vlovich123 5 hours ago
    I feel like Debug should be lazily emitted altogether when it’s first used - just a special marker that’s never expanded since 99% of the Debug implementations aren’t used and having the rest marked #[cold] as inline is obviously wrong. Of course implementing it in practice sounds exceptionally difficult.

    That being said, even the justifying performance improvement PR was itself a mix of improvements and regressions

    • dwroberts 1 hour ago
      The attributes on it seem irrelevant if it’s never called since it is presumably removed by dead code elimination?
    • bloppe 3 hours ago
      Lazily emitted from what? You would need information about the type at runtime to derive the implementation, and normally that information is not available at runtime. The debug implementation itself is probably within spitting distance of any other representation that would be sufficient to lazily emit the debug implementation.
      • vlovich123 2 hours ago
        You say runtime but that's ambiguous in this context. My point is that what gets emitted for Debug is just "type X lazily implements Debug" and you never even generate the equivalent Rust code for it until someone uses it and at that point you mark it with cold & inlineable if it's a small function.

        > within spitting distance of any other representation

        If marking Debug as #[inline] vs #[noinline] has an effect, it's clearly already true that there's a lot of time being spent on emitting the Debug trait eagerly and having it go through all the compiler stages.

        • bloppe 7 minutes ago
          Seems like you could either generate the implementation "eagerly" at compile time (possibly with tree shaking to eliminate it automatically if it's not actually referenced anywhere), or you could generate it "lazily" at runtime, which might be useful if you think it's very unlikely but technically possible that the function might ever be called, and you won't know until runtime.

          Note that Rust already does dead code elimination at a couple different levels by default, and generic link-time optimization is always an option on top of that, so I assumed you weren't just referring to that.

      • fluoridation 1 hour ago
        Pretty sure what is meant by "lazily" is that it should just be added by default and then those that are not used should be optimized away by the linker.
        • bloppe 4 minutes ago
          Rust already does dead code elimination at a couple different levels by default, and generic link-time optimization is always an option on top of that, so I assumed GP was talking about something that was not already the status quo.
      • Lvl999Noob 3 hours ago
        You can do it at compile time. Or more likely, link time. If the implementation is ever actually used, it is kept in the binary. Otherwise it is used. The compiler can do the same, keeping all derives as just markers until it finds a place that actually uses it then firing off a background worker to compile the derive impl.
        • K0nserv 2 hours ago
          The compiler already does this.

          Compare https://rust.godbolt.org/z/osMY8777W and https://rust.godbolt.org/z/azGjbbzc8. I'm not sure why the inlining doesn't happen in the case when we do use the Debug implementation though.

        • saghm 2 hours ago
          My instinct is that this would end up being more costly than it's worth. You'd be essentially adding extra bookkeeping and logic for every single type that needs to remain alive as long as the type is still possible to reference.

          Moreover, how would you deal with downstream usage by dependents? Should I be able to make my dependency's type implement Debug (which is at least in spirit a violation of the orphan rule, and would make the bookkeeping/extra logic described above explode for any non-trivial dependency tree)?

          People already find the Rust compiler too slow. If there's budget for adding more expensive checks, I don't think I'd want it to be spent on something like this.

          • vlovich123 2 hours ago
            The vast majority of Debug traits never get used. All the standard library containers (collections, smart pointers, refcell, etc) have Debug traits auto-emitted if the underlying type has Debug. Even if the type never gets formatted.

            I really doubt lazily code genning the Debug trait makes things worse or is super expensive to track. The main tension is more how this interplays with LTO (which I'm guessing would be unable to do this) and designing a generic language feature here for macros.

            A rough sketch of the idea would be that you can annotate functions with #[lazy] and the Rust compiler knows to save off the function body for later and only keep the function definition around (no codegen). Then if it does need to codegen because it gets invoked from a non-#[lazy] function, the function is treated as inlineable and #[cold] so that if it's not inlined it's put in a cold region of the binary. It's non-trivial because you somehow need to put the codegen back into the original crate's object file which is probably tricky.

            I don't understand the concern regarding the orphan rule or "downstream usage by dependents" - that's not relevant here and no language semantics change.

  • scottlamb 4 hours ago
    I wonder if they ever considered a table-driven approach for `#[derive(Debug)]`, as `facet` [1] does. That would have been my first instinct for something this formulaic where binary size and compilation time matter more than execution speed. But my impression is facet hasn't quite realized its promise on those fronts, so maybe the table-driven approach in std was similarly tried and rejected.

    [1] https://crates.io/crates/facet

  • zamazan4ik 6 hours ago
    Or just try to avoid all of these optimization guesses by using Profile-Guided Optimization (PGO), that inserts/deletes all inlines based on actual application runtime profile.
    • woodruffw 6 hours ago
      The codebase in question (uv) uses PGO already. I suspect there isn’t a general way to guarantee that PGO ensures that only the “right” things get inlined.
  • api 5 hours ago
    There's a joke that the LLVM heuristic for whether to inline a function is "return true;" LLVM tends to inline aggressively.

    You can control this behavior with opt level "s" or "z" or "#[inline(never)]", but be aware that too little inlining can have large negative performance impacts.

    It's hard to get inlining exactly right without profile guided optimization.

    • pjmlp 1 hour ago
      Same applies to something somehow related, devirtualisation.

      Across all languages with compilers that do it, and even with a mix of PGO data, a data set that brings another dynamic dispatch workflow into action might mess it all, by going into the else part of the optimization.

  • payne-research 2 hours ago
    [flagged]