Hi everyone,
I'm an undergrad computer science student in my final year of college. I'm currently in the process of researching and choosing a topic for my senior project. Very broadly, I want to do something related to compilers, and more specifically compiler optimisations. I've been going through papers and articles trying to find ideas and Cranelift came up a couple of times in relation to formal verification and e-graphs. My interest was piqued, and I think Cranelift is a very cool project!
I was wondering if someone here could point me towards some things in Cranelift that I could look into and see if they'd make for a final year project. I figured it would be best if I asked people familiar with the project lest I make things unnecessarily hard for myself by trying to do something against the grain of what's already there. What I have in mind is something along the lines of researching, implementing, benchmarking, formalising, and evaluating some (subset of an) optimisation but I'm open to other, maybe less well-defined, suggestions too because nothing's set in stone yet.
Obviously, I don't expect to be mentored or that upstream takes my work seriously (though I wil be glad if it ends up being helpful, of course). I'm just looking for leads on something I can learn about and work on for the next 6-ish months.
Any help would be greatly, greatly appreciated. Thank you!
@Hari Mohan thanks for reaching out! You're in good company -- a few undergrad theses have already been written on various topics (one on the "chaos testing" infra in Cranelift, one on extensions to aegraphs -- one of @Alexa VanHattum's students for the latter). I think there are lots of interesting open problems.
The most active research front right now is on verification; there are also a number of us working on it (Alexa, @Michael McLoughlin, others) but perhaps there is a separable piece that we could point to, I'm not sure (I'd defer to Alexa/Michael as "project managers" on that).
You mentioned optimization so one thing that might make a very interesting, and also tractable, undergrad thesis is some sort of systematic study of the "optimization gap" between Cranelift and, say, LLVM or V8 or SpiderMonkey. For any given input (say Wasm via some LLVM-based Wasm engine, V8, and Wasmtime+Cranelift), look at the machine code that comes out from each and try to categorize/root-cause reasons for any perf gap. It's a more analysis-heavy thesis topic but one could then try to implement anything missing, or maybe fake it (do it by hand) for certain benchmarks, to categorically understand the tradeoffs. We've wanted a deeper understanding of "what's left to do" for a while and so this would actually help us a bunch too.
Will say more if I think of any other ideas! Best of luck...
@Hari Mohan a couple more ideas:
cranelift-frontend is basically a register allocator if you squint, and we have historically had a bunch of bugs due to missing stack maps due to things like bugs in the liveness analysis that the safepoint spiller uses. it would be really awesome to implement a "safepoint" checker that is similar to regalloc2's register allocation checker but specifically for safepoints and stack mapshappy to go into more details if you anything interests you...
Thank you both so much for replying; I really appreciate your input! Sorry for the delay in getting back to you; I had an unexpectedly busy week, and I wanted to take the time to look into your suggestions before replying.
@Chris Fallin I find your suggestion very interesting and it's very similar to what I had in mind. I read the performance comparison section of the copy-and-patch paper cited on the Cranelift homepage and also found some references to benchmarks in blog posts and GitHub repos from Wasmer. If I understand correctly, there are benchmarks in common use (CoreMark, PolyBenchC) on which it has been observed that there is an execution time gap between Cranelift and other optimising compiler back-ends (TurboFan, LLVM). You are saying that it would be useful to compare, for each of the individiual test cases in the benchmarks, the end-result machine code produced to identify the specific optimisations/categories of optimisations for which Cranelift does not yet have rewrite rules. Once this survey is complete, the second part would be to evaluate the tradeoff between the improved execution time vs. the additional complexity / compile time cost before attempting to implement the best candidates. This last step would include extracting a suitably general rewrite rule, formalising it in ISLE, and benchmarking it again for empirical evidence of the performance gain.
Is that what you meant? I just want to make sure I have everything right in my head.
@fitzgen (he/him) I think your first idea is in many ways similar to @Chris Fallin 's idea above, but instead of comparing machine code across Wasm compilers, it automatically extracts rewrite rules from a Wasm program by comparing the given Wasm to its superoptimised canonical form. I didn't know about this approach before and I think it's fascinating, though I worry that in the context of a final-year project, implementing these rewrite rules in ISLE might not be enough work for the 6 months. Perhaps I'm not fully understanding the scope of what's involved? I took a look at https://github.com/bytecodealliance/wasmtime/pull/10979 to understand the process. The actual code changes in the PR are quite small, though there is quite a large proof included in the PR description. I'm unable to tell how much of the ISLE code is boilerplate vs. how much is unique to this proof. I can imagine that if a lot of it is specific to this rewrite and if most of the Souper-identified rewrites require this level of rigour to verify, then there is quite a bit of work to be done. Is that the case?
I'm not as familiar with the compiler internals referenced in your second idea, and I think I'd like to stick with optimisations for this project. Perhaps one for when I've graduated!
@Hari Mohan a few thoughts on the "perf analysis and followups" track:
@Chris Fallin Thanks for clarifying! I was talking to my supervisor about this and they think it's a good project for me as well. While discussing how I'd go about working on this, a couple more questions came up and I was hoping you could help clarify them.
rustc_perf measures much larger pieces of code, like parsers and renderers, so it seems too complex to use it for more focused comparisons.cg_clif or Rust-LLVM. I'd encourage diversity of benchmarks -- my experience (with regalloc and the mid-end optimizer at least) is that it's easy to overfit if you only have a few.gcc-as-a-benchmark has a big code footprint that is very "flat")Last updated: Oct 11 2026 at 04:10 UTC