We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Admit when you’re wrong
almirsarajcic
People in the Elixir community shouldn’t be butthurt about DHH’s Campfire benchmarks. We should use them as an opportunity to look at the state of the ecosystem and review some of the assumptions that brought us here.
We chose Elixir for reasons beyond raw performance. We wanted supervision, crash recovery and the other things that make OTP useful, and we accepted compromises to get them. But we should be able to reconsider those compromises when the way we build software changes. The reasons I chose Elixir years ago might not be enough for the application I want to build now.
# Campfire: HTTP throughput
Requests per second · higher is better
| Workload | Rails | Django | Laravel | Elixir | Go | Rust |
|----------------|------:|-------:|--------:|-------:|-------:|-------:|
| Room page | 242 | 170 | 164 | 722 | 3,860 | 36,260 |
| Messages page | 402 | 196 | 175 | 1,053 | 5,573 | 40,872 |
| Sidebar | 541 | 615 | 715 | 1,275 | 19,753 | 34,672 |
| Search | 424 | 315 | 305 | 1,156 | 7,053 | 33,299 |
| Post a message | 225 | 154 | 137 | 801 | 4,767 | 6,896 |
16 concurrent clients · AMD Ryzen AI MAX+ 395
Four hardware threads per app
Source: basecamp/once-campfire · checked 5 October 2026
Published implementations, not our measurements.
Source: Campfire’s published comparison.
The published Campfire comparison gives us something worth investigating. On the room page, it reports 722 requests per second for Elixir and 36,260 for Rust. Those are results for particular implementations under particular conditions, and the measured Rust version was already optimized. They don’t tell us what the best possible Elixir implementation could do, or which language an agent handles best with an equal amount of time and guidance.
Still, I don’t think we should see that gap and immediately start looking for reasons why it doesn’t matter.
One response is that the agents wrote bad Elixir. Fair enough. There are concrete problems in the implementation, and people are already working on them. But if agents struggle to write good Elixir, that also matters to someone deciding whether to use Elixir in a new application.
Explaining what an Elixir expert would have done doesn’t remove the need for that expert. Someone still has to recognize the problem, explain it to the agent, review the correction and check whether it actually helped. That’s time, tokens and knowledge that we should count when comparing stacks.
I want to make product decisions and have agents do more of the implementation. I want to explain what the application should do, how it should behave and what its users need. I don’t want to keep going into the code to explain architectural decisions that the agent should have been able to make. If I have to keep nudging it to review its own work and find mistakes I already suspect are there, that isn’t the workflow I’m trying to get to.
That’s what I’m trying to build with Kogen: I shape the feature and its user experience, specify technical decisions when they matter, and let it handle implementation, checks, independent review and rework. It’s still in development, so I’m describing what I’m working towards. How much help agents need with a particular stack is directly relevant to whether I can make that work.
Maybe that’s too much to expect from agents today. But if they get closer to it in another stack, I want to know.
I like Elixir, and I’ve never written a line of Rust. If agents consistently produce better software in Rust with the tokens and subscriptions I have, I would choose Rust. The same goes for Go or whatever else works better for the application. My preference for writing Elixir matters less when I’m doing less of the writing, while the cost of getting a working application still matters very much.
That doesn’t mean I can pick the biggest number in a table and stop thinking. I still need the application to behave correctly, and someone has to maintain it when requirements change or something breaks. Those costs belong in the comparison too. But we can’t keep making the comparison as if I’m going to write every line myself.
There are people in the community responding with useful work. Zach Daniel’s changes address database reads serialized through one process, repeated work and WebSocket disconnects. Kurt Melby’s changes investigate caching, pooled readers and other improvements, with a clear explanation of why his local results aren’t directly comparable with the published benchmark.
I’d like to see more of that. Show how Campfire should be written in Elixir, explain the decisions and measure what changes. If a compatibility requirement forced a compromise, document it. If removing it improves performance but changes behavior, document that too. Then we can decide whether the compromise makes sense for the application we’re actually building.
We should also use those fixes to improve what agents produce next time. I think we have some responsibility here. If we know how an application should be built but that knowledge mostly appears in replies after an agent gets it wrong, we need to make it easier to find and use before the next implementation.
That includes ElixirDrops. A useful post can explain why a decision was made, what went wrong with the alternative and when the advice doesn’t apply. Documentation, complete applications and examples of actual failures can give agents more to work with than a rule telling them to write idiomatic Elixir.
Tooling deserves the same attention. I’d like us to look at things as ordinary as test output: what does an agent need to understand a failure, how much irrelevant output does it have to read, and can we make fixing the problem cost fewer tokens? We spend time improving the experience for developers. If agents are doing more of the work, their ability to use our tools matters too.
I don’t know whether publishing more examples will improve future model training. I do think we can test whether supplying better examples and instructions improves the agents we have now. Give them the material, see what they build, and check whether the same mistakes keep happening.
But we have to leave room for the result we don’t want. We might improve the implementation, improve the guidance, give Elixir a fair comparison and still decide that another stack is a better choice for some new applications. We should be able to say that on an Elixir website without treating it as a betrayal.
The circlejerk makes that conversation harder. If we only accept benchmarks when Elixir does well, or answer every weakness with a list of unrelated strengths, we’re making it harder for developers to judge the tradeoffs. Someone considering Elixir needs to know where it causes problems as well as where it helps.
I’ve invested a lot in Elixir, so this applies to me too. I don’t want that investment to become a reason to keep recommending it after my reasons have stopped making sense. I’m not married to any stack, and I’d rather reconsider my choices than encourage someone else to spend years making the same investment without questioning it.
For my next application, I want to know how much help the agent needs, what the resulting software does well and what it costs to keep it working. Elixir may still be my choice. I haven’t done the comparison needed to decide that yet.
I’m going to try building Campfire myself and write about what I learn in a separate post.
copied to clipboard