Skip to content

saving >$400/mo on atproto infra

(eventually) trying to make operating atproto infra cheaper

nate
Jul 31, 20269 min read12 reads
19
13

so, i work for a startup and they pay me to maintain open source. the stability and freedom to explore hobby projects (that maybe feed back into work™️) is a privilege i don't take for granted

i am a chronic maker of things. when things go well, they cross-pollinate in constructive ways... but things don't always go well (or at least, go well fast enough). or even if they do go well, there's some lurking untenable nature to the things that i am blind to.

lately, atproto libraries and services have been the things i make*1 outside of work. for a while, i didn't have that many things, so i wasn't too worried about a fly.io machine here, a hetzner node there. i have a history w IaC and claude knows its HCL very well, so its almost too easy to spin up new stuff.

of course, this mode of excess has a practical limit for those with normal finances. i am from a middle-class michigan family, my only money is from my startup tech job at prefect.io. i have disposable income, but not quite as much as i was beginning to spend on compute and storage for atproto stuff.

this is a short recap of my excess, audit and reduction of costs

up through jan 2026

i already had a fair number of web apps up on the internet:

at the time, the most expensive 2 were

  • plyr.fm: there's moderation on every upload, a whole staging deployment (copy of the prod app, fly.io) ~$30/mo
  • coral: i needed a perf machine to do NER, ~$20/mo

so about $50/mo in total. not horrendous

but then..

feb-may 2026

i just start deploying all sorts of shit

in late feb, i deploy my first relay and jetstream instance*2 (reading back through that, wow have i learned a lot about atproto sync since then!) and about a week later, i deploy zlay.waow.tech*3 (an instance of my zig relay). both of these are deployed on Hetzner nodes in the USA*4. this will become important later

so now i had relays, so i wanted to see how well they're doing!

how do you even evaluate a relay?

i ask... and @mackuba.eu answers with pulsar:

but i didn't really understand how exactly it worked. "% of max what?" etc. plus it was in ruby 🫪 so i made my own!

relay-eval

comparing what each atproto relay sees

(fun fact, you can listen to the relays on plyr.fm/radio/firehose)

then, got nerdsniped by @brookie.blog and @bnewbold.net talking about how an independent typeahead would be great but maybe hard

typeahead — community actor search for atproto

Fast handle and display-name search across the whole atproto network, not just one appview. A drop-in replacement for bluesky’s typeahead endpoint.

at first, i underestimated this one. oh "just listen to the firehose, index people as they do stuff" i say! "just use a single fly machine and turso.tech!" i say. just "copy all bluesky's moderation decisions" i say! we'll come back to this one as well :/


i then deploy tangled.org/zzstoatzz.io/my-prefect-server to (yet another US node! on) Hetzner to manage all of my cron jobs, which were then simply grabbing stuff from github to build a lil dashboard to keep track of OSS triage and such. we'll also come back to this

atmosphere conf happens in late march, which is invigorating and makes me want to get even more involved with more things

while literally at the conference in Vancouver, I:

jeez. even writing this out is painful. it snowballed so badly. april and may, i finally slow down but haven't audited my costs yet.

a warning shot comes in late may, as typeahead gets overwhelmed and i realize that i should not just throw more compute at it. this is the beginning of a good trend but i needed to come to jesus first.

the comedown

on june 16th, i finally realize*5 that i really really need to know the scope and scale of my spending, as at this point its spread across fly.io, hetzner, neon and turso and i know its getting to be a lot.

so.... yea i was at about $709/mo on atproto projects

yes i know! i know. its bad. like 0th world, 99.9th percentile bad. believe it or not, i am mostly quite frugal in my life. i think this just got so out of hand because of my frenetic excitement about atproto, the fact that the the costs were spread across platforms, and because of some serious blind spots i'll get to in just a second.

worse yet, i/others had actually started to use these things! especially in the case of typeahead, plyr.fm, pub-search and my relays (which had listReposByCollection, unlike most indie relays). so just saying "sorry y'all" and taking stuff offline felt not ok.

where were all of these costs coming from?!?

unfortunately, a huge chunk of this was entirely avoidable*6.

guess what those top 3 have in common?

yep. US hetzner nodes. biiiiiiiiiiiiiig old blind spot. not ideal

trying to get back to a sane place

on june 17th, i migrated my 3 hetzner nodes to fsn1 (Falkenstein, Germany), immediately cutting $352 dollars off my monthly bill.

just moving them to the EU saves 70% on their compute spend on the spot, and also the bandwidth budget goes from 3TB to 20TB a month. relays, since they subscribe to all ~3k PDS hosts and broadcast the firehose, were exceeding that 3TB every month, causing overages.

there's a long tail of pretty cheap stuff, some were old fly.io things i just deleted, some were expensed work-related things.

this was a serious wake up call for me! i had been less rigorous at deployment time than i ought have been. i think it was a combo of

  • excitement to participate as a relay operator -> let guard down
  • allowed myself to be sold snake oil by claude*7 who dramatically undersold the cost delta of deploying to US hetzner nodes

so i was down to ~$350/mo, which is still a lot. so i kept going

typeahead and pub-search were next, which are similarly structured.

both ingest the firehose to get updates, had a poor fly.io machine doing O(corpus) work per query, and using turso (read as SaaS for globally replicated SQLite) for persistence, cloudflare for pages and handling requests to the backend

typeahead is almost inherently work-intensive, as it intends to be a full-network actor index that's constantly trying to keep up with account status commits and moderation decisions.

pub-search also bc standard.site is kinda popping off and indexing those records*8 and offering fast/durable keyword and semantic search of the actual post content at scale is non-trivial.

the slow road to fig-ification

i use this (made-up) phrase loosely to refer to @bad-example.com 's pattern of serving network scale services off their raspberry pis at home, ie communal services via home infra

well actually, typeahead had this shape until that warning shot / response i mentioned. i realized that there was more of a lambda pattern thing i could do*9, where i can avoid doing corpus-proportional work*10 by doing periodic index preparation "offline", and having a live overlay to keep what i serve up to date.

and i thought

huh, if only i knew a way to do periodic jobs that could keep my indices fresh

....

i literally have spent the last 4 years maintaining prefect, which helps people define and run periodic jobs. of course, there are so many options here, but its not every day that a truly motivating dogfooding use case for your occupational focus appears naturally out of your hobby horse..

so naturally i:

  • first rewrote the prefect server in zig so i could downsize my hetzner node for the server (~parity, 250MB RSS -> 15MB)*11
  • moved all compute-heavy workloads to be prefect jobs that i can schedule or trigger with events
  • deleted the fly machines i had for compute-heavy periodic work

now

i now spend about $250/mo out of pocket on atproto projects*12

which i now track and publish to my PDS (thinking that someday i'd like to solicit donations for public infra, and having an auditable history feels like a good way to stay reputable for donors)

compute heavy index maintenance for typeahead and pub-search now happen on my laptop at home*13, which gets sent work from my zig prefect server running on a small EU hetzner node, and publishes snapshots to R2 that a now slim fly.io backend can surface to the edge to satisfy the needs of the blogosphere and actor searches.

i still have a ton of headroom on my machine (which someday soon i'd like to move to a proper homelab situation), and my IO isn't crazy either (under the ~1TB ISP limit by a lot), so future compute heavy stuff can really just go on my machine for now. of course, power outages will be a thing, but i've decent fallbacks*14 for that.

conclusion

i am running out of steam here. all this steeping in my own embarrassment at excess makes me want to go for a walk on the 606 or play guitar or make 3 quesadillas and eat them over the cutting board.

however, i'm pleased w some of the outcomes of all of this:

  • i learned a ton about how relays do and do not work, mostly while my indigo relay was totally borked. having to go dig around in the microcosm discord or ask @bad-example.com if my grafana dashboard makes any sense at all. similarly, learning more of @sri.xyz 's work as an operator but also vsky.network, and joining @firehose.club and responding to fig's call for feedback from independent operators for their talk at atmosphere conf
  • i nerded out w @calabro.io about zig because of the zig relay, which prompted me to make zat, joining the ranks of the atproto sdks, and learn a bunch about io_uring and musl*15
  • typeahead.waow.tech now serves 50+ atproto applications (about 10 with non-trivial traffic) with actor search, both for login and actor search within apps (e.g userinput.app)
  • i have learned how to make network-scale, community services much cheaper. not only have i learned how, but i've also drastically increased my intolerance for waste (energy and money). e.g. no more sugary deployment patterns for big stuff, no more python runtimes where zig will do etc

everything i mention here, i'll continue to maintain and make cheaper as long as i can. i have much to learn about databases, and i am starting to see why apparently every bluesky team member is working on databases somehow*16. also kafka probably

if you'd like to sponsor future work (e.g. permissioned data might be a significant boon for the economic viability of atproto, and pds.zat.dev already implements spaces along with @trezy.codes @ngerakines.me and @chadtmiller.com projects [happyview, atproto-pds, pds.js], so that apps can explore permissioned data UX now) then please reach out! i currently work at prefect as i mentioned, and have chosen to pay for these things out of pocket of my own volition bc i care about atproto, but am interested in doing atproto-stuff full time sustainably.


  1. why atproto stuff has been the stuff i'm interested in working on lately could be its own post. but suffice to say, i think healthy and stable information channels are both critically important to society and extremely scarce. i think atproto can help fix this! and as a software infra guy, i am compelled to proof stuff out ↩
  2. these were on the same Ashburn, Virginia cpx31 Hetzner node, which got bumped to a cpx41 and later a cpx51 (as i experimented w different operating config, like replay windows, go mem limit etc) ↩
  3. this was a cpx41 in Hillsboro OR ↩
  4. the repeated pattern here is terraform deploying single-node k8s on hetzner via k3s. just familiar and easy for me personally ↩
  5. both for the obvious reason that i know its gone too far, and because i was/am considering making a change that'd throw my financial stability/freedom into question ↩
  6. meaning, in a way that's totally separate from properly doing budgeting like a normal person. i know, i know, i know ↩
  7. claude/codex is always trying to sell me snake oil, i know. but my takeaway was that i need to exercise more caution when im outside of my wheelhouse (i'd only been using hetzner for a couple months and mostly via HCL and not personally reading docs) ↩
  8. which may or not contain the textContent i need to offer search of post content, and may or may not be complete nonsense, like the hundreds of thousands of random webpages bridgyfed sends or literal car patents ↩
  9. i haven't totally gotten into kafka yet, but maybe soon ↩
  10. meaning, things like full table scans when you have 10 million entities, which you literally have in the case of typeahead, and could feasibly have in the case of pub-search ↩
  11. i say this like im joking, bc its funny w the bun thing, but i really did do this. i had been planning it for a couple months, before the bun stuff even happened ↩
  12. the rest here are not atproto-related or are work-related or subsidized ↩
  13. ubuntu system76, Intel i9, 24 cores, 64 GB RAM, 1 TB SSD + 2 TB NVMe ↩
  14. typeahead just serves the last snapshot, but if it gets stale, it can switch over to the bsky endpoint heh. and pub-search has a similar deal, it just gets stale if my periodic job doesn't happen.  ↩
  15. just in time for the zeitgeist on zig to sour as andrew kelley went full mean girls on jared sumner over bun, among other things ↩
  16. hyperbole but afaik paul is doing some foundation DB craziness and jim is doing jetstream v2 segment stuff, and idk how those intersect! need to go deeper ↩

Did you enjoy this article?

Recommend it — Standard Reader surfaces well-loved writing to more readers across the network.

Across the AtmosphereDiscussions