Skip to content

Performance

Apple M4 Max. PostgreSQL and MySQL run in Docker on the same machine; SQLite writes to the local SSD with synchronous=NORMAL. Medians; see the benchmark code for ranges.

PostgreSQL MySQL 8.4 SQLite (modernc)
Bulk insert, 10k jobs per call ~250k jobs/s ~77k jobs/s ~300k jobs/s
Claim + finish, 50 per fetch ~45k jobs/s ~15k jobs/s ~80k jobs/s
One server, 100 workers, no-op handler ~21k jobs/s ~6k jobs/s ~70k jobs/s
Enqueue to handler start p50 3.3ms, p99 6.6ms ~5–20ms with redisbus, PollInterval without p50 0.16ms in the same process

The MySQL numbers are dominated by commit latency on this setup (binlog with sync_binlog=1, 4-6ms per commit); a server with a faster disk will do noticeably better.

Compared with River v0.47 on the same PostgreSQL, alternating rounds between the two libraries (bench/):

Scenario kiln River (defaults) River (1ms fetch cooldown)
Bulk insert 251k jobs/s 134k (InsertMany), 222k (InsertManyFast)
8 goroutines inserting one job at a time 5.1k jobs/s 4.4k jobs/s
Drain 50k no-op jobs, 100 workers 21.2k jobs/s 1.0k jobs/s 20.3k jobs/s
Drain 20k jobs with a 1ms handler 19.9k jobs/s 1.0k jobs/s 13.7k jobs/s
Enqueue to handler start, p50 3.3ms 55ms 4.7ms

River's defaults fetch at most once every 100ms, which is what caps it around 1k jobs/s; with that cooldown lowered the no-op drain is a tie, and the p99 latency of the two was too noisy on this machine to call either way.

Under a production-like load

The numbers above are single-purpose benchmarks. To see kiln in something closer to production, a small shop backend runs the same code on all four stores: PostgreSQL and MySQL with an API process and two workers, SQLite and memstore in one process. Placing an order writes it and, in the same transaction, enqueues a flow: a payment job under Limit{Max: 4, Rate: 20, Per: time.Second, Burst: 5}, then a confirmation e-mail and an invoice that wait for it. The payment gateway is a fake that rejects more than 20 requests per second or 4 at a time, fails 15% of charges and leaves PIX payments pending. Servers run with a 2s heartbeat, DeadAfter 15s and LeaderTTL 6s; MySQL polls every 200ms and has no bus.

v0.3.1 v0.4.0
Busiest second at the gateway, 150 orders from 15 clients up to 40 requests (429s) 24–25 on every store, no 429
Same, with the limits row held for 1s mid-run (PostgreSQL) 40 25
Payment job, enqueue to start, MySQL p50 684ms 72ms
Continuation, parent done to child start, MySQL p50 61ms 46ms
Same two on PostgreSQL 6ms / 5ms 7ms / 6ms
Export resumed after its worker got kill -9, MySQL 41s 14s
Same with SQLite, where the one process restarts 38s 24s

Rescued jobs now restart within 100ms of being rescued; the rest of those times is noticing that the worker is gone (DeadAfter, plus LeaderTTL when the dead server was the leader). Across the runs, 25 buyers racing for 5 units of stock always ended with 5 orders, 20 rejections and no job left behind for a rolled-back order, and every charge reached the gateway exactly once.

Under chaos

soak/ runs worker processes against PostgreSQL or MySQL for as long as you ask, while it kills them with SIGKILL, stops them with SIGTERM and restarts the database, then checks that every job reached a final state, none ran more than its attempts plus the runs cut short, limits and rates held, and the store needs no repair. The jobs mix plain ones, retries, jobs that always fail, fan-in flows that read their parents' outputs, slow jobs, and jobs under a mutex, a Max, a Rate and both.

A 15 minute run on PostgreSQL, 4 workers, 50 enqueues per second:

Jobs 57,211
Chaos 31 SIGKILLs, 22 SIGTERMs, 3 database restarts of about 1.5s
Jobs lost 0
Jobs run more often than allowed 0
Peak concurrency per limited key at or under its Max
Store after the run limit counters at 0, nothing for Sweep to repair

A 24 hour run on PostgreSQL 17, 4 workers, 20 enqueues per second, on a 2 vCPU / 2 GB DigitalOcean droplet with kiln v0.8.0:

Jobs 2,233,828
Chaos 2,886 SIGKILLs, 2,139 SIGTERMs, 360 database restarts
Jobs lost 0
Jobs run more often than allowed 0
Orphaned jobs rescued after a SIGKILL 2,307
Peak concurrency per limited key at or under its Max, over about 85,000 runs each
Starts per window on the Rate keys at most 6 of 8 allowed in 1s, 32 of 34 in 10s
Store after the run limit counters at 0, nothing for Sweep to repair, drained in 7s
Memory of the harness and its 4 workers 44 to 58 MB across the 24 hours

The workers logged store errors around each database restart (claims, heartbeats, leases), as expected; none of them cost a job.

cd soak && go run . -duration 24h -rate 20

soak/vps runs the same thing on a server with nothing but Docker: a compose file with its own PostgreSQL, which keeps the summary, the worker logs, CPU and memory every 5 minutes and a dump of the database in soak/vps/evidence.

See soak/README.md for the flags. -container none skips the database restarts when other tests share the database.