Skip to content

Operating kiln

ServerConfig Default What it controls
PollInterval 1s How often an idle pool looks for work when no notification arrives
HeartbeatInterval 5s How often a server reports that it is alive
DeadAfter 60s Silence after which a server's running jobs are retried elsewhere
LeaderTTL 15s Lease of the leader that runs recurring jobs, rescue, sweep and pruning
ShutdownTimeout 30s How long Run waits for running jobs after its context is cancelled
KillGrace 5s Extra wait after cancelling jobs that ignored the shutdown
Timeout 30m Default per-job timeout

Recovering from a dead worker. A server that exits cleanly finishes its running jobs, or hands them back, before Run returns. A server that dies is noticed after DeadAfter, and within one more HeartbeatInterval its jobs are rescued: they go straight back to their queues, and the rescued attempt counts toward MaxAttempts. With the defaults, a job whose worker was killed starts again after roughly 60 to 70 seconds. A job rescued twice in a row waits for its backoff before the next attempt, so a job that takes down the process running it is not handed from worker to worker without a pause. If the dead server was also the leader, a new leader takes over after LeaderTTL and waits DeadAfter + HeartbeatInterval before rescuing anything, which adds about 20 seconds. Lower DeadAfter to notice dead workers sooner (it must stay above 3*HeartbeatInterval + KillGrace + 5s), and keep long jobs resumable with SetParam checkpoints.

Deploys. On SIGTERM a server stops claiming and waits up to ShutdownTimeout for its running jobs. A job still running after that is cancelled with ErrShutdown and goes back to its queue without using up an attempt, so another server runs it again from the start (or from its SetParam checkpoints). Set ShutdownTimeout above your longest job, and give your orchestrator a grace period longer than ShutdownTimeout + KillGrace (Kubernetes' terminationGracePeriodSeconds, Docker's stop_grace_period), or the process is killed first and its jobs wait DeadAfter to be rescued.

One queue per service. A server claims only the kinds it has handlers for, so two services can share a database and even a queue without running each other's jobs. Sharing the queue has a cost, though: the claim query walks past the other service's ready jobs to reach its own kinds, about 34ms per 100,000 of them on PostgreSQL. Give each service its own queues, and the walk disappears.

Connections. pgstore opens its own pool (MaxConns, 8 by default) plus one connection for LISTEN, on top of your application's pool: budget up to 9 connections per process that opens a store. mysqlstore and sqlitestore use the *sql.DB you pass, keeping one of its connections for their own writes, so leave SetMaxOpenConns at 2 or more.

Polling. An idle server runs one indexed query per pool every PollInterval, and, when it gets no notifications (MySQL without a Bus), one more every 100ms to release delayed jobs on time, besides its heartbeat and the leader's periodic work. Each query is cheap, but they add up: an API process and two idle workers on MySQL ran about 70 queries per second with the default PollInterval and 200 with 200ms. A Bus removes the 100ms check and lets PollInterval stay long.

Rate limits. A key's Rate and Burst bound the times at which the store admits its jobs, that is, makes them claimable: no window of length Per admits more than Rate + Burst of them, even when admission falls behind and catches up. Admission reads the database clock when it runs, after any wait for a lock, so a slow admission does not bunch jobs up, and starts follow admissions by the claim latency. What admission cannot see is a stall between an admission and its commit, such as a slow disk flush or a paused database container: the jobs admitted just before the stall become claimable when it ends, together with those admitted right after it, and can start together, briefly above the rate. Checking the rate again at claim time would add a write to every claim of a rate-limited job, so kiln does not; leave the downstream some headroom over the configured rate.

Retries and alerts. A job that runs out of attempts stays in failed until someone requeues or deletes it, and the jobs waiting on it stay in awaiting until then. Give each kind a retry window longer than the longest outage you expect from what it calls: the window is the sum of the waits between attempts, and when those add up to fifteen seconds, a payment API that is down for four minutes fails every job that ran meanwhile. Registered with Exponential(2*time.Second, 2*time.Minute) and given MaxAttempts(12), a kind waits six to twelve minutes in all, since each wait is drawn between half and all of its value. Then alert on the jobs that fail anyway: kilnotel.Observe reports them as the kiln.jobs.failed gauge (kiln_jobs_failed in Prometheus), and the dashboard's Failed tab requeues them all at once, which also releases what waits on them.

Health. Server.Healthy() fails when the server has fenced itself off, its heartbeat is stale or its results are piling up; use it for readiness and liveness probes. Server.Stats() has the counters.