Operating kiln¶
ServerConfig |
Default | What it controls |
|---|---|---|
PollInterval |
1s | How often an idle pool looks for work when no notification arrives |
HeartbeatInterval |
5s | How often a server reports that it is alive |
DeadAfter |
60s | Silence after which a server's running jobs are retried elsewhere |
LeaderTTL |
15s | Lease of the leader that runs recurring jobs, rescue, sweep and pruning |
ShutdownTimeout |
30s | How long Run waits for running jobs after its context is cancelled |
KillGrace |
5s | Extra wait after cancelling jobs that ignored the shutdown |
Timeout |
30m | Default per-job timeout |
Recovering from a dead worker. A server that exits cleanly finishes its running jobs, or hands them
back, before Run returns. A server that dies is noticed after DeadAfter, and within one more
HeartbeatInterval its jobs are rescued: they go straight back to their queues, and the rescued attempt
counts toward MaxAttempts. With the defaults, a job whose worker was killed starts again after roughly
60 to 70 seconds. A job rescued twice in a row waits for its backoff before the next attempt, so a job that
takes down the process running it is not handed from worker to worker without a pause. If the dead server
was also the leader, a new leader takes over after LeaderTTL and waits DeadAfter + HeartbeatInterval
before rescuing anything, which adds about 20 seconds. Lower DeadAfter to notice dead workers sooner (it
must stay above 3*HeartbeatInterval + KillGrace + 5s), and keep long jobs resumable with SetParam
checkpoints.
Deploys. On SIGTERM a server stops claiming and waits up to ShutdownTimeout for its running jobs.
A job still running after that is cancelled with ErrShutdown and goes back to its queue without using
up an attempt, so another server runs it again from the start (or from its SetParam checkpoints). Set ShutdownTimeout above your longest job, and give your orchestrator a
grace period longer than ShutdownTimeout + KillGrace (Kubernetes' terminationGracePeriodSeconds,
Docker's stop_grace_period), or the process is killed first and its jobs wait DeadAfter to be
rescued.
One queue per service. A server claims only the kinds it has handlers for, so two services can share a database and even a queue without running each other's jobs. Sharing the queue has a cost, though: the claim query walks past the other service's ready jobs to reach its own kinds, about 34ms per 100,000 of them on PostgreSQL. Give each service its own queues, and the walk disappears.
Connections. pgstore opens its own pool (MaxConns, 8 by default) plus one connection for
LISTEN, on top of your application's pool: budget up to 9 connections per process that opens a store.
mysqlstore and sqlitestore use the *sql.DB you pass, keeping one of its connections for their own
writes, so leave SetMaxOpenConns at 2 or more.
Polling. An idle server runs one indexed query per pool every PollInterval, and, when it gets no
notifications (MySQL without a Bus), one more every 100ms to release delayed jobs on time, besides its
heartbeat and the leader's periodic work. Each query is cheap, but they add up: an API process and two
idle workers on MySQL ran about 70 queries per second with the default PollInterval and 200 with
200ms. A Bus removes the 100ms check and lets PollInterval stay long.
Rate limits. A key's Rate and Burst bound the times at which the store admits its jobs, that
is, makes them claimable: no window of length Per admits more than Rate + Burst of them, even when
admission falls behind and catches up. Admission reads the database clock when it runs, after any wait
for a lock, so a slow admission does not bunch jobs up, and starts follow admissions by the claim
latency. What admission cannot see is a stall between an admission and its commit, such as a slow disk
flush or a paused database container: the jobs admitted just before the stall become claimable when it
ends, together with those admitted right after it, and can start together, briefly above the rate.
Checking the rate again at claim time would add a write to every claim of a rate-limited job, so kiln
does not; leave the downstream some headroom over the configured rate.
Retries and alerts. A job that runs out of attempts stays in failed until someone requeues or
deletes it, and the jobs waiting on it stay in awaiting until then. Give each kind a retry window
longer than the longest outage you expect from what it calls: the window is the sum of the waits
between attempts, and when those add up to fifteen seconds, a payment API that is down for four
minutes fails every job that ran meanwhile. Registered with Exponential(2*time.Second,
2*time.Minute) and given MaxAttempts(12), a kind waits six to twelve minutes in all, since each
wait is drawn between half and all of its value. Then alert on the jobs that fail anyway:
kilnotel.Observe reports them as the kiln.jobs.failed gauge (kiln_jobs_failed in Prometheus),
and the dashboard's Failed tab requeues them all at once, which also releases what waits on them.
Health. Server.Healthy() fails when the server has fenced itself off, its heartbeat is stale or its
results are piling up; use it for readiness and liveness probes. Server.Stats() has the counters.