Operations
Run the published image against your bucket, watch a handful of metrics, and recover by starting new instances.
Run it
Images are published for linux/amd64 and linux/arm64. Pin the patch tag; there is no latest tag.
docker pull ghcr.io/burrr-ai/gitcask:0.0.7
docker run --rm -p 8080:8080 \
-e AWS_ACCESS_KEY_ID -e AWS_SECRET_ACCESS_KEY \
-v "$PWD/gitcask.toml:/etc/gitcask/gitcask.toml:ro" \
ghcr.io/burrr-ai/gitcask:0.0.7S3 keys are read from the environment at startup by default. Set credentials = "default" to use the refreshing AWS SDK chain: environment, profile, ECS task role, then instance metadata.
[store]
bucket = "your-bucket"
[store.s3]
credentials = "default"
region = "us-east-1"
endpoint = "" # use AWS S3 rather than the local rustfs default
force_path_style = false- serve
- Git, the JSON API and LFS.
- maintain
- Checkpoints, compaction and fsck, driven by pending markers.
- events
- The webhook bridge.
- roles = []
- All three. Any number of hosts can share one bucket.
First look
/healthz and /readyz never touch S3. A 200 does not mean the bucket is fine.
curl -si "$BASE/healthz"
curl -si "$BASE/readyz"
curl -s "$BASE/metrics"
gitcask --config gitcask.toml wal pending
gitcask --config gitcask.toml repo info owner/repo
curl -s "$BASE/owner/repo/api/overview"
curl -s "$BASE/owner/repo/api/tasks"- wal pending
- LISTs pending/. A manual check, never on a request path.
- repo info
- Runs a full sync and pulls packs local. Read api/overview first.
- api/tasks
- Per instance. Keep its hostname with the log request_id.
Metrics and alerts
Starting points; tune durations to your baseline. Sum counters across instances of a role.
| Metric | Normal | Alert on |
|---|---|---|
gitcask_push_refused_total{reason} | Flat outside deploys | One connectivity or unpack immediately; body only as a sustained rate; one draining outside a deploy window |
gitcask_store_requests_total{op,outcome} | ok, ordinary not_found / precondition_failed | One retryable_error or error immediately |
gitcask_store_retries_total{op} | Usually flat | Still climbing after 5 minutes |
gitcask_publish_local_apply_failed_total | 0 | Any increase |
gitcask_pending_marker_put_failures_total | 0 | Any increase; check that repository by hand |
gitcask_pending_markers | 0 when idle | > 0 for two whole maintenance intervals; page if pinned at the page limit |
gitcask_maintainer_heartbeat_timestamp{host} | Within two intervals of now | time() - value > max(2 * maintenance.interval, 5m) |
events_bridge_lag_entries{repo} | 0 after catch-up | > 0 for longer than events.sweep_interval |
gitcask_cache_disk_used_fraction | Below cache.disk_high_watermark | Above the watermark for ≥ 2 × cache.evict_interval |
gitcask_runtime_stall_total | Flat | Any increase; urgent when the log line shows inflight > 0 |
gitcask_repo_missing_objects{repo} | 0 | > 0 is a data-integrity incident, immediately |
From symptom to cause
| Symptom | Look at first |
|---|---|
| Clone or push is slow | Band-2 local copy is missing packs / local copy ready, then wal.materialize, wal.download_pack and the push's receive.body span |
| A push gets 503 | /readyz first: 503 is drain. Otherwise store retries and {"error":"store_unavailable","retryable":true} |
| A pushed ref is missing on another instance | Same bucket and prefix, wal.freshness_ttl = "0s". If both hold, it is a bug: keep the evidence |
| The disk is full | df -h on cache.dir, the watermark, and pack-using work in api/tasks |
| Webhooks are not arriving | events_bridge_lag_entries, then events_bridge_gap_total and events_bridge_sweep_found_total |
| A fresh instance is slow | One materialize task, then fast requests, is a normal cold cache |
Each symptom has a step-by-step check in docs/OPERATIONS.md.
Capacity
The default size is 2 vCPU. Scale by adding instances of the same role, not by raising worker counts.
| Signal | Action |
|---|---|
Store, disk and locks healthy; latency and gitcask_http_inflight high on several instances | Add serving instances |
gitcask_pending_markersdoes not converge and every maintainer's workers stay busy | Add maintainers |
Store retries and store.* elapsed_ms worsen as instances are added | Investigate the store; more serving multiplies S3 requests |
Deploys and recovery
Draining
SIGTERM drains in two phases. Route deploys on /readyz; /healthz stays 200 while the process lives.
| Phase | /readyz | What changes | Bound |
|---|---|---|---|
| 1. Maintenance drain | 200 | No new maintenance unit; the running unit is interrupted. Serving is normal | 30 s |
| 2. Serving drain | 503 + Retry-After: 15 | New fetch, push and LFS object work refused; in-flight requests finish | server.drain_timeout ("20s"), then 2 s more for the load balancer |
Incidents
| Incident | Do |
|---|---|
| S3 outage | Reduce write load and restore S3. Verify with a real refs read, then a small push read back from another instance |
| All instances lost | Start new instances with the same config and bucket: gitcask-server --config gitcask.toml |
| Bucket corruption or deletion | Stop writes and restore the original keys from S3 versions or replicas. Never hand-craft a manifest |
| A force push moved a ref | Do not roll the bucket back; push the old OID as a new ref update |
Rewind a ref
gitcask --config gitcask.toml wal ls owner/repo
gitcask --config gitcask.toml wal show owner/repo 42
gitcask --config gitcask.toml wal materialize owner/repo --at-seq 41 --out /tmp/repo-restore
git -C /tmp/repo-restore fsck --full
git -C /tmp/repo-restore push --force "$REPO_URL" \
<old_oid>:refs/heads/<branch>