gitcaskDocsGitHub

Operations

Run the published image against your bucket, watch a handful of metrics, and recover by starting new instances.

Run it

Images are published for linux/amd64 and linux/arm64. Pin the patch tag; there is no latest tag.

docker pull ghcr.io/burrr-ai/gitcask:0.0.7
docker run --rm -p 8080:8080 \
  -e AWS_ACCESS_KEY_ID -e AWS_SECRET_ACCESS_KEY \
  -v "$PWD/gitcask.toml:/etc/gitcask/gitcask.toml:ro" \
  ghcr.io/burrr-ai/gitcask:0.0.7

S3 keys are read from the environment at startup by default. Set credentials = "default" to use the refreshing AWS SDK chain: environment, profile, ECS task role, then instance metadata.

gitcask.toml
[store]
bucket = "your-bucket"

[store.s3]
credentials = "default"
region = "us-east-1"
endpoint = ""                 # use AWS S3 rather than the local rustfs default
force_path_style = false
serve
Git, the JSON API and LFS.
maintain
Checkpoints, compaction and fsck, driven by pending markers.
events
The webhook bridge.
roles = []
All three. Any number of hosts can share one bucket.

First look

/healthz and /readyz never touch S3. A 200 does not mean the bucket is fine.

curl -si "$BASE/healthz"
curl -si "$BASE/readyz"
curl -s "$BASE/metrics"

gitcask --config gitcask.toml wal pending
gitcask --config gitcask.toml repo info owner/repo
curl -s "$BASE/owner/repo/api/overview"
curl -s "$BASE/owner/repo/api/tasks"
wal pending
LISTs pending/. A manual check, never on a request path.
repo info
Runs a full sync and pulls packs local. Read api/overview first.
api/tasks
Per instance. Keep its hostname with the log request_id.

Metrics and alerts

Starting points; tune durations to your baseline. Sum counters across instances of a role.

MetricNormalAlert on
gitcask_push_refused_total{reason}Flat outside deploysOne connectivity or unpack immediately; body only as a sustained rate; one draining outside a deploy window
gitcask_store_requests_total{op,outcome}ok, ordinary not_found / precondition_failedOne retryable_error or error immediately
gitcask_store_retries_total{op}Usually flatStill climbing after 5 minutes
gitcask_publish_local_apply_failed_total0Any increase
gitcask_pending_marker_put_failures_total0Any increase; check that repository by hand
gitcask_pending_markers0 when idle> 0 for two whole maintenance intervals; page if pinned at the page limit
gitcask_maintainer_heartbeat_timestamp{host}Within two intervals of nowtime() - value > max(2 * maintenance.interval, 5m)
events_bridge_lag_entries{repo}0 after catch-up> 0 for longer than events.sweep_interval
gitcask_cache_disk_used_fractionBelow cache.disk_high_watermarkAbove the watermark for ≥ 2 × cache.evict_interval
gitcask_runtime_stall_totalFlatAny increase; urgent when the log line shows inflight > 0
gitcask_repo_missing_objects{repo}0> 0 is a data-integrity incident, immediately

From symptom to cause

SymptomLook at first
Clone or push is slowBand-2 local copy is missing packs / local copy ready, then wal.materialize, wal.download_pack and the push's receive.body span
A push gets 503/readyz first: 503 is drain. Otherwise store retries and {"error":"store_unavailable","retryable":true}
A pushed ref is missing on another instanceSame bucket and prefix, wal.freshness_ttl = "0s". If both hold, it is a bug: keep the evidence
The disk is fulldf -h on cache.dir, the watermark, and pack-using work in api/tasks
Webhooks are not arrivingevents_bridge_lag_entries, then events_bridge_gap_total and events_bridge_sweep_found_total
A fresh instance is slowOne materialize task, then fast requests, is a normal cold cache

Each symptom has a step-by-step check in docs/OPERATIONS.md.

Capacity

The default size is 2 vCPU. Scale by adding instances of the same role, not by raising worker counts.

SignalAction
Store, disk and locks healthy; latency and gitcask_http_inflight high on several instancesAdd serving instances
gitcask_pending_markersdoes not converge and every maintainer's workers stay busyAdd maintainers
Store retries and store.* elapsed_ms worsen as instances are addedInvestigate the store; more serving multiplies S3 requests

Deploys and recovery

Draining

SIGTERM drains in two phases. Route deploys on /readyz; /healthz stays 200 while the process lives.

Phase/readyzWhat changesBound
1. Maintenance drain200No new maintenance unit; the running unit is interrupted. Serving is normal30 s
2. Serving drain503 + Retry-After: 15New fetch, push and LFS object work refused; in-flight requests finishserver.drain_timeout ("20s"), then 2 s more for the load balancer

Incidents

IncidentDo
S3 outageReduce write load and restore S3. Verify with a real refs read, then a small push read back from another instance
All instances lostStart new instances with the same config and bucket: gitcask-server --config gitcask.toml
Bucket corruption or deletionStop writes and restore the original keys from S3 versions or replicas. Never hand-craft a manifest
A force push moved a refDo not roll the bucket back; push the old OID as a new ref update

Rewind a ref

gitcask --config gitcask.toml wal ls owner/repo
gitcask --config gitcask.toml wal show owner/repo 42
gitcask --config gitcask.toml wal materialize owner/repo --at-seq 41 --out /tmp/repo-restore
git -C /tmp/repo-restore fsck --full
git -C /tmp/repo-restore push --force "$REPO_URL" \
  <old_oid>:refs/heads/<branch>