# gitcask > An open-source (Apache-2.0), stateless git server that keeps every repository as a write-ahead log in S3-compatible object storage. No database and no leader; any instance serves any repository, and cost scales with pushes rather than with the number of repositories. English | [한국어](https://github.com/burrr-ai/gitcask/blob/main/README.ko.md) # gitcask gitcask is a stateless git server that uses S3-compatible object storage as its only persistent layer. Each repository is stored in the bucket as a write-ahead log; servers hold only caches. There is no database and no leader election, and operating cost is proportional to pushes rather than to the number of repositories, so idle repositories cost nothing to keep. It is intended for platforms that create and delete repositories programmatically, one per user project. - **git over HTTP** — clone, fetch, push and LFS work with standard git clients. - **JSON API** — read trees, commits and diffs; commit files, create branches and merge, all without a clone or working directory. - **Independent initialization** — [copy a pinned Gitcask tree into a pristine repository](/llms/initialize.md) as one root commit, with independent objects and safe retries. - **Webhooks** — each ref change is delivered once, from a durable cursor, and can be replayed. - **Stateless servers** — every instance can serve every repository. A new instance starts serving refs within a few seconds. - **Built-in authentication** — gitcask verifies platform JWTs with a public key or introspects opaque tokens with their issuer, then applies repository scopes. No user database is involved. ## Where it came from Vicent Martí described this architecture in Cursor's [*Git at any scale*](https://cursor.com/blog/git-at-any-scale) — Cursor calls it Continuity and runs a large monorepo on it. Tobi Lütke reproduced it as [walgit](https://github.com/tobi/walgit), and gitcask is a fork of walgit. Cursor noted that the design "scales in both directions": one huge repository, or a vast number of small ones. gitcask is built for the latter: many small repositories, created and deleted by a platform. What changed relative to walgit, and why, is recorded in [docs/DIRECTION.md](/llms/direction.md). gitcask exists because Cursor published the design and Tobi published the code. Thank you both. ## Try it in five minutes ```sh docker compose up --build --wait TOKEN=$(docker compose run --rm token --config /etc/gitcask/gitcask.standalone.toml token mint \ --key /run/secrets/gitcask-private.pem --principal local-demo --scope local/demo:admin --ttl 1h) curl -fsS -X PUT -H "Authorization: Bearer $TOKEN" http://127.0.0.1:8080/local/demo git clone "http://ignored:$TOKEN@127.0.0.1:8080/local/demo.git" cd demo && git commit --allow-empty -m first && git push -u origin HEAD:main ``` This starts a local S3-compatible store (rustfs) and one gitcask process. The token is a demo JWT signed with a throwaway key. In production the platform signs tokens with its own Ed25519 key and gitcask holds only the public key. Git sends the token in the Basic-auth password field; the username is ignored. Note on token lifetimes: a token a person pastes into git is saved by the OS credential helper and reused. If it has to live long — weeks rather than hours — narrow its scope accordingly, to one repository with minimal permission. Backend tokens minted per request can expire in minutes. ## How it works A repository lives under `repos///` in the bucket as a write-ahead log. A push uploads an immutable pack and a log entry, then updates a small manifest with a compare-and-swap. The CAS is the only point of consensus; no election or quorum is involved. When two instances race, one succeeds and the other retries. A read begins with a conditional GET of the manifest. In the common case the store answers 304 and the local copy is used; otherwise the server applies the new entries before responding. A push that has been acknowledged on one instance is therefore immediately visible on all of them. If every server disappears, the data remains in the bucket. A new instance pointed at it serves refs within a few seconds and downloads packs on the first object request. Maintenance work — checkpoints, compaction, integrity audits, garbage collection — is driven by markers written on push, so repositories without recent pushes are not visited at all. The full design, including the reasoning behind each decision and the round-trip cost model, is documented in [AGENTS.md](/llms/architecture.md). ## Scope Some things are intentionally left to the calling platform: - **Repository listing and search.** Listing would require bucket scans, which this design avoids; the platform's database is expected to hold the authoritative list. - **Users and login.** gitcask receives an opaque principal string and verifies a signature. Identity, permissions and teams remain in the platform, so there is no account data to synchronise. - **CI, issues, UI.** The webhook can drive any CI system and the API can back any UI. The boundary and its rationale are described in [docs/PRODUCT.md](/llms/product.md). Data can always be taken out: `git clone --mirror` exports a repository (with `git lfs fetch --all` for LFS content), and a consistent copy of the bucket — taken while writes are paused, or via S3 versioning or point-in-time replication — is a complete backup: a fresh deployment pointed at it serves as-is. ## Running it Published images are available at `ghcr.io/burrr-ai/gitcask:0.0.7` for `linux/amd64` and `linux/arm64`. Pin the patch tag or the image digest recorded in the [release](https://github.com/burrr-ai/gitcask/releases); there is no floating `latest` tag. ```sh docker pull ghcr.io/burrr-ai/gitcask:0.0.7 docker run --rm -p 8080:8080 \ -e AWS_ACCESS_KEY_ID -e AWS_SECRET_ACCESS_KEY \ -v "$PWD/gitcask.toml:/etc/gitcask/gitcask.toml:ro" \ ghcr.io/burrr-ai/gitcask:0.0.7 ``` Provide a production config with your S3 bucket/region/endpoint and JWT or introspection authentication. The cache at `/var/lib/gitcask` is disposable. The image runs as UID 1000; mounted cache directories must be writable by that user. By default, S3 keys are read from environment variables at startup. Set `store.s3.credentials = "default"` to use the AWS SDK credential chain (environment, profile, ECS task role, then instance metadata) with automatic refresh. On ECS, the SDK picks up the task role automatically; attach one and use: ```toml [store] bucket = "your-bucket" [store.s3] credentials = "default" region = "us-east-1" endpoint = "" # use AWS S3 rather than the local rustfs default force_path_style = false ``` `access_key_env` and `secret_key_env` are ignored in this mode. Presigned URLs, including those used for `accel_redirect`, work only when the SDK can obtain signing credentials. Platforms with opaque tokens use `server.auth_mode = "introspect"` and configure their introspection endpoint in [`gitcask.example.toml`](/llms/config.md). To accept direct tokens and trusted proxy requests on the same listener, opt into `introspect_forwarded`; see [precedence and proxy setup](/llms/security.md#mixed-direct-and-trusted-proxy-authentication). Browse the API reference at `GET /docs` or fetch its OpenAPI schema at `GET /openapi.json`. Both are public by default; set `[server] public_docs = false` to require authentication. Existing `/api/v1/docs` and `/api/v1/openapi.json` links permanently redirect to the new paths. `/metrics` always requires authentication. The exact top-level docs paths leave owners named `docs` and `openapi.json` available under `/{owner}/{repo}`. ```sh # build (needs the Rust version in rust-toolchain.toml, plus protoc) cargo build --release -p gitcask-cli # or: docker build -t gitcask . # one machine: gitcask with authentication on :8080, backed by local rustfs docker compose up --build -d curl http://127.0.0.1:8080/healthz ``` - [`gitcask.standalone.toml`](https://github.com/burrr-ai/gitcask/blob/main/gitcask.standalone.toml) — the single-process configuration. A good starting point. - [`gitcask.example.toml`](/llms/config.md) — every configuration key, with defaults and comments. - [`deploy/nginx.conf.example`](https://github.com/burrr-ai/gitcask/blob/main/deploy/nginx.conf.example) — an optional nginx in front, for public TLS and byte offload. Roles (`server.roles`): `serve` (git, API, LFS), `maintain` (checkpoints, compaction, fsck), `events` (the webhook bridge). An empty list enables all three. Any number of hosts can share one bucket. ## Developing ```sh just test # fast hermetic tier (< 1 min) just e2e # real git against a running server (~20 s) just ci # what a merge requires: warnings, clippy, test, e2e scripts/smoke.sh . 8090 # end-to-end against local rustfs pnpm run docs # the documentation site (site/) on localhost:3000 cargo test -p gitcask-server --test sim # fault injection: crashes, partitions, stale reads ``` ``` crates/ gitcask-proto protobuf schema (wal.proto), log framing, store keys gitcask-store the ObjectStore trait; S3 backend, leases, retries gitcask-git bare repositories on disk, receive-pack, pack ingest, refs gitcask-wal sync levels, publish (group commit + CAS), checkpoints, tasks gitcask-server axum: smart HTTP, LFS, auth, JSON API, maintainer, events bridge gitcask-config gitcask.toml (+ GITCASK__ env overrides), fail-closed validation gitcask-cli gitcask serve|import|migrate|compact|wal|repo|token ``` Further documentation: [docs/IMPORT.md](/llms/import.md) (pinned full-history import), [docs/INITIALIZE.md](/llms/initialize.md) (parentless tree snapshots), [docs/PRODUCT.md](/llms/product.md) (product boundaries), [docs/DIRECTION.md](/llms/direction.md) (this fork's decisions), [docs/OPERATIONS.md](/llms/operations.md) (the runbook), [docs/MIGRATION.md](/llms/migration.md) (moving from Gitea), [docs/ROUNDTRIPS.md](/llms/roundtrips.md) (the cost model), [docs/EVENTS.md](/llms/events.md), [docs/INTEGRITY.md](/llms/integrity.md), [docs/LFS.md](/llms/lfs.md). ## Contributing Development, tests and DCO requirements are in [CONTRIBUTING.md](/llms/contributing.md). Please report vulnerabilities privately as described in [SECURITY.md](/llms/security.md). Participation is governed by the [Contributor Covenant](https://github.com/burrr-ai/gitcask/blob/main/CODE_OF_CONDUCT.md). # GOAL — what gitcask is for Context: for everyone (humans and agents) working on this repository. Read this before `AGENTS.md`. This page is **what we want**; `AGENTS.md` is **how we build it**; `docs/DIRECTION.md` records the decisions that shaped this fork. When a choice is not obviously right, come back here and ask: which option serves this goal better? ## The one sentence **A share-nothing git backend for platforms with a very large number of small repositories — one per user project — created and deleted programmatically, with an object store as the *only* source of truth and disposable serving instances.** comwit is gitcask's first user and the workload that shaped it; the architecture is not tied to comwit. ## What that means, unpacked 1. **The object store is the only source of truth.** The bucket *is* the repository. Every push is an immutable object + one CAS'd manifest write; every instance is a disposable cache that revalidates with one conditional GET. No database inside gitcask, no leader, no node identity. Wipe every instance and lose nothing but warmth. (Cursor's *Git at any scale* / Continuity is the design we follow — `docs/reference/cursor-git-at-any-scale.md`.) 2. **Share-nothing, elastic.** Any host pointed at the bucket can serve any repository, including its object work. Coordination is only through object-store primitives (CAS, leases, content-addressed immutable objects). Consistency is never "eventual": push acknowledged ⇒ the next request anywhere sees it. 3. **Many small repositories, not one big one.** The workload is tens of thousands of users × ~20 projects each, a few MB to a few hundred MB per repository, pushed frequently by an agent during a session and idle for days after. Every cost must scale with *pushes*, never with the *number of repositories*: no pass over all repositories, no listing of the bucket on any path that runs periodically. Packs always fit on the instance; there is no remote-pack path. 4. **gitcask knows nothing about users.** Identity, permissions, the list of repositories and their metadata live in the calling platform's database. gitcask verifies EdDSA JWTs with a public key/JWKS or introspects opaque tokens with their issuer, then applies the same repository scopes; it never stores users, sessions, revocations, or a signing key. It exposes create/delete for the platform to call. There is no login, no token store, no repository listing. Direct tokens and a trusted proxy can share one process with explicit `introspect_forwarded` mode ([security contract](/llms/security.md#mixed-direct-and-trusted-proxy-authentication)). 5. **All the features a git host needs, and only those**: smart HTTP v0/v2 (ls-refs, fetch with filter/shallow/deepen, receive-pack atomic/delete/tags/push-options/report-status-v2), LFS, `/` namespaces, ref events to a webhook, a JSON API for browsing and deterministic Git writes, tasks/narration so nothing ever waits silently. Not in scope: code review, merge queues, CI, issues, branch protection, per-repository policy — those live in the product built on gitcask. 6. **Predictable for the systems that build on it**: stable, immutable, cacheable artefacts (packs, sha-addressed API answers); O(1) ref lookups; latency that does not depend on which instance you hit; a provenance log you can rewind to any push (`gitcask wal materialize --at-seq`) — the raw material for "restore my project to yesterday". 7. **Use the tools; don't reinvent them.** Upstream `git` for upload-pack, index-pack and repack; Rust + tokio + axum for the server; the object store as it is (conditional writes, range reads, multipart); a plain proxy in front. Anything git can do, git does; gitcask decides only *where the bytes live*. ## How we know we are there (acceptance) | Claim | Bar | |---|---| | Cold instance is useful in seconds | `ls-remote` of any repository < 1 s on a fresh instance | | Fresh clone / fetch / push | ordinary smart HTTP works with the client's `git`; `scripts/smoke.sh` passes end to end against rustfs | | Cache is a cache | stop the server, delete `cache.dir`, start it: every repository clones again from the bucket | | Push | acknowledged only after the bucket ACKs; one CAS per batch; 5 store requests on the happy path | | Consistency | push then fetch anywhere sees it; concurrent pushers: exactly one winner (the simulation suite) | | Cost model | a maintainer pass touches only repositories with a pending marker; no periodic LIST of `repos/` anywhere | | Security | in `jwt`/`introspect` modes every route but `/healthz`, `/readyz`, and the static API docs (when `server.public_docs = true`) requires a valid token and repo routes require scope; introspection-service failures are 503, never grants; `introspect_forwarded` selects one verified identity per request (AGENTS D48); `forwarded` remains available; `none` binds loopback only | | Transient store errors | 5xx / throttling on any store operation is retried with backoff and never surfaces as a failed push on its own | | Data completeness | every object reachable from an advertised ref is in the pack set; `fsck` reports violations | ## What we deliberately do **not** optimise for - A single repository larger than the instance's disk. Packs are always fully local; if a repository outgrows the host, the answer is a bigger host or a size limit in the calling platform, not a remote-pack path. - Human-facing hosting features (web UI, sign-in, branch protection, bundle-uri for CI clones). The calling platform owns the product; gitcask owns the bytes. - Forking git or inventing an object format: weird stuff happens *around* git, never inside it. # PRODUCT — what we build and what we refuse to This document fixes the **boundaries**: what belongs in the gitcask core versus the calling platform, and what is open source versus cloud-only. Every feature request is judged by the two rules in §2. When the rules do not produce an answer, the rules are incomplete — make the decision and write it down here. Boundaries kept in people's heads drift. | Document | The question it answers | |---|---| | `GOAL.md` | What we aim for technically | | `docs/DIRECTION.md` | What this fork changed relative to walgit; operating decisions D1–D14 | | **this document** | How far the product goes | ## 1. Position **A headless git backend.** The storage, transport, read and commit layer of a git platform, used only through its API. Offered both self-hosted (open source) and as a cloud service. It has no UI, no issues, no pull requests, no CI, and no identity. comwit is to gitcask what an application is to Postgres. Supabase turned Postgres into an API; gitcask does that for git. ### Why this seat is empty An AI coding platform has to store *N repositories per user × tens of thousands of users*. Today it has three answers, all bad: | Alternative | The problem | |---|---| | Delegate to users' GitHub accounts (OAuth) | Users must have a GitHub account; ownership and rate limits sit in someone else's hands | | Self-host Gitea/GitLab | Cost scales with repository count — exactly what broke for comwit | | Buy hosted headless git (code.storage, used by Lovable and Bolt as of 2026-09) | Proprietary, hosted only; hot storage priced at $0.005/GB/hour (≈ $3.6/GB/month) plus egress, because its architecture keeps replicated repositories on server disks and uses object storage only as a cold tier. User code lives on someone else's infrastructure | | Build it yourself | Collapses the moment history, branches or LFS show up | gitcask has one weapon: **cost scales with pushes, not with the number of repositories** (D40's pending markers; confirmed in a 50k-repository spike — 118 vs 120 store requests over five idle minutes). The market where that weapon matters is platform operators with exploding repository counts, and they already have their own UI and issue tracker. So we sell storage, transport and reads — nothing above them. For the same reason we do not aim to *replace Gitea*. A hundred-user team is fine on Gitea, and our weapon means nothing there. ## 2. The two judging rules ### Rule A — core or platform? > **Same repository state + same request → always the same result: core. > The answer differs per organisation: platform.** | Operation | What the answer depends on | | |---|---|---| | Commit a file | parent commit + new content | core | | Create/delete a branch or tag | ref name + oid | core | | Merge two refs | two trees + their merge base | core | | Archive | one commit | core | | "May this PR be merged?" | approvals, CI status, branch protection — **differs per organisation** | platform | | Issues, reviews, notifications | data that does not live in the repository | platform | | "May this user see this repository?" | the organisation's permission model | platform | GitHub's own API splits along this line: `POST /repos/{o}/{r}/merges` (a plain git merge, no PR involved) is the core side; `PUT /repos/{o}/{r}/pulls/{n}/merge` (approvals, checks, policy) is the platform side. **Secondary test**: can it be done with git plumbing alone, no working directory? `update-ref`, `hash-object` + `commit-tree`, `merge-tree` — if those suffice, it is core. ### Rule B — open source or cloud? > **Everything needed to run your own repositories on your own infrastructure: OSS. > Everything needed to run someone else's repositories for them: cloud.** | | OSS core | Cloud | |---|---|---| | git transport (clone/push/LFS/v2) | ● | | | Read API (refs/resolve/tree/blob/commits/commit/compare) | ● | | | Write API (branches/tags, archive, batch commits, merge) | ● | | | Repository create/delete | ● | | | JWT verification, token introspection, trusted forwarding, repository scopes, offline token CLI (§6) | ● | | | Event webhooks, metrics, `/healthz`, size limits | ● | | | Migration tooling | ● | | | Multi-tenancy (tenant isolation, bucket placement) | | ● | | Metering, billing, tenant quotas | | ● | | Web console, dashboards | | ● | | Autoscaling, multi-region, managed backups, SLA, on-call | | ● | The boundary is also a **repository boundary**: `gitcask` (public) / `gitcask-cloud` (private). Cloud code never lives in this repository — GitLab's `ee/` layout is the cautionary tale: closed logic leaks into the core and the split becomes impossible later. ## 3. The litmus test > **comwit must run on the OSS core alone.** If comwit runs in production without any cloud feature, the OSS is real and adoption can happen. The moment "we'd need a cloud feature to run comwit" is uttered, **the line was drawn wrong** — move that feature down into the OSS. Being our own biggest user keeps the OSS honest automatically. ## 4. Out of scope — with the reason attached A bare "we don't do that" erodes within six months. Each refusal carries its reason. | Item | Why not | |---|---| | **CI / action runners** | Job queues, runner registration, run history, secrets — all state outside the repository, tens of writes per second. S3 CAS (one write per second per object) cannot carry that, so a database walks in, and at that moment gitcask is Gitea. **The event webhook already connects any CI system** — that is our answer. | | **Issues, PRs, reviews, wikis, releases** | Data that does not live in the repository (Rule A). | | **Users, orgs, teams, login, SSO, SCIM** | We own no identity (§6). Owning none, the grey zone does not exist. | | **Repository listing and search** | It would need LISTs — the exact bottleneck this architecture removed (D40). The caller's database already has the authoritative list. **This surprises OSS adopters, so the README explains it explicitly.** | | **A web UI** | The platform owns the product; gitcask owns the bytes (`GOAL.md`). | | **PITR / backup tooling / export** | Already covered by what exists: a force-pushed-away ref comes back via `wal ls`'s `old_oid` + `git push -f`; a backup is a bucket copy; an export is `git clone --mirror`. **The answer is documentation, not code.** | ## 5. Layers and how they separate ``` OSS ──── gitcask the git engine — auth · transport · WAL · read/write API Closed ─ gitcask-cloud multi-tenancy · metering · billing · operations ``` The line along which things *can* separate is **"does it read packs (git objects)?"** | Level | How | Why | |---|---|---| | **Process split** | none | Credential verification and the handlers' permission checks live in one process, with one path-to-permission table. Optional trusted forwarding follows the explicit security boundary in §6. | | **Logical split** | roles (`serve` / `maintain` / `events`) | Want pure storage? Point only `serve` hosts at the bucket. The code boundary is `crates/gitcask-server/src/web/`. | | **Physical split** | static bytes go to the edge | Raw blobs and archives are immutable and servable without packs — `X-Accel-Redirect` (D23). The real split point for a SaaS whose cost is bandwidth. | Note the intuition that "clone/push is light, the API is heavy" is **backwards**. `upload-pack` walks the object graph to build a fresh pack and `receive-pack` indexes and connectivity-checks, so both need every pack local (D41). `refs`/`resolve` touch zero objects — they are already fully diskless via `SyncLevel::Refs`. ## 6. We own no identity **The only thing we store about a caller is one opaque `principal` string.** We do not know who they are, their email, or whether they still exist. With no user table there is **nothing to synchronise and therefore nothing to conflict with the caller's own authentication.** Gitea demands a Gitea user before a repository can exist, and at that moment the caller enters user-sync hell: sign-ups, deletions, suspensions and renames must be mirrored in two places, and every mismatch orphans a repository. **"A git backend with no user synchronisation" is the differentiator.** Git clients speak HTTPS Basic, so the EdDSA JWT rides in the password slot (the username is ignored); the API takes the same token as a Bearer header. The platform signs `sub` (an opaque principal), `scopes`, `exp`, `iat` and `jti` with its own Ed25519 private key; gitcask verifies with only the public key or a cached JWKS from `[auth.jwt]`. There is no HS256, no issuance endpoint, no callback, no static token list. Self-hosters and CI use the offline `gitcask token keygen|mint` commands. JWT verification holds only public keys; revocation uses short expiry and issuer key rotation. Introspection keeps identity outside gitcask too: the platform owns its opaque tokens and can revoke them instantly, with gitcask observing revocation when its bounded answer cache expires (or on every request with zero positive TTL). gitcask has no user DB, no sessions, no revocation list, no usage history. Deployments with a trusted IdP proxy can choose `forwarded`, or explicitly enable `introspect_forwarded` to accept direct opaque tokens and proxy requests on one listener. This remains credential verification in the OSS core, with one identity per request and no platform policy engine. See the [precedence and proxy trust contract](/llms/security.md#mixed-direct-and-trusted-proxy-authentication). ## 7. License **Apache-2.0**, switched from MIT on 2026-09-01, before publication. Both are permissive; Apache adds an explicit patent grant and patent-retaliation termination, which is what enterprise legal teams look for in infrastructure software — Kubernetes and Neon's storage engine ship the same way. gitcask is a fork of walgit (MIT); the original copyright and permission notice are preserved in `NOTICE`, as the MIT terms require. Relicensing a permissive-licensed fork this way is legal and routine; it is copyleft (GPL/AGPL) code that a fork cannot relicense. AGPL or BUSL would fence off cloud vendors, but early on the risk of not being adopted dwarfs the risk of being copied. And the engine alone cannot be sold as a SaaS — multi-tenancy, provisioning, billing and operations all live on the closed side, which is a natural moat regardless of the license. ## 8. Current gaps (2026-09-20) | Item | Status | |---|---| | Upstream correctness fixes since the fork (empty-pack tip verification; warnings gate blind under forced colour) | done (T37, PR #2). Upstream is reviewed monthly for such fixes; they are re-implemented, not cherry-picked | | Opaque-token authentication (`auth_mode = "introspect"`, RFC 7662) for platforms without a JWT signer | ✅ merged (PR #4) | | Ref-level push restrictions (protected branches, fast-forward only) | open — needed before agents get write tokens; the platform cannot enforce this after the fact because the push has already landed | | git transport, read API, repository CRUD, event webhooks, metrics, size limits | done | | Write API — branch/tag CRUD, archive | done (T28): reuses the WAL publish path, `expected_old_oid` CAS, immutable archives | | Write API — batch file commits, merge | done (T32): one request = one commit = one pack; conflicts are 409 + the conflicting paths | | Authentication | done (T34): EdDSA JWT verified in-process, repository scopes, Basic/Bearer, public key/JWKS, offline token CLI | | Migration (bulk Gitea → gitcask) + runbook | done (T27): `gitcask migrate gitea`, resumable, LFS included, `docs/MIGRATION.md` | | Publication readiness (de-comwit framing, one-command compose, SECURITY/DCO) | done (T33); the human checklist at the end of `README.md` remains | **Parked** — revisit when the SaaS starts: - **R2 validation**: Cloudflare R2 has free egress, which decides whether a bandwidth-priced business works at all. It is S3-compatible, so this is likely an endpoint question rather than a new backend — the one thing to verify against a real account is that conditional PUT (our CAS) behaves. Local rustfs proves nothing here. - **The cloud control plane**: multi-tenancy, metering, billing, console. It observes traffic and storage and **puts prices on them — gitcask itself never knows about billing.** # gitcask — architecture and operating manual for contributors (humans and agents) Context: **everyone, first.** Read this before touching the repository: the constraints the design answers (§1), how the WAL works (§2), the principles a PR is judged against (§3), every design decision in force (§4), and the working rules (§5). Reading order starts at `GOAL.md`; §0 maps every other document. > **No backwards compatibility (pre-1.0).** We owe nothing to previous shapes of this system. Do not keep > aliases, fallbacks, shims, deprecated routes, old config keys, old proto fields or migration code. When a > decision changes the shape, **delete the old shape in the same change**: routes, clients, tests, docs, config > keys. Data in the bucket is the one exception — WAL/manifest/log formats stay append-only and replayable, > because the bucket is the repository; everything else is disposable. If you find yourself writing "still > accepted for …" or "legacy", stop and remove the thing instead. gitcask serves git over smart HTTP (v0/v2), receive-pack, upload-pack, LFS and a JSON API, written in Rust, from disposable hosts whose only durable state is an object-store bucket. The design follows Cursor's *Git at any scale* (Continuity) as implemented by walgit, reshaped for platforms with many small repositories and identity and metadata outside gitcask; comwit is the first such user. `README.md` tells the story, `GOAL.md` the target, `docs/DIRECTION.md` what this fork changed and the operating decisions (D1–D14 there are *operating* decisions, distinct from the design decisions numbered in §4 here); this file keeps the rules. ## 0. Document map (one home per fact — link, don't duplicate) | Doc | Who / when to read it | |---|---| | `GOAL.md` | Everyone, first. What gitcask is for, the acceptance table, what we do not optimise for. | | `docs/DIRECTION.md` | Everyone. Why this fork differs from walgit, what was removed and why, the operating decisions (D1–D14). | | `docs/PRODUCT.md` | Everyone. The product boundary: core vs platform, OSS vs cloud, the two judging rules, and what we will not build — with the reason for each. | | `scripts/smoke.sh`, `scripts/clippy-count.sh` | The end-to-end smoke test against rustfs and the deterministic clippy counter. | | `AGENTS.md` (this) | Everyone. Constraints §1, WAL §2, principles §3, decisions §4, working rules §5. | | `README.md` | The introduction: why (the Cursor lineage), what it does, how it works briefly, running it, invariants. | | `docs/ROUNDTRIPS.md` | **Anyone touching a protocol that talks to the bucket** (publish, sync, checkpoints, compaction/leases, store backends). Round trips are the cost model; correct is not sufficient. | | `docs/LFS.md` | Anyone touching LFS (`lfs.rs`) or importing a repository whose LFS history lives elsewhere. | | `docs/INTEGRITY.md` | Anyone touching import, the maintainer's `fsck` unit, or seeing `connectivity: missing object` on a push. | | `docs/IMPORT.md` | Full-history pristine import: pinning, receipts, HTTPS acquisition and preservation boundaries. | | `docs/INITIALIZE.md` | Pristine repository initialization: pinned trees, independent objects, authorization, retries and LFS/submodule boundaries. | | `docs/EVENTS.md` | Anyone changing WAL-derived ref events, the webhook bridge, consumer semantics or event cursors. | | `docs/OPERATIONS.md` | Operators and on-call responders. Metrics, symptom-first diagnosis, capacity, incidents, recovery and routine checks. | | `docs/CONTRACT.md` | When you touch a crate boundary. The cross-crate contract; *extend, don't rename*; code wins where they differ. | | `docs/reference/cursor-git-at-any-scale.md` | Pointer + key excerpts of the source design (the full post is Cursor's copyright). Read the post once before touching WAL/publish/sync. | | `gitcask.example.toml` | Every config key with its default and a comment. Change it with the code. | | `gitcask.standalone.toml` | The one-machine shape: one JWT-verifying `gitcask-server` on :8080 → rustfs. | | `deploy/nginx.conf.example` | Optional public TLS and `X-Accel-Redirect` byte offload in front of gitcask. | | `Dockerfile` | An OCI image. | | `site/` | The website (D54): the value landing, hand-written docs pages, and the Markdown above published verbatim at `/llms.txt`. `pnpm run docs` at the root serves it locally; read `site/design.md` before changing a screen. | --- ## 1. Constraints we design for ### 1.1 The machine: ephemeral, shared-nothing - **Instances are ephemeral and shared-nothing**: no stable identity, no node-to-node networking, no gossip. Every serving host can perform a repository's refs and object work, including during a deploy. - **CPU may be throttled between requests** on serverless platforms; background work must stay off the serving runtime and run on the bounded bulk runtime. - **Object store facts**: ~60–80 ms per GET, ~100 MB/s per connection (stripe for more), conditional GET ~15 ms; same-object overwrite is serialized (~1 write/s) — a single CAS'd object is a throughput cap. ### 1.2 What follows - **An instance must become useful in seconds**, cold: refs in < 1 s (one manifest GET + snapshot + tail). - **Everything an instance computes that is immutable is cached for everyone**: in-process LRU and, where a second instance would otherwise recompute, the bucket itself (`cache/api/v1/*.json`, `cache/archive/v1/*`). Wiping every instance loses nothing but warmth. - **No silent waiting.** Long work is a *task* (id, log, lock, attachable stream) and is narrated to the client: SSE envelope for API clients, sideband band-2 lines for git. "Cloning into… and then nothing" is a bug. ### 1.3 Security contract (`Config::validate` fails closed) - Five auth modes (`server.auth_mode`): **`none`** (everyone is `anon` with write and admin — `validate` refuses unless `server.listen` is loopback), **`jwt`**, **`introspect`**, **`forwarded`**, and **`introspect_forwarded`** (D48). JWT mode accepts EdDSA tokens from a Git Basic password or API Bearer header, verifies `[auth.jwt]` public-key PEM or cached JWKS plus issuer/audience/times, and applies `/:read|write|admin` scopes. Introspect mode sends the same Basic password or Bearer credential to `[auth.introspect].url` as JSON `{"token":"…"}`, with a service bearer secret read from `secret_env` at startup. The issuer returns `active`, opaque `principal`, repository `scopes`, and optional `ttl` seconds; it owns tokens and revocation. Both modes share permission checks. Introspection answers are bounded, SHA-256-keyed process caches; service failures return 503 + `Retry-After: 5`, never 401 or stale grants after expiry. Scope misses are 404; missing/invalid credentials are 401 with `WWW-Authenticate: Basic realm="gitcask"`. Client `X-Gitcask-Principal`/`-Write`/`-Admin` headers are ignored. In forwarded mode the authenticating proxy supplies `X-Gitcask-Principal`; `X-Gitcask-Write: 1` and `X-Gitcask-Admin: 1` grant those permissions. If `GITCASK_FORWARD_SECRET` is set, `X-Gitcask-Forward-Secret` must match it. `Authorization` is ignored. Combined mode requires both configurations and selects exactly one identity per request; [SECURITY.md](/llms/security.md#mixed-direct-and-trusted-proxy-authentication) defines precedence and the proxy boundary. - `/healthz` and `/readyz` are open at the application. With `server.public_docs = true` (default), `/docs` and `/openapi.json` and their permanent redirects under `/api/v1/` are open too; `false` requires valid credentials for all four docs paths. `/metrics` always requires auth. Everything else requires valid credentials. - **An edge announces byte offload, per request, in `X-Gitcask-Capabilities`** (D39): `accel-redirect` means static bytes may be served by `X-Accel-Redirect`, honoured only when `server.accel_redirect = true` **and the TCP peer is loopback**). Hit directly, nothing is assumed. - Tests never write the user's global git config (private `GIT_CONFIG_GLOBAL`). ### 1.4 Requirements (the bar) - Full git surface: smart HTTP v0/v2, ls-refs with prefixes, fetch with filter/shallow/deepen/sideband-all, receive-pack (atomic, delete, tags, push options, report-status-v2), LFS batch/basic transfer, `/` namespaces. Every read on every instance is as fresh as a fetch: push acknowledged ⇒ the next request anywhere sees it. - Cost must not scale with ref count on any hot path (O(1) `refs`, O(k) `resolve`, paged ref lists, prefix filtered advertisements). --- ## 2. The WAL — source of truth, and how we squeeze value out of it ### 2.1 Objects and the commit point (`repos///…`) | Object | Role | |---|---| | `manifest.pb` | Tiny, **CAS-rewritten**: `head_seq`, live pack set `PackRef[]` (checksum, sizes, tier, has_rev/bitmap), log segments, checkpoint pointer, `revision`, `updated_at`. **The linearization point.** Nothing is visible before its CAS; everything after is idempotent and replayable. | | `log/.pb` | Immutable, uvarint-framed `LogEntry` frames: PUSH / REF_UPDATE (ref transaction + pack pointer), COMPACT (new pack, `supersedes[]`). Strictly increasing `seq`. One small object per publish batch. | | `wal/.pack/.idx/.rev/.bitmap/.commit-graph` | Immutable packs, content-addressed by pack checksum: push packs (tier 0), compaction outputs (tier 1), plus the side-files a reader needs. | | `checkpoints//checkpoint.pb`, `refs.pb` | Folded state at `seq`: live pack set + full `RefSnapshot`. Cold start = snapshot + tail, never full replay. | | `leases/.pb` | CAS lease with TTL heartbeat, such as `compact`. The only cross-instance mutex. | | `cache/api/v1/.json` | Shared render cache of immutable web API answers. | | `cache/archive/v1/.` | Shared immutable prefix-free `git archive` result (one per commit/format), served through the static-byte path. Free-form `?prefix=` variants are never stored here. | | `fsck.pb` | Last connectivity audit (`FsckReport`), written by the maintainer's `fsck` unit and exported as a metric (`docs/INTEGRITY.md`). | | `gc.pb` | Last completed bucket-GC compaction/checkpoint cursors (`GcState`); overwritten after conditional deletes finish, safe to lose and recompute (D43). | | `events/cursor.json` | Durable acknowledged WAL sequence of the events bridge; advanced only after the webhook acknowledged (D32). | | `lfs/objects///` | LFS objects (sha256-addressed, immutable). | | `pending//` (bucket root) | Empty marker written after every successful manifest CAS; the maintainer's work queue (D40). | Schema `crates/gitcask-proto/proto/gitcask/v1/wal.proto`; the S3 backend and test-only in-memory store share one contract suite (`crates/gitcask-store/tests/contract.rs`, incl. compose). ### 2.2 Write path receive-pack (ours, `gitcask-git/src/receive.rs`) → receive the pack body to request EOF into an unlinked spool file (`LocalRepo::spool_pack`, bounded by `server.max_push_bytes`) while the full sync runs; over HTTP/1.x a push being received gets no status, header or band-2 byte before request-body EOF (D52; refusals decided before or during reception are complete responses and may come earlier) → index the pack locally (`git index-pack --stdin --fix-thin --keep --rev-index --threads=0`, `--fsck-objects` when `wal.fsck_objects`) in a per-ingest scratch git dir (a rejected push leaves nothing behind) → connectivity per config (`spawn_blocking`) → `pack PUT ∥ idx PUT ∥ log PUT` → **manifest CAS** (group commit per repo per instance, `wal.batch_window`) → commit local ref txn → `ok` to the client → best-effort `pending//` marker PUT (a failed marker never fails the push). On 412: refetch, re-validate every old value (moved ref ⇒ `ng`), re-seq, retry with jittered backoff. Publish is CAS-safe for concurrent writers by construction. Never ACK before the bucket ACKs. ### 2.3 Read path — sync levels (`RepoHandle::sync_*`, `gitcask-wal/src/handle.rs`, `sync.rs`) Every request: conditional GET of `manifest.pb` (skippable for `wal.freshness_ttl`) → 304 serve / 200 apply. | Level | Brings | Used by | |---|---|---| | **Refs** | checkpoint `RefSnapshot` + every log entry's ref txn → `packed-refs`. No packs. | `info/refs`, `ls-refs`, web `refs`/`resolve`/overview, read_log | | **Full** (`sync_full()`) | Refs + every live pack local (striped parallel downloads) | upload-pack, receive-pack, web object endpoints, compaction | Long syncs register as `materialize` tasks and stream progress; pack work runs on the **bulk runtime** and never takes the refs phase's lock (D19). Local disk uses idle and watermark eviction. ### 2.4 Getting the most out of the WAL (the strategies) - **Refs-first everything.** Ref advertisement, peeled tags (`RefUpdate.new_peeled` recorded by the writer so replicas advertise `^{}` without objects), web `refs`/`resolve` — all from snapshot + tail, no objects. - **Side-files published with packs**: `.idx`, `.rev`, `.bitmap`, and split commit-graph layers. Every pack writer produces `.rev` (`pack.writeReverseIndex`; git ≥ 2.47 on the server); a published pack ≥ 250 k objects without one gets it from the maintainer (`rev-index` unit) — without it git rebuilds the reverse index in memory on every `pack-objects` (2.85 s per fetch on a 60 M-object repository). - **Checkpoints fold the log** when `wal.snapshot_every_entries` OR `wal.checkpoint_interval` OR `wal.checkpoint_tail_bytes` fires — the checkpoint is the unit of serving state. Writing one is refs-level work (`sync_refs_only`, no packs downloaded), evaluated only for repositories with a pending marker. - **Provenance for free**: every push and repack is a log entry; `gitcask wal ls|show|materialize --at-seq`. - **Never LIST on a hot path**; 404s are free; probe, don't list. Immutable objects get `Cache-Control: public, max-age=31536000, immutable` + strong ETag + Range everywhere (D10 static contract). ### 2.5 Compaction (WAL + git), leader by lease - Maintainers run **geometric** folding of fresh packs (tier 0 → tier 1, `git repack -d --geometric --write-midx`) under `leases/compact.pb`; the result is a COMPACT entry; followers download the new pack and drop superseded ones after in-flight readers finish. Triggers: `compaction.trigger_packs`, `trigger_bytes` — and at least **two** fresh packs (one pack folds into itself). - Superseded packs are retained `compaction.retention_superseded` (provenance window) then GC'd. ### 2.5b Self-healing by construction Everything the maintainer produces — checkpoints, compactions and retention — is a **pure function of (config, WAL state)**. The maintainer does not run schedules; for each repository it computes the *desired state* and performs **one bounded unit of the most important missing work** at a time (checkpoint → compaction → rev-index → fsck audit → bucket GC), as a task, under a lease, until the repository is idle. A deleted or corrupt artefact is "missing" and rebuilt identically; config changes take effect by re-planning; there are no one-off backfill scripts. **Which repositories** (D40): never all of them. A pass lists `pending/` (at most `maintenance.max_repos_per_pass` markers, with their versions), works each marked repository to idle, then deletes its marker conditionally on the listed version — a push that arrived meanwhile rewrote the marker, so the delete fails and the repository is seen again next pass. A repository nobody pushes to is never visited; time-based triggers (checkpoint age, fsck interval) apply only to marked repositories. Cost is proportional to pushes, not to repository count. The only other periodic LIST is `maintain/` once per 10 minutes; it scales with live instance count and expires stale maintainer heartbeats. ### 2.6 Tasks, progress, narration (`gitcask-wal/src/tasks.rs`, `crates/gitcask-server/src/sse.rs`, `smart.rs`) Any long work = a task: unique id, per-instance log (`GET …/tasks`), `(repo, kind)` lock (a second start joins), replayable packet stream (`GET …/tasks/{id}`). Packets: `notice`, `progress {label,done,total?,unit,percent?}`, `task`, terminal `result` | `error`. Web: any JSON endpoint that cannot answer immediately streams the SSE envelope when the client accepts `text/event-stream`; fast answers stay plain cacheable JSON. Git: v2 `fetch` advertises `sideband-all` and narrates on band 2 (`remote: * …`). `no-progress` is honoured. --- ## 3. Principles (what a PR is judged against) One idea: **the bucket is the repository; everything else is a cache or a reader of the log.** Ten consequences. A change that feels natural on a conventional git host (a table, a queue, a cache server, a webhook from the push path, a bigger disk) is usually a violation here. A violation means the principle is wrong — amend it with a decision in §4 — or the PR is; never "fix later". | # | Principle | The tell in a PR | The question to answer | |---|---|---|---| | **I** | **No state outside the object store.** Disk and memory are caches. | A database, Redis, SQLite, a file that must survive a restart, an env var that encodes data. | "If every instance is wiped now, what is lost?" — must be "warmth". | | **II** | **The manifest CAS is the only commit point.** Immutable objects are never overwritten (`PutMode::Overwrite` only on the manifest, leases, fsck.pb, gc.pb, events/cursor, maintainer heartbeats, render cache). | A second "commit" (a flag file, a list update that makes data visible), an ACK before the bucket's. | "What does a client on another instance see between the PUT and the CAS?" — nothing new. | | **III** | **Side effects are readers of the WAL, never steps of a write.** Events and notifications tail the log from a durable cursor. | A webhook/HTTP call from `receive.rs`, `publish.rs`, `smart.rs`. | "If this side effect fails, does the push?" — no. "Is it replayable from the cursor?" — yes. | | **IV** | **Every read revalidates; there is no eventually.** | A cache that outlives the manifest's generation, a TTL invented for a repo-scoped answer, a read that skips `sync_*`. | "After `push` returns `ok`, can any instance serve the old state?" (`cargo test -p gitcask-server --test sim`). | | **V** | **Packs are local; never a hard-coded host.** | An object path that skips full synchronization, a hostname in `crates/`. | "Has the full pack set been synchronized before this reads objects?" | | **VI** | **Never block the async runtime; bulk bytes never share a lane with the control plane.** | `Command::new(...).output()` or `std::fs` big reads in an `async fn` outside `spawn_blocking`/the bulk runtime; a queued writer on `RepoHandle::rw` from an install path. | "On which thread does this run, and what holds `sync_mutex`/`rw` while it does?" (e2e `blocking_work_in_the_install_path_does_not_stall_requests`). | | **VII** | **No LIST on a hot path; count the round trips.** | A `.list(` in request handling; a new GET "just to check"; a protocol change without a `docs/ROUNDTRIPS.md` row. | Before/after depth in the commit; the sim asserts request budgets (`FaultStore::stats().ops`). | | **VIII** | **Standalone first; the edge announces, the app never assumes.** | Reading an `X-Gitcask-*` request header without checking `X-Gitcask-Capabilities`; a feature only testable behind nginx. | "Does this work on `gitcask.standalone.toml` with nothing in front?" | | **IX** | **No silent waiting.** | A new op that blocks a handler, a loop without a task, a `sleep` a client would wait on. | "Where does the client see this taking time?" — a task kind, a progress packet or a band-2 line. | | **X** | **Keep gitcask small.** Upstream `git` does git things; gix only where measured faster and correct; one config file, one auth story, one SDK; scope is `GOAL.md §4`. Proto append-only. | A new dependency without a why; a reimplementation of something `git` does; an alias/shim/"legacy" branch; a removed proto field. | "Which line of GOAL §4 is this for?" | --- ## 4. Decisions in force (append, never silently change) - **D1** Rust, tokio + axum, HTTP/2 (h2c or ALPN), streaming both ways, gzip request bodies. - **D2** gix in-process where it is correct and measured; upstream `git` for upload-pack and delta-compressing repack. - **D3** `ObjectStore` trait with CAS version tokens, conditional GET, range, compose and edge-fetch capabilities. Production backend scope is governed by D44. - **D4** protobuf on the wire and in the bucket; schema versioned, append-only. - **D5** Repo identity `/[.git]`, prefix `repos///`, creation = CAS create of the manifest. - **D6** Manifest CAS is the only commit point. **D7** No node identity, no elections; leases for exclusivity. - **D8** `gitcask.toml` only (+ `GITCASK__` env overrides). **D9** One binary, roles by config (`serve`, `maintain`, `events`; `maintain` includes compaction). - **D10** One static-serving code path for every immutable byte (ETag/304/If-Range/Range/416/HEAD/immutable; UI assets precompressed at build; store objects never compressed at request time). - **D12** (**superseded by D46**) `server.auth_mode` was `none` | `forwarded`; a front proxy authenticated clients and sent principal/write/admin grants to gitcask. - **D13** Long work is a task and is narrated (§2.6); no endpoint may block silently. - **D14** (retired) No embedded web UI: gitcask is an API + git server; any UI lives in the calling platform. - **D15** Repo-scoped API lives at `/{o}/{r}/api/…` (refs, resolve, tree, blob, commits, commit, compare, overview, ops, tasks). The browser lane is `/{o}/{r}/api-browser/…`; `/api/v1` is non-repository discovery. No aliases. - **D19** **The serving runtime is untouchable.** (1) control-plane store objects (manifest, log, checkpoints, leases, render cache) do not queue behind bulk bytes; S3 bulk GETs use presigned HTTP while control-plane operations use the SDK path; (2) pack materialization never runs on the serving runtime (`sync::on_bulk_runtime`) and never queues as a writer on `RepoHandle::rw` — the refs phase needs only `sync_mutex`, pack removal is `try_write()` (a queued writer on a tokio RwLock blocks every new reader; one 24-minute clone once starved every info/refs for minutes). The bulk runtime has `cache.bulk_threads` async workers (default 2); blocking filesystem/git work uses its blocking pool. - **D20** **One API, two lanes**: `/{o}/{r}/api/…` is the direct lane; `/{o}/{r}/api-browser/…` is the cross-origin browser lane (CORS only for `server.cors_origins`). - **D23** (**authentication clause superseded by D46**) **An edge was load-bearing for authentication and optional for bytes.** A front proxy terminates TLS, authenticates clients, routes by `//`, and may offload static bytes by `X-Accel-Redirect` — only when it injects `X-Gitcask-Capabilities: accel-redirect`; the app's answer supplies `X-Gitcask-Store-Url` / `-Authorization` / `-Key` and the edge slices and caches. `deploy/nginx.conf.example` is the reference; nothing in `crates/` knows a hostname. - **D26** **Routing is by repo prefix, nothing else.** Everything whose path starts with `//` (git smart HTTP, `.git` suffix, LFS, `/{o}/{r}/api*`) is one routing unit; an edge maps `^//[./?]` → a host. "Which machine serves a repo" is decidable from the first path segments alone (the `Server` header shows it). **D27** Lanes are a segment *after* the repo prefix (`/api`, `/api-browser`); the prefix is still the only routing key. - **D31** **Draining must not stop serving.** *Phase 1* (SIGTERM): the maintenance loop starts no new unit and the running unit is **interrupted at once** (the next pass redoes it) while the instance serves everything normally, `/readyz` 200; bounded 30 s. *Phase 2*: `/readyz` 503 + Retry-After, new fetch/push/LFS refused with 503 before any work, in-flight requests get `server.drain_timeout`, exit. Test `tests/drain.rs`. - **D32** **Events are produced from the WAL by one small service, never by the push path** (`docs/EVENTS.md`). The **events bridge** (`roles=["events"]`) tails each repo's log from a durable per-repo cursor (`events/cursor.json`), converts committed PUSH/REF_UPDATE entries to `ref` events, POSTs each batch to `events.webhook_url` (JSON array; `X-Gitcask-Delivery` = sha1 of the body; `X-Gitcask-Signature: sha256=` when `webhook_secret` is set), advances the cursor: published iff durable, a crash loses nothing, lag = `head_seq − cursor`. Writers and the WAL crate contain no event code. Wake-ups: `POST /_events/notify` from a bucket notification (at-least-once) + a periodic sweep as backstop. Dedup key `(repo, seq, ref_name)`. - **D40** **The maintainer is driven by pending markers, never by a scan** (2026-08-30, `docs/DIRECTION.md` D1). The publish path writes `pending//` after the manifest CAS (best-effort, `Overwrite`); `run_pass` lists `pending/` oldest-first and works repositories concurrently; a repo that reaches compaction runs to idle in that pass, while a checkpoint-only repo yields after one unit with its marker intact. Completed markers are deleted with their listed versions. There is no `registry.list()`, no LIST of `repos/` and no per-repository HEAD on any periodic path. Repository listing and metadata belong to the calling platform's database. - **D41** **Packs are always local** (2026-08-30). `SyncLevel` is `Refs` | `Full`; there is no remote reader, no bucket mount, no prewarm, no tier-2 base or history pack. A repository that does not fit the instance is a sizing problem, not a code path. - **D42** **Transient store errors are retried in the store** (2026-08-30): 5xx / 429 / throttling / connection failures on GET, HEAD, LIST, DELETE, multipart and unconditional PUT are retried with full-jitter backoff up to `store.max_retries`; conditional PUTs are never retried at that layer (an ambiguous success cannot be told apart from a lost write) — `coord::cas_update` owns CAS retry. - **D43** **Bucket GC is a lowest-priority, per-repository WAL reader** (2026-08-31). A maintainer runs it only for a pending repository after a COMPACT or checkpoint newer than `gc.pb`, after fsck, under `leases/gc.pb`. It lists that repository prefix once, retains the current manifest pack set, packs referenced by WAL entries inside `compaction.retention_superseded`, and packs in retained checkpoints; then it revalidates the manifest generation and conditionally deletes superseded pack families, folded logs, and old checkpoint directories by their listed versions. From the same listing, it also conditionally deletes `cache/api/v1/` and `cache/archive/v1/` objects older than `cache.shared_retention`; these two prefixes are the exception to the otherwise out-of-scope repository-prefix orphans. The latest checkpoint and every checkpoint inside the retention window remain. `wal materialize --at-seq` discovers those retained checkpoints; a sequence whose inputs were collected fails explicitly as beyond retention. `gc.pb` advances only after all deletes, is not a commit point, and may be lost and recomputed. There is no LIST across repositories and no time-driven visit; repository-prefix orphans unrelated to a committed supersede/checkpoint (other than the two shared cache prefixes above) are out of scope. - **D39** **gitcask is a standalone program; any deployment is packaging.** `gitcask-server --config x.toml` (a thin bin = `gitcask serve`; `gitcask` with no subcommand serves too) works on one machine against one bucket in loopback-only `none` mode. It serves HTTP/1.1 + h2c; the front proxy terminates public TLS. Byte offload is announced per request in `X-Gitcask-Capabilities`, never assumed. `cache.dir` defaults to `/tmp/gitcask`. A missing `--config` file is fatal (exit 2); `--config /dev/null` is the explicit defaults+env form. - **D44** **S3 is the sole production object-store backend** (2026-08-31). The in-memory and fault stores are test-only behind the `testing` feature. GCS-specific configuration, dependencies and protocol branches were removed; another production backend must justify and implement the complete store contract before returning. - **D45** (**superseded by D46**) **Authentication was a stateless, separate gate** (2026-09-01). `gitcask-gate` was the public process; gitcask remains on its unchanged `forwarded` contract. Mint mode issues short-lived HS256 JWTs carrying only an opaque principal and repository scopes; static mode is configuration-only for local play and CI. The gate stores no identities, sessions, revocations or usage history, strips every client `X-Gitcask-*` header, and streams all bodies except the bounded LFS batch JSON. It shares `gitcask.toml` with the server. D9's one-binary rule applies to repository roles; the authentication boundary is intentionally a second process. - **D46** **gitcask verifies asymmetric tokens itself; issuance and identity stay outside** (2026-09-01; extended by D47). `server.auth_mode` is `none` | `jwt` | `forwarded`; `jwt` is the one-process standalone/public path. Only EdDSA (Ed25519) is accepted. The issuer keeps the private key; gitcask has a public-key PEM or cached JWKS, refreshes JWKS only on `kid` miss, and retains the last successful set on refresh failure. Claims are opaque `sub`, repository `scopes`, `exp`, `iat`, `jti` (plus configured `iss`/optional `aud`; `nbf` is honoured). Git sends the token as the Basic password and APIs use Bearer. Existing handler `require_read`/`require_write`/ `require_admin` calls are the sole path-permission table; scope misses return 404. There is no issuing endpoint: platforms sign the format themselves and self-hosters use offline `gitcask token keygen|mint`. The gate crate, HS256/static credentials, shared forward secret in the standard deployment, and second process are deleted. `forwarded` remains only for deployments that already have a trusted IdP proxy. gitcask stores no users, sessions, revocations, or token-use history. An edge terminates TLS and may offload bytes, but is not required for authentication. - **D47** **Opaque tokens are verified by issuer introspection** (2026-09-20; extends D46). `server.auth_mode = "introspect"` uses the RFC 7662 active-token model with gitcask's explicit JSON contract: POST `{"token":"…"}` to an HTTPS (or loopback HTTP) URL, authorized with a service Bearer secret named by `auth.introspect.secret_env`. A 200 answer carries `active`, non-empty `principal`, the same repository `scopes` as JWT, and optional `ttl` seconds. JWT and introspection share Basic/Bearer extraction and the handlers' permission path; no identity or token-issuance service enters gitcask. A strictly bounded 10,000-entry FIFO cache uses SHA-256 token digests, never credentials; positive answers live `min(ttl, cache_ttl)` (default 30 s, maximum 10 min), inactive/invalid token answers use `negative_cache_ttl` (default 3 s). Concurrent misses share one call, including failure; the in-flight table is also capped at 10,000 and cancellation releases waiters. Shared answers carry the cache's absolute expiry; a late follower gets 503 rather than an expired grant. Zero-TTL answers coalesce only the current flight. Duplicate known response fields are uncached service errors; a present non-integer `ttl`, including null, is an invalid token answer. FIFO gives constant-time eviction without another cache dependency. These answers are warmth: restarting loses only latency, never authoritative identity or revocation state. Platform revocation is observed after the cached grant expires; zero TTL disables positive caching. Non-200, malformed JSON, oversized (>64 KiB) responses and transport errors are uncached 503 + `Retry-After: 5`; 401/403 from the issuer warns that the service secret was rejected. The request timeout defaults to 2 s and is capped at 10 s. The bounded cache and flight table are the only new state; no bucket requests are added. `forwarded` remains for trusted IdP-proxy deployments. - **D48** **Introspection and trusted forwarding may share one listener** (2026-09-21; extends D47). Explicit `server.auth_mode = "introspect_forwarded"` requires valid introspection settings and a non-empty, printable ASCII `GITCASK_FORWARD_SECRET` without whitespace at startup. A present proxy-secret header selects forwarding only after constant-time verification; any invalid, empty or repeated secret fails closed, with no token fallback. An absent secret selects the existing bounded introspection client and repository scopes, ignoring forwarded grants. A trusted proxy identity wins over `Authorization` and must supply a principal; privileges never merge. Standalone modes retain their contracts. No bucket requests, routes, ports or identity state are added. The precise error semantics and mandatory proxy header stripping live in [SECURITY.md](/llms/security.md#mixed-direct-and-trusted-proxy-authentication). - **D49** **S3 credentials may come from the AWS SDK default chain** (2026-09-23). `store.s3.credentials = "static"` remains the default and reads the configured key environment variables once at startup, including an optional session token. `"default"` uses the SDK's refreshing credential chain (environment, profile, ECS task role, IMDS); the configured key environment variable names are ignored. Region, endpoint and path-style settings apply in either mode. Presigned URLs and edge byte offload require the SDK to obtain signing credentials. - **D50** **API documentation is public by default** (2026-09-23). Exact `GET /docs` and `GET /openapi.json` serve the bundled Scalar UI and compile-time OpenAPI schema without credentials when `server.public_docs = true` (default). `false` requires authentication on both paths and the permanent `/api/v1/docs` and `/api/v1/openapi.json` redirects. `/metrics` remains protected. These one-segment paths do not consume repository names: `/{owner}/{repo}` remains available with `docs` or `openapi.json` as an owner. Decision identifiers are stable; gaps in the numbering are intentional. - **D51** **Pristine repositories can be initialized from a pinned Gitcask tree** (2026-09-28). `POST /{o}/{r}/api/initialize` creates one parentless commit from a full source commit OID, with caller-supplied identity, time and message, and independently uploads its complete tree closure into the destination's normal packs. The existing WAL publisher checks pristine manifest state across the batch and every CAS retry; its same manifest CAS records one bounded initialization receipt. Exact authorized retries return that original result without changing later writes or opening the source. Source read and destination write use the same resolved principal. No source history, durable alternates, shared snapshot, external fetch, checkout or second commit point is introduced. Same-format SHA-1 and SHA-256 are supported; mismatches and LFS pointer trees are rejected. Gitlinks remain external references. The complete contract and all-writers upgrade requirement are in [docs/INITIALIZE.md](/llms/initialize.md). - **D52** **Over HTTP/1.x a push being received is answered only after the request body ends** (2026-09-29). A proxy that forwards the request body over HTTP/1 may stop forwarding it once the upstream answers: behind an AWS ALB (HTTP/2 to the client, HTTP/1 to gitcask) a side-band-64k push whose banner went out after the ref commands stalled forever in the body read, while the same pack without side-band took 0.8 s. The pack is received to request-body EOF into an anonymous (already unlinked) file under the repository's `objects/pack/`, bounded by `server.max_push_bytes`; gzip bodies are decoded to the end of the body. The full sync runs concurrently with the upload. One rule in `receive_pack` decides when the response starts: a side-band push over **HTTP/2** (full-duplex streams) narrates live from the start; over **HTTP/1.x**, and for every push without side-band, a push that is being received gets no status, header or band-2 byte before request-body EOF. Refusals decided before or during reception (auth, URL, drain, parse, oversize, read error) are complete responses and may come earlier, which is safe behind a proxy because nothing is awaited from the body afterwards. The HTTP version is that of the connection reaching gitcask (the last hop): a direct or standalone git client over `http://` and every HTTP/1.1 proxy hop take `after_eof`; `live` is reachable only when a proxy speaks prior-knowledge h2c to gitcask (or terminates HTTP/2 over TLS in front of it and forwards h2c). The side-band task writes its whole response (banner, sync narration, index-pack, connectivity, publish, report) to one non-blocking channel; the rule only chooses when that channel starts being copied to the client, so narration produced during an HTTP/1.x upload is buffered and replayed in order after EOF, and the client shows its own upload progress meanwhile. Deployment caveat: a proxy that speaks HTTP/2 to gitcask but HTTP/1.x to the client must be proven to keep forwarding the request body after an early response; otherwise terminate at a proxy that talks HTTP/1.1 to gitcask. Reception holds no ingest lock, a reception failure cancels the sync (dropping its read guard), counts as `gitcask_push_refused_total{reason="body"}` and is answered as a complete refusal (streamed when already live), and `receive.body` records `http_version` and `response_start` (`live` | `after_eof`) while `git.spool_pack` records received `bytes` and its `outcome` (`eof`, `too_large`, `error`, `cancelled`) on every exit. No bucket requests change. - **D53** **Full-history pristine import is pin → independently acquire → manifest CAS** (2026-10-05). The platform persists the stateless resolve snapshot before invoking the bulk import and owns durable jobs/retries/permissions; gitcask owns Git objects and one same-CAS manifest receipt. Heads/tags and symbolic or detached HEAD are preserved with their complete reachable closure, never a source push or durable alternate. Pristine is checked across every batch/CAS retry. Exact authorized retries return the receipt without reopening source or changing later target writes; target-read receipt reconciliation requires no source grant. External v1 is isolated public HTTPS SHA-1 only, DNS pinned and redirects refused, with resource/deadline bounds; LFS history is refused and gitlinks never recurse. Snapshot initialize remains a distinct valid operation. Append-only detached-HEAD checkpoint and receipt fields are preserved by all writers. See [docs/IMPORT.md](/llms/import.md); no engine job DB, identity store or new auth path is added. - **D54** **The website is a separate app: value for people, the canonical Markdown for agents** (2026-10-07; hosting superseded by D55). `site/` is a Next.js app deployed to Cloudflare Workers through OpenNext; it is not built into, linked with or served by gitcask, which stays an API + git server. The landing page argues why to use gitcask with one interactive object per claim; `/docs` pages are hand-written diagrams and tables for people. Text stays where it lives: before every build `site/scripts/llms.mjs` publishes `README.md`, `GOAL.md`, `SECURITY.md`, `CONTRIBUTING.md`, this file, `docs/*.md` and `gitcask.example.toml` verbatim as `/llms.txt`, `/llms-full.txt` and `/llms/.md` (`docs/reference/` is excluded: it excerpts Cursor's post). The Markdown remains the only home of each fact; a site page that shows a fact is a view of it and changes with it (§5). Every number on the site traces to these documents. Pages are prerendered and served from the static-assets incremental cache, which the deploy populates (`opennextjs-cloudflare deploy`, or `populateCache remote` before `wrangler deploy`); no page reads files at runtime. Its design system is the comwit UI template's tokens and components (`site/design.md`). No crate, config key, route or bucket request changes. - **D55** **The website deploys from local files through the Comwit Cloud CLI** (2026-10-07; supersedes D54's hosting). `site/` uses `@brrrd/adapter` to turn `next build` into `dist/brrrd/`, including prerendered pages and the generated Markdown assets. `comwit deploy --package dist/brrrd` uploads that package to the `gitcask-site` app in the `gitcask` Cloud project and serves `gitcask.comwit.io`. There is no Git integration or automatic deployment; the existing site CI only validates the build. OpenNext, Wrangler, and their configuration are removed. `site/README.md` owns the target IDs, CLI commands, adapter/Next version constraint, and custom-domain setup. The site remains independent of the gitcask server and bucket. --- ## 5. Working rules - **No backwards compatibility (pre-1.0, banner at top):** change the shape and delete the old one in the same commit — no aliases, shims, deprecated routes/keys/fields. Only bucket formats (WAL, manifest, log, checkpoint protos) stay append-only/replayable. - Keep this file current: append decisions with a number and a date; never delete history, replace it with the decision that superseded it. - **Never block the async runtime**: no blocking git/fs work (repack, midx, commit-graph, gix reopen, large reads/copies) on a tokio worker — `spawn_blocking`; every `refresh()` on an async path is `refresh_async()`. Pack materialization runs on the **bulk runtime** (`sync::on_bulk_runtime`). The runtime watchdog logs "async runtime stalled" with `inflight` and `tasks_running`: `inflight = 0` at a late tick ⇒ the platform paused the process, `inflight > 0` ⇒ a real stall — look at `lock_wait_max_ms` and `gitcask_lock_wait_seconds{lock}`. Bulk bytes never queue on the serving runtime; S3 bulk GETs use their own presigned-HTTP path. - **Correct is not sufficient.** Every protocol change (publish, sync, leases, checkpoints) is also judged on critical-path round trips against the bucket — read `docs/ROUNDTRIPS.md`, update its budget table, put before/after depth in the commit, keep verification on the failure path, assert request budgets in the sim. - **Standalone first (D39, D46, D47):** repository features and JWT/introspection authentication work by hitting one gitcask process directly with no external edge. Bytes are streamed by gitcask. Anything another edge takes over is announced per request in `X-Gitcask-Capabilities`; never infer an edge from config, never hardcode a hostname in `crates/`. - **S3 is the production store.** Every store feature runs against the S3 contract suite (`just test-s3` against rustfs) and the test-only memory implementation. - **Use the rig before prod** (`docker compose up --build -d` → authenticated gitcask on :8080, or simply `scripts/smoke.sh . 8090` against an existing rustfs). - No new auth paths (§1.3). No LIST on hot paths. No unbounded buffering of packs in memory. No silent long operations (make it a task, narrate it). - Every immutable response: `immutable` + strong ETag + Range; every ref-dependent response: SWR + ETag. - Before changing the wire/store formats: proto is append-only; manifests/log entries must stay replayable by old readers within the retention window. - Config: `gitcask.example.toml` documents every key; change it with the code. - Site: a page under `site/` that shows a fact (a number, a route, a status code, a config key) is a view of the Markdown that owns it; change both in the same commit (D54). - Test tiers: `just test` (fast, < 1 min), `just e2e`, `just warnings` (no unused/dead-code rustc warnings; the deliberate `unsafe_code` warns are not part of that gate), and `scripts/clippy-count.sh` (the `[workspace.lints]` set is deliberately *warn*-level and the tree carries historical warnings; the gate is **no regression against the base branch**, and `#[allow]` is never added to pass it); `scripts/smoke.sh` against rustfs before merging anything that touches smart HTTP, publish, sync or auth; the **simulation suite** `cargo test -p gitcask-server --test sim` (fault links per instance over one truth store: crash, partition, stale, lost response, orphan scenarios + randomized seeds `GITCASK_SIM_SEEDS`/`GITCASK_SIM_SEED`); `just test-slow` (ignored benches); `tests/e2e.sh` against a running server (`GITCASK_E2E_BASE_URL`). Never `cargo test --workspace --no-fail-fast` in a session; wrap ad-hoc cargo in `timeout`. # Round trips are the cost model Context: **who needs to read this, and when.** Anyone (human or agent) touching a protocol that talks to the bucket: the publish path (`gitcask-wal/src/publish.rs`), sync levels and freshness (`sync.rs`, `handle.rs`), checkpoints, compaction and its lease (`gitcask-server/src/ops.rs`, `coord.rs`), LFS, or the `ObjectStore` backends themselves — and anyone writing or triaging simulation scenarios (`crates/gitcask-server/tests/sim.rs`), because a sim fix that is correct but adds a happy-path request is a regression, not a fix. Read it *before* designing the change, use §5 when writing the commit/PR, and keep §2 current. It exists because gitcask runs on tmpfs host instances that own nothing but a bucket: every user-visible latency is a sum of sequential store requests, and the manifest is a single CAS'd object with a hard write-rate cap — so the number and shape of round trips *is* the performance design. Referenced from AGENTS.md (read order and §5 working rules) and from the sim harness briefs. **Correct is necessary, not sufficient.** gitcask's only durable primitive is a bucket with ~60–80 ms per request, one serialized overwrite per second per object, and no transactions beyond single-object CAS. Every protocol we design on top (publish, sync, compaction, checkpoints, leases) is judged on two axes at once: 1. **Safety/liveness** — the simulation suite (`crates/gitcask-server/tests/sim.rs`) and the contract tests. 2. **Critical-path round trips** — how many *sequential* bucket requests a user-visible operation needs, and how many requests in total (cost, contention on CAS'd objects). A change that is correct but adds a sequential round trip to a hot path is a regression. A fix for a liveness bug that moves work onto the *failure* path and leaves the happy path at the same depth is the right shape. This document is the thinking tool; apply it to every protocol change and say so in the commit. ## 1. The latency model | Primitive | Cost | Notes | |---|---|---| | GET / PUT small object | 60–80 ms p50/p99 | one request = one round trip | | Conditional GET (`If-None-Match`) → 304 | 15–18 ms | the cheapest "is anything new?" | | HEAD | ≈ GET | never add one on a happy path (`download_object` skips it when `PackRef` already has the size) | | 404 | free | probe, don't list | | LIST | slow, paged, eventually-ish | never on a hot path (rule in AGENTS §5) | | CAS overwrite of one object | serialized, ~1 write/s | a CAS'd object is a throughput cap; 412 is the normal contention signal | | Range read of a big object | ~100 MB/s per connection | stripe for more; bulk bytes on their own pool | | Compose | 1 request, no data transfer | ≤ 32 sources | ## 2. Budgets to defend (happy path, sequential depth → total requests) | Operation | Depth | Requests | Where | |---|---|---|---| | Authentication (introspect mode) | 0 on a cache hit; 1 HTTP request to the introspection URL per cache miss | 0 store requests; ≤ 1 HTTP request per effective cached TTL per token per instance while resident; concurrent misses coalesce. Issuer TTL may shorten `cache_ttl`; inactive answers use `negative_cache_ttl`; evictions and uncached service errors may require another call | `auth/introspect.rs` | | Any read (`info/refs`, ls-refs, web refs/resolve) | 1 cond GET (or 0 within `freshness_ttl`) | 1 | `sync.rs::freshness_check` | | Cold Refs sync | 1 manifest GET → 1 round (checkpoint refs ∥ log tail segments) | no checkpoint: 1 + tail (2 with one segment); checkpoint: 2 + tail | `registry.rs::open`, `sync.rs` | | Push request / publish (`process_batch`) | 1 freshness GET → pack PUT ∥ idx PUT ∥ log PUT (1 round) → manifest CAS (1 round) → best-effort pending-marker PUT | request: 6; already-synced publish: 5 | `publish.rs` | | Receive-pack tip validation (including empty packs) | Local lookups / connectivity walk after the existing Full sync, before publish; missing tips are refused before log PUT or manifest CAS | 0 extra; warm push remains 6 requests with a pack, 4 for ref-only | `smart.rs::receive_pack_process` | | Ref write API | 1 freshness GET → optional tag pack PUT ∥ idx PUT ∥ log PUT → manifest CAS → best-effort pending-marker PUT | ref-only: 4; annotated tag: 6 | `web/api/write.rs`, existing `publish.rs` | | Commit / merge write API | 1 freshness GET → new-object pack PUT ∥ idx PUT ∥ log PUT → manifest CAS → best-effort pending-marker PUT; a fast-forward publishes no pack; already-merged returns after freshness | commit / merge commit: 6; fast-forward: 4; already merged: 1; cold packs use the existing Full sync budget | `web/api/commit.rs`, existing `publish.rs` | | Initialize API (warm source and destination) | destination freshness GET → source freshness GET → pack PUT ∥ idx PUT ∥ log PUT → manifest CAS → pending-marker PUT; exact replay stops after destination freshness | first initialization: 7 requests / 5 rounds; exact replay: 1 / 1. Cold source adds existing Full sync work; no LIST or per-blob store reads. Receipt and pristine guard add 0 publisher requests | `web/api/initialize.rs`, existing `publish.rs`; `tests/initialize.rs::publication_failures_ambiguous_success_and_request_budgets` | | Import resolve, internal warm source/target | target permission/open (no freshness needed), source freshness GET; cold opens add ordinary manifest/snapshot/tail reads | 1 warm request; refs-level, no LIST or pack GET | `web/api/import.rs::resolve_source` | | Full-history import, warm source/target | target freshness GET → source Full freshness GET → pack PUT ∥ idx PUT ∥ log PUT → manifest CAS → best-effort pending marker | 5 sequential rounds, 7 requests; cold Full sync adds existing pack/side-file GET budget | `web/api/import.rs`, `publish.rs`; `tests/import.rs::fault_lost_response_crash_and_roundtrip_budgets` | | Full-history import exact replay / matching receipt | target freshness GET, no source lookup/materialization | 1 warm request | `web/api/import.rs`; same fault/budget test | | Import external public HTTPS | target freshness GET → scoped HTTPS acquire/validate → pack PUT ∥ idx PUT ∥ log PUT → manifest CAS → pending marker | 4 bucket rounds, 6 bucket requests; source HTTPS traffic is separately byte/deadline bounded | `import_transport.rs`, `web/api/import.rs` | | Archive API, no `prefix` (warm packs) | 1 freshness GET → archive GET (plain/304) or HEAD → Range GET; shared-cache miss adds a free 404 probe → archive PUT → serving read | hit plain/304: 2; hit range: 3; miss plain: 4; miss range: 5; cold packs use the existing Full sync budget | `web/api/archive.rs`, `static_object.rs`, `sync.rs` | | Archive API, with `prefix` (warm packs) | 1 freshness GET; `git archive` writes an unlinked local temporary file and the response streams from it; no archive-cache GET/PUT | 1; cold packs use the existing Full sync budget | `web/api/archive.rs`, `static_object.rs`, `sync.rs` | | Compaction publish | pack/side-file uploads → log PUT → manifest CAS; no pending marker | unchanged | `publish.rs::publish_compact_impl` | | Checkpoint | 1 cond GET (freshness) → refs PUT ∥ checkpoint PUT → manifest CAS | 3 rounds, 4 requests | `checkpoint.rs` | | Lease acquire | 1 GET → 1 CAS put (or 1 Create when absent) | 2 | `coord.rs::try_acquire` | | Maintainer pending pass | 1 bounded LIST, then each repository keeps its existing unit depth | existing requests per repo; up to `maintenance.workers` repos overlap | `maintain.rs::run_pass` | | Maintainer heartbeat GC | 1 LIST → heartbeat GETs → conditional DELETEs for expired hosts | 1 + live instance count + expired instance count; at most once per 10 min | `maintain.rs::heartbeats` | | Events backstop sweep | 1 `pending/` LIST → existing `catch_up` requests per distinct pending/cached repo | events-role periodic path only; no push/read critical-path requests | `bridge.rs::sweep` | | Bucket GC due check | 1 `gc.pb` GET, only after higher-priority units and only when a compacted pack/checkpoint exists | 1 | `gc.rs::due` | | Bucket GC unit | lease GET → lease CAS → manifest GET ∥ `gc.pb` GET → 1 repo-prefix LIST → retained log/checkpoint GETs (32-way) → manifest revalidation GET → conditional DELETEs (32-way) → `gc.pb` PUT → lease DELETE | 8 fixed + one GET per physical log/checkpoint + one conditional DELETE per collected object; shared-cache expiry reuses the listing (0 additional LIST, N cache DELETEs); no manifest CAS or user-visible critical-path requests | `gc.rs::collect` | | Web overview pending state | pending marker HEAD in parallel with opening the repo | +1 request, no added cold-open depth | `web/status.rs::overview` | | Publish, local commit (2026-08-23) | unchanged in round trips: after the manifest CAS the ref txns are applied to the local copy **before** the new manifest version is advertised, both under `sync_mutex` (the refs phase of every sync); the reverse order let a reader cache the old refs under the new version, and without the lock a concurrent sync replayed the same entry (two `update-ref`, a lock collision). A landed CAS is answered `ok` whatever the local apply does — the next sync replays (one conditional GET that then returns 200, no extra write). | 0 extra | `publish.rs::process_batch` | | Orphan log slot (failure path only) | +1 fresh manifest GET, +HEAD per probe, +Create at next seq | — | `publish.rs::claim_log_slot` | `healthy_request_round_trip_budgets` in `crates/gitcask-server/tests/sim.rs` pins the healthy MemoryStore counts at push **6**, warm refs **1**, cold refs with one tail segment **2**, and checkpoint **4**. Cold open used to spend an extra unconditional manifest GET (3 requests, 3 sequential rounds); it now applies the manifest it already fetched directly (2 requests, 2 rounds). `claim_log_slot`, `cas_landed`, and `put_immutable_create` add probes only after Create/CAS failure, so those failure helpers do not change their rows' healthy counts. Pending markers intentionally change a healthy push from 3 to 4 store-request rounds (5 to 6 requests): the marker must follow the manifest CAS so an uncommitted push cannot wake maintenance. Its failure is logged and counted but never changes the already-committed push result. Initialization adds one source freshness GET to the existing six-request write API budget. The source is opened only after the destination receipt check, so a replay does not touch the source. A conflict may add a destination freshness GET to identify an identical winning initializer; this verification stays on the failure path. The publisher captures manifest, CAS version and ref verification under the existing sync mutex, and advances the pristine predicate with the accepted batch: no extra bucket request or new CAS object. Existing push/read/checkpoint budgets are unchanged. Forwarding the negotiated SHA-256 object-format capability to upstream v2 upload-pack is entirely local and also adds zero bucket requests. Maintainer parallelism does not add a round trip or change any CAS'd object's write rate. It raises the instance's maintenance request rate to approximately `maintenance.workers / per-repository latency` while leaving each repository's checkpoint, compaction and lease protocol unchanged. Bucket GC adds no push/read critical-path requests and never scans `repos/`. Its one LIST is scoped to a repository already selected by `pending/`; log/checkpoint reads and conditional deletes are bounded to 32 concurrent requests. Only `gc.pb` is overwritten, once after the delete plan completes. Keep this table current; when you change a protocol, update the row and put the before/after depth in the commit message. The sim harness can enforce it: `FaultStore::stats().ops` counts exact store requests per link, so a scenario can assert "a push on a healthy link is ≤ N requests" as a regression test. ## 3. Rules of thumb - **Depth before count.** Two PUTs in parallel cost one round trip; the same two in sequence cost two. `tokio::join!` independent writes; never `await` uploads one after the other. - **Let the conditional write be the read.** `PutMode::Create` → 412 *is* "it exists"; `Update(v)` → 412 *is* "someone moved it" and may return the current version. Don't GET to decide what a conditional write will tell you for free. - **Verification goes on the failure path.** e.g. `put_immutable_create` HEADs only after a 412; `cas_landed` re-reads the manifest only after a non-412 error. The happy path must not pay for rare cases. - **Carry state in the object you already fetch.** The manifest is the one GET every request makes: anything a reader needs at refs level (pack set + side-file inventory, checkpoint pointer, log segments, revision, writer) belongs in it, so no second request is needed to know *what* to fetch. - **Re-use version tokens you hold.** After a CAS success you have the new generation: don't re-GET to learn it. After a 412 you may have `current`: skip the GET when the store provides it. - **Batch at the CAS.** The single-flight publisher (group commit, `wal.batch_window`) turns N concurrent pushes into one log PUT + one CAS. - **Never pay per ref or per pack on a hot path.** O(1) manifest, O(k) resolve; pack data by range, never by "download and see". - **Immutable means cache without revalidation while present**, on every layer (process LRU, bucket `cache/api/v1`, HTTP `immutable`). Bucket copies may expire by retention and be recomputed byte-identically. One instance computing something all instances need = a shared-cache write, not N recomputations. - **Jitter every retry** that targets a shared CAS'd object; a synchronized retry storm is a self-inflicted serialization on the 1 write/s object. - **Measure, don't guess**: `gitcask-traces` shows the store span tree per request (count + critical path); sim `Stats::ops` gives exact counts; the measurement table holds the numbers. ## 4. Anti-patterns seen (don't reintroduce) - HEAD/GET "to be safe" before a PUT that is conditional anyway. - Sequential `.await` on independent uploads. - Re-sync (GET) after a CAS success to "refresh". - LIST to find the latest checkpoint/segment (the manifest knows). - An undelimited LIST over `repos/` to enumerate repositories: it walks every object in the bucket (122 k keys, 8–9 s, store deadline retries, 2026-08-22). Enumerate *prefixes* (`list_prefixes`), probe manifests, cache. - Fixing a liveness bug by adding a wait or a probe to the *happy* path. - A retry loop without backoff+jitter on `manifest.pb`. ## 5. Checklist for a protocol change (paste into the PR/commit) - Depth/requests before → after for each affected row in §2. - What moved to the failure path, and how often that path runs (measured or reasoned). - Which CAS'd object's write rate changes. - Sim scenario(s) covering the new failure mode; `Stats::ops` budget assertion if the hot path changed. # gitcask cross-crate contract Context: **the original cross-crate interface contract (2026-08-18/19, written so eight owners could build the crates in parallel)**, kept as the reference for names and shapes. Rule still in force: *extend, do not rename* — a type or function listed here is relied on by another crate. **Where this file and the code disagree, the code is right and this file is stale**; verify with `rg`/`cargo doc` before relying on a signature. Known supersessions (2026-08-20 sweep): synchronization has refs-only and full-local levels (`sync_refs`, `sync_refs_only`, and `sync_full`, `AGENTS.md §2.3`); the server router is `AGENTS.md D15/D20/D26/D27`. Read when you touch a crate boundary; update the relevant block when you extend one. Shared interfaces between crates. Implement exactly these names/shapes; extend freely, do not rename. Original owners (parallel batch): StoreS3, StoreCoord, GitEngine, Wal, Server, Cli. Read `AGENTS.md` first (design §1–§2, decisions §3; the original layout/phases/config draft is the measurement log). ## Existing (do not rewrite; extend only) - `gitcask-proto`: prost types from `proto/gitcask/v1/wal.proto` (Manifest, LogSegmentRef, LogEntry, PackRef, RefTransaction/RefUpdate, Checkpoint(+Ref), RefSnapshot/Ref, Lease); `keys::*`; `frame::{encode_entries,decode_entries}` (uvarint-framed log encoding); `time::*`. - `gitcask-store`: `ObjectStore` trait (`Version` opaque CAS token, `ObjectMeta{key,size,version,last_modified}`, `GetOptions{if_none_match,if_match,range}`, `GetResult::{NotModified,Object}`, `PutMode::{Overwrite,Create,Update(Version)}`, `PutBody::{Bytes,Stream,File}`, `PutOptions`, `StoreError::{NotFound,PreconditionFailed{current},Retryable,InvalidArgument,Other}`, `ObjectStoreExt`, `Prefixed`, `memory::MemoryStore`, `util::{collect,once,file_stream,backoff,retry}`), modules `coord.rs`, `s3.rs`, and test-only `fault.rs` / `memory.rs`. - `gitcask-config`: `Config` for gitcask.toml (+ `GITCASK__` env overrides, `PORT`). `AuthMode::{None,Jwt,Introspect,Forwarded,IntrospectForwarded}`; `AuthConfig::{jwt,introspect}`. `IntrospectConfig` holds `url`, `secret_env`, `cache_ttl`, `negative_cache_ttl`, `timeout`; the secret value is never config data. `required_forward_secret() -> anyhow::Result` reads and validates the required proxy secret for combined mode without including its value in errors. `Config::validate` checks both auth configurations. Request selection and errors: [SECURITY.md](/llms/security.md#mixed-direct-and-trusted-proxy-authentication). ## gitcask-git (owner: GitEngine) ```rust pub struct RepoId { owner: String, name: String } // FromStr("owner/name" | "owner/name.git"), Display "owner/name". Validation: each part ASCII [A-Za-z0-9._-], // no leading '.', not "..", 1..=100 chars. fn owner(), name(), store_prefix() (gitcask_proto::keys::repo_prefix), // local_dir(root:&Path)->PathBuf (= root/owner/name.git). pub enum ObjectFormat { Sha1, Sha256 } // From, <-> gix_hash::Kind, as_str() /// Bare git repo on local disk in standard layout (objects/pack/*.{pack,idx}, loose refs + packed-refs, HEAD, /// config with repositoryformatversion / extensions.objectformat) readable by gix AND upstream git. /// Clone-able handle (Arc inside), thread-safe. pub struct LocalRepo; impl LocalRepo { pub fn init(root: &Path, id: &RepoId, format: ObjectFormat) -> Result; pub fn open(root: &Path, id: &RepoId) -> Result, GitError>; pub fn id(&self) -> &RepoId; pub fn path(&self) -> &Path; pub fn object_format(&self) -> ObjectFormat; pub fn gix(&self) -> gix::Repository; // per-thread handle from shared ThreadSafeRepository pub fn refresh(&self) -> Result<(), GitError>; // re-read odb/refs after pack/ref changes // ---- packs /// Read `pack` to EOF into an unlinked file under objects/pack/ (freed on drop), bounded by `max_bytes`. /// No ingest lock. receive-pack runs it alongside the sync; over HTTP/1.x it finishes before the /// response starts (D52). pub async fn spool_pack(&self, pack: R, max_bytes: Option) -> Result; pub struct SpooledPack { /* private */ } // .bytes() /// = git index-pack: write objects/pack/pack-.{pack,idx,rev}; thin packs resolved against /// the odb (--fix-thin); verify checksum; opts.fsck => object-level validation. Empty pack => Ok(None). /// opts.max_bytes is ignored: spool_pack enforced it (receive-pack passes None). pub async fn ingest_spooled(&self, pack: SpooledPack, opts: IngestOptions) -> Result, GitError>; /// Stream either all source refs (CLI) or pinned OID closure (HTTP) through normal ingestion. pub async fn import_pack_from(&self, source: &Path, tips: Option<&[String]>, opts: IngestOptions) -> Result, GitError>; /// spool_pack(pack, opts.max_bytes) then ingest_spooled, for local streams (import, API writes). pub async fn ingest_pack(&self, pack: R, opts: IngestOptions) -> Result, GitError>; pub struct IngestOptions { pub fsck: bool, pub max_bytes: Option, pub thin: bool } pub struct IngestedPack { pub checksum: gix_hash::ObjectId, pub pack_path: PathBuf, pub idx_path: PathBuf, pub pack_size: u64, pub idx_size: u64, pub object_count: u64 } /// Atomically move downloaded files into objects/pack/ (rename), then refresh. pub async fn install_pack(&self, pack: &Path, idx: &Path, extra: &[PathBuf]) -> Result<(), GitError>; /// Delete .pack/.idx/.rev/.bitmap. Caller guarantees no readers (wal holds a lock). pub fn remove_pack(&self, checksum: &gix_hash::oid) -> Result<(), GitError>; pub fn packs(&self) -> Result, GitError>; pub struct PackInfo { pub checksum: gix_hash::ObjectId, pub pack_size: u64, pub idx_size: u64, pub object_count: u64, pub has_rev: bool, pub has_bitmap: bool } pub fn pack_path(&self, checksum: &gix_hash::oid) -> PathBuf; // objects/pack/pack-.pack (idx: set_extension) // ---- refs /// All refs sorted by name incl. peeled tags + HEAD symbolic target. `From` both ways with /// gitcask_proto::v1::RefSnapshot. pub fn refs(&self) -> Result; /// Atomic all-or-nothing. check_old => verify old_oid (zero = must not exist). Supports HEAD symbolic update. /// Error GitError::RefConflict{name, expected, actual}. pub fn apply_ref_txn(&self, txn: &gitcask_proto::v1::RefTransaction, check_old: bool) -> Result<(), GitError>; /// Replace ALL refs + HEAD (write packed-refs directly; must be fast for 500k refs). pub fn load_ref_snapshot(&self, snap: &gitcask_proto::v1::RefSnapshot) -> Result<(), GitError>; pub fn pack_refs(&self) -> Result<(), GitError>; // ---- objects pub fn has_object(&self, oid: &gix_hash::oid) -> bool; /// Every object reachable from tips exists. `stop_at_existing_refs` => stop at objects reachable from /// current refs (rev-list --objects --not --all). Error GitError::MissingObject{oid}. pub fn check_connectivity(&self, tips: &[gix_hash::ObjectId], stop_at_existing_refs: bool) -> Result<(), GitError>; // ---- protocol, server side /// Raw passthrough: spawns `git upload-pack --stateless-rpc` with `GIT_PROTOCOL` set. pub async fn upload_pack_raw(&self, protocol: Protocol, body: R, out: W) -> Result<(), GitError>; /// v2 ls-refs from the ref snapshot; efficient prefix filtering. pub fn ls_refs(&self, args: &LsRefsArgs) -> Result, GitError>; pub struct LsRefsArgs { pub ref_prefixes: Vec, pub symrefs: bool, pub peel: bool, pub unborn: bool } /// v0 advertisement with capabilities. pub fn advertise_refs_v0(&self, service: Service, out: &mut Vec) -> Result<(), GitError>; pub enum Service { UploadPack, ReceivePack } // FromStr("git-upload-pack"|"git-receive-pack") // ---- upstream git helpers pub async fn git(&self, args: &[&str]) -> Result; // cwd=repo, GIT_DIR set pub async fn repack(&self, opts: RepackOptions) -> Result; pub struct RepackOptions { pub mode: RepackMode /* Geometric{factor} | Full */, pub write_bitmap: bool, pub write_midx: bool, pub keep: Vec } pub struct RepackResult { pub new_packs: Vec, pub removed: Vec } } pub mod pkt; // pkt-line read/write, flush/delim/response-end, sideband encode; Protocol::{V0,V2} from // GIT_PROTOCOL header; command/arg parsing for v2 (ls-refs, fetch, object-info) pub mod receive; // parse receive-pack request: caps + commands ("old new refname\0caps"), push-options, // => (gitcask_proto::v1::RefTransaction, ReceiveCaps{report_status_v2, side_band_64k, // atomic, quiet, push_options, agent, object_format}); pack bytes follow in the same body. // `report_status(caps, unpack: Result, per_ref: &[(name, Result<(),String>)], out)` writer // producing report-status(-v2), sideband-framed when requested. pub enum GitError { Io, Gix(Box), Pack, RefConflict{name,expected,actual}, MissingObject{oid}, Fsck(String), Subprocess{cmd,status,stderr}, InvalidInput(String), Protocol(String) } ``` `build_v2_fetch_request(req: &UploadPackRequest, format: ObjectFormat)` encodes the explicit object-format protocol capability before the v2 delimiter, then the fetch arguments. The server supplies its synchronized repository's format. ## gitcask-store::coord (owner: StoreCoord) ```rust /// Generic read-modify-write CAS loop on a protobuf object. `f(None)` when absent. Returning `None` from `f` /// aborts with Ok(None). Retries on PreconditionFailed (re-reading) up to `max_retries`, on Retryable with /// backoff. Returns the written meta + value. pub async fn cas_update(store: &dyn ObjectStore, key: &str, max_retries: u32, f: F) -> Result, CoordError> where F: FnMut(Option<&T>) -> Result, CoordError>; /// Read a protobuf object with its version. Ok(None) if absent. pub async fn get_message(store: &dyn ObjectStore, key: &str) -> Result, CoordError>; pub async fn get_message_if_changed(store, key, known: &Version) -> Result, CoordError>; /// Lease = gitcask_proto::v1::Lease at `key`, acquired by Create or by Update over an expired lease. pub struct LeaseGuard; // holds store handle, key, holder id, current Version; Drop => best-effort release impl LeaseGuard { pub async fn heartbeat(&mut self, ttl: Duration) -> Result<(), CoordError>; // CAS-extend expires_at pub async fn release(self) -> Result<(), CoordError>; // CAS delete pub fn spawn_heartbeat(self: Arc>, every: Duration, ttl: Duration) -> tokio::task::JoinHandle<()>; pub fn holder(&self) -> &str; pub fn expires_at(&self) -> SystemTime; } pub async fn try_acquire(store: DynStore, key: &str, holder: &str, purpose: &str, ttl: Duration) -> Result, CoordError>; // None = held by someone else and not expired pub async fn acquire(store, key, holder, purpose, ttl, wait_up_to: Duration) -> Result, CoordError>; pub fn instance_id() -> &'static str; // explicit instance name/id, hostname+pid, or uuid; computed once pub enum CoordError { Store(StoreError), Decode(prost::DecodeError), Aborted, RetriesExhausted{key, attempts}, Other } ``` ## gitcask-store backend (owner: StoreS3) ```rust // s3.rs pub struct S3Store; impl S3Store { pub async fn new(cfg: &gitcask_config::StoreConfig) -> anyhow::Result; } // lib.rs pub async fn open_store(cfg: &gitcask_config::Config) -> anyhow::Result; // by cfg.store.backend, applies Prefixed(cfg.store_prefix()) ``` Contract tests: `crates/gitcask-store/tests/contract.rs` with a `run_contract(store: DynStore)` suite executed for memory always, and for S3 when `GITCASK_TEST_S3_ENDPOINT` is set (bucket `GITCASK_TEST_BUCKET`, default "gitcask-test"). The memory store and fault wrapper are exposed only with the `testing` feature. ## gitcask-wal (owner: Wal) ```rust pub struct Registry; // one per process: DynStore + Arc + cache_root; DashMap> impl Registry { pub fn new(store: DynStore, cfg: Arc) -> Arc; /// Open existing (materialize local copy lazily). Err(WalError::NotFound) if manifest.pb absent. pub async fn open(&self, id: &RepoId) -> Result, WalError>; /// CAS-create manifest.pb (PutMode::Create). Err(WalError::AlreadyExists). pub async fn create(&self, id: &RepoId, format: ObjectFormat) -> Result, WalError>; pub async fn open_or_create(&self, id: &RepoId, format: ObjectFormat) -> Result, WalError>; /// Repositories materialized by this process (cache/event sweep inspection only). pub fn cached_repos(&self) -> Vec; pub fn store(&self) -> &DynStore; pub fn config(&self) -> &Arc; /// Disk cache maintenance: evict idle repos and relieve disk pressure. pub async fn evict_idle(&self) -> Result; } pub struct RepoHandle; impl RepoHandle { pub fn id(&self) -> &RepoId; pub fn local(&self) -> &LocalRepo; pub fn store(&self) -> &Prefixed; // repo-scoped pub fn manifest(&self) -> Arc; // last known pub fn manifest_version(&self) -> Option; /// Freshness check (conditional GET on manifest.pb; honors wal.freshness_ttl) + catch-up (download new /// packs, apply log entries after our seq, apply COMPACT: install new pack, remove superseded). Returns a /// read guard; while any guard is alive no pack is removed locally. Every request calls this first. pub async fn sync_full(&self) -> Result, WalError>; pub async fn sync_refs(&self) -> Result, WalError>; pub async fn sync_refs_only(&self) -> Result, WalError>; pub async fn try_compaction_lease(&self) -> Result, WalError>; pub async fn try_checkpoint_lease(&self) -> Result, WalError>; /// Force full re-materialize from store (repair). pub async fn rematerialize(&self) -> Result<(), WalError>; /// Publish a push. `pack` was produced by LocalRepo::ingest_pack on this handle's local repo (already on /// disk). Steps: upload pack+idx to wal/.{pack,idx} (skip if exists) ‖ verify txn old values against /// synced refs; then CAS: append LogEntry to log (new segment object per batch on regional buckets), /// cas_update manifest (head_seq+1, packs+=, log_segments+=); on PreconditionFailed: re-sync, re-verify /// old values (conflict per ref → whole push rejected unless !atomic and per-ref reporting), retry. /// Then apply refs locally. Coalesces concurrent publishes on this handle (wal.batch_window/max_batch). pub async fn publish_push(&self, pack: Option, txn: RefTransaction, meta: HashMap) -> Result; pub struct PublishResult { pub seq: u64, pub per_ref: Vec<(String, Result<(), RefError>)> } pub async fn publish_ref_update(&self, txn: RefTransaction, meta) -> Result; /// Same normal publisher, with a pristine-state precondition on every CAS /// attempt and earlier accepted requests in the batch. Assigns receipt.seq /// and records the receipt in the manifest CAS alongside branch + HEAD. pub async fn publish_initialization(&self, pack: IngestedPack, receipt: Initialization, meta: HashMap) -> Result; /// COMPACT entry: new pack (already local, e.g. from LocalRepo::repack) superseding `supersedes`. pub async fn publish_compact(&self, new_pack: PackInfo, supersedes: Vec, tier: u32) -> Result; /// Write checkpoint at current head (refs snapshot + pack set), then CAS manifest (checkpoint=, min_seq=, /// log_segments trimmed). Idempotent. pub async fn write_checkpoint(&self) -> Result; /// Read log entries [from_seq, to_seq] from the store (provenance/rewind tooling). pub async fn read_log(&self, from_seq: u64, to_seq: Option) -> Result, WalError>; pub fn last_access(&self) -> Instant; pub fn touch(&self); } pub enum WalError { NotFound, AlreadyExists, Store(StoreError), Coord(CoordError), Git(GitError), Publish{msg,retryable}, Corrupt(String), Retry{attempts}, Io(std::io::Error) } // WalError::is_retryable() preserves transient store failures across publish batching and crate boundaries. pub enum RefError { NonFastForward, Conflict{expected,actual}, Rejected(String), Missing } ``` `TaskRecord`, `TaskOutcome`, and `Progress` implement both `Serialize` and `ToSchema`; the server's task JSON and OpenAPI components are generated from the same cross-crate types. `is_pristine(&Manifest)` checks zero head sequence, no packs/logs/checkpoint/receipt. `on_bulk_runtime(worker_threads, future)` exposes the existing bounded bulk runtime for server-side initialization work; it keeps running if the caller disconnects. `Manifest.initialization: Option` is an append-only protobuf field carrying one bounded operation key, canonical request hash and original result. Every subsequent manifest rewrite must preserve it; see [INITIALIZE.md](/llms/initialize.md). ## gitcask-server (owner: Server) ```rust pub struct AppState { pub cfg: Arc, pub store: DynStore, pub registry: Arc, pub auth: Arc } pub fn router(state: Arc) -> axum::Router; pub struct auth::Authenticator; impl auth::Authenticator { pub async fn authenticate(&self, headers: &HeaderMap) -> Result; pub async fn require_read(&self, headers: &HeaderMap, owner: &str, repo: &str) -> Result; pub async fn require_write(&self, headers: &HeaderMap, owner: &str, repo: &str) -> Result; pub async fn require_admin(&self, headers: &HeaderMap, owner: &str, repo: &str) -> Result; } pub fn auth::mint_token(private_key_pem: &str, issuer: &str, audience: Option<&str>, principal: &str, scopes: &[String], ttl: Duration) -> anyhow::Result; pub fn auth::generate_key_pair_pem() -> anyhow::Result<(String, String)>; /// Bind, serve (HTTP/1.1 + h2c), graceful shutdown on SIGTERM/SIGINT/`shutdown` future. pub async fn serve(state: Arc, shutdown: impl Future + Send) -> anyhow::Result<()>; // gc::due/collect: lowest-priority per-repo maintenance; one repo-prefix LIST, // manifest-generation revalidation, then version-conditional deletes under leases/gc.pb. // Routes (all under /{owner}/{repo}[.git]): // GET /info/refs?service=git-upload-pack|git-receive-pack (v0 advert or v2 capability advert per Git-Protocol) // POST /git-upload-pack POST /git-receive-pack (Content-Encoding: gzip supported; streaming both ways) // GET /HEAD GET /objects/info/packs (404 unless dumb enabled) // POST /info/lfs/objects/batch PUT/GET /info/lfs/objects/{oid} POST /info/lfs/verify // PUT / (create repo, write permission) DELETE / (admin permission) // Non-repo: GET /healthz /readyz /metrics /docs /openapi.json // GET /api/v1/docs and /api/v1/openapi.json permanently redirect to /docs and /openapi.json. // server.public_docs = true (default) opens all four docs paths; false gates all four. // /metrics is always authenticated. Exact one-segment docs routes leave /docs/{repo} available. // Auth: EdDSA JWT from Git Basic password / API Bearer is verified against public PEM or cached JWKS and // repository scopes are checked by require_read/write/admin. Introspection uses the same scopes; // IntrospectForwarded selects one scheme per request. Permission checks follow the resolved principal // (scoped token or forwarded grants), never the configured mode alone. Endpoints // synchronize refs only or the complete local pack set (AGENTS.md §2.3). ``` ## Pristine import extensions (2026-10-05) - `gitcask-config::Config.import: ImportConfig` bounds refs, acquired objects, transferred/output pack bytes, resolve timeout and acquisition timeout; [IMPORT.md](/llms/import.md) owns the defaults and wire contract. - `gitcask-git::RefSnapshotData.head_oid` carries detached HEAD only. Symbolic HEAD uses `head_target` and resolves its OID from refs. Both conversions to/from protobuf and offline replay preserve this distinction. - `gitcask-git::LocalRepo::import_pack_from_supervised` uses normal bounded spool and idx/rev/pack adoption with async cancellable index-pack; import-owned process scopes retain acquisition/scratch through bounded termination verification before timeout unwind. - `gitcask-git::isolated_command(path)` clears inherited Git environment/config/helpers/hooks/protocols for scratch Git acquisition. The shared `LocalRepo::import_pack_from` feeds normal index-pack ingestion. - `gitcask-wal::RepoHandle::refs_snapshot()` captures refs/HEAD under the existing publisher sync lock after freshness sync, with no additional bucket request. - `gitcask-wal::RepoHandle::revalidate_refs()` always performs refs-level freshness checking even with optional read TTLs, for committed-result reconciliation. - `gitcask-wal::RepoHandle::publish_import(pack, txn, ImportReceipt, meta)` adds the pinned transaction and manifest receipt in the existing publisher CAS, with pristine validation at every batch/CAS retry. `Manifest.import_receipt` and `RefSnapshot.head_oid` are append-only protobuf extensions. - Server resolve/import/target-read receipt routes use the existing API lanes, principal checks, SSE/tasks and bulk runtime. Cloud owns durable orchestration, not the Registry or ObjectStore. ## gitcask-cli (owner: Cli) `gitcask --config gitcask.toml `: `serve` | `compact owner/name [--once]` | `repo create|info` | `wal pending|ls|show|materialize --at-seq` | `synth --out DIR --size s|m|l [--commits N --files M]` | `import --from GITDIR owner/name` | `token keygen|mint`. Also `Dockerfile`, `compose.yaml` (rustfs + gitcask), `justfile`, `gitcask.example.toml`, `tests/e2e.sh` (real git vs. server on memory store and on rustfs). # gitcask — direction and decisions This document fixes what gitcask is, where it differs from its origin (walgit), and which decisions are settled. How far the *product* goes (core vs platform, OSS vs cloud) is decided by `docs/PRODUCT.md`. Read this before touching code — human or agent. `AGENTS.md` has been rewritten for this fork and the two documents agree; if a conflict ever appears, **this document wins** and the other one gets fixed. gitcask was renamed from repohub on 2026-09-01 before publication because the old name collided with an existing project and did not express its Bitcask lineage. ## 1. What gitcask is **A stateless git storage and serving layer with an object store (S3) as the source of truth.** A repository is a WAL in S3; a server is a cache that may be wiped at any time. - git itself is not modified. `git upload-pack` / `index-pack` / `repack` and the rest of the plumbing are called as-is; gitcask decides only *where the bytes live* (packs in S3 under `wal/`, the current pack set in `manifest.pb` behind a CAS). - The server's local disk is a cache. Idle repositories are evicted (`cache.evict_idle_after`) and a container that restarts empty rebuilds everything from S3. No persistent volume. - Provenance: Cursor's *Git at any scale* (`docs/reference/cursor-git-at-any-scale.md`), reproduced by Tobi Lütke as walgit (2026-08-23), taken and renamed. History and the license file were not imported. ## 2. The workload — the opposite of walgit's | | walgit (origin) | gitcask (us) | |---|---|---| | Goal | one company's 57 GiB monorepo served from a machine smaller than it | SaaS: thousands–tens of thousands of unrelated users × ~20 **small** repositories each | | Repo size | tens of GB; one pack larger than the disk | a few MB – a few hundred MB; packs always fit on disk | | Repo count | dozens, placed by hand in a config file | created automatically on sign-up; unbounded | | Concurrent push | hundreds of developers into one repository | one user per repository, plus that user's agent | | Auth | company IdP login, global write/admin flags | platform-owned tokens; gitcask verifies JWTs or introspects opaque tokens | | Metadata | none (replaced by S3 enumeration) | **comwit's RDB** (repository list, owner, last_push_at) | Because of this inversion, everything walgit built for *one big repository* is removed, and everything that scales with *repository count* is reworked. The split has held since the fork. As of 2026-09-13 upstream's stated goal is "fast for monorepos": its September work replaced bundle-uri with native packfile-uri delivery of proven static packs, verified multi-pack bitmaps and pack-retirement proofs, while its maintainer still lists every repository in the bucket each pass. Nothing there moved toward the many-small-repositories workload. ## 3. Settled decisions 1. **Identity and the permission model live outside gitcask.** The platform judges sessions/permissions and owns tokens; gitcask verifies EdDSA JWTs with a public key/JWKS or introspects opaque tokens (AGENTS D47), then applies the same repository scopes. There is no login, no user DB, no sessions, no token issuance, no OIDC, no TLS. 2. **The repository list and metadata live in comwit's RDB.** The enumerate-every-repository APIs (`registry.list()` + S3 HEAD) are gone; the maintainer takes its work from push-driven markers, never from an S3 scan. 3. **comwit creates repositories.** `auto_create_on_push` is off. comwit inserts the RDB row, then calls `PUT /{owner}/{repo}`. 4. **comwit reads user files via `git clone`** (option A). The repository-browsing JSON API (`/{o}/{r}/api/tree|blob|commits…`) is not used yet but is **kept for the next stage**. 5. **Rust stays; this is a fork, not a rewrite.** Upstream (tobi/walgit) is not tracked as git history. Its commits are reviewed about monthly for correctness fixes that also apply here; those are re-implemented against this tree (crate names, receive path and pack lifecycle have diverged too far for cherry-picks) and credited by upstream sha in the commit message. 6. **The cache disk is node-local SSD (emptyDir).** No tmpfs, no FUSE mounts, no remote pack serving. Packs are always downloaded whole. 7. **LFS is on** (D5). Big files go through LFS; the object cap is 1 GiB. ## 4. Keep / remove, by feature ### Kept (the core) - Smart HTTP push/fetch (`smart.rs`, `gitcask-git/receive.rs`, calling `git upload-pack`) - WAL publish (CAS + group commit), sync, checkpoints, compaction, evict (`gitcask-wal`) - The S3 backend (`gitcask-store`) - The maintainer loop (`maintain.rs`) — reworked to be marker-driven - rev-index / fsck, drain - EdDSA JWT verification or opaque-token introspection + repository scopes, plus `forwarded` mode for existing proxies - Repository create/delete (`admin.rs`: `PUT/DELETE /{o}/{r}`) - Operational status API + SSE (`/{o}/{r}/api/overview|ops|tasks`, `sse.rs`) - The repository browsing API (`web/api/`) - The webhook bridge (`events.rs`, `bridge.rs`, `[events]`), `/_events/notify` - health / metrics / telemetry - LFS + `static_object.rs` (on, D5) - CLI: `serve`, `repo *`, `wal *`, `compact`, `import`, `migrate`, `synth`, `token` ### Removed | Group | What | Why | |---|---|---| | Auth | OIDC/browser login (`web/login.rs`), token **issuance endpoints**, install scripts (`setup.rs`, `/services/public/*`, `setup.json`), built-in TLS (`tls.rs`) | identity and issuance belong to the platform; gitcask only verifies credentials | | Bundles | the `gitcask-bundle` crate, `bundles.rs`, CLI `bundle`, `[bundles]`, bundle-uri advertisement | small repositories clone instantly | | Big-repo machinery | in-process upload-pack (`upload_gix.rs`), the remote reader (`remote.rs`), `store_mount`, prewarm, tier-2 base/history packs, `cache.mode = budget` | packs always fit locally | | Placement & forwarding | `[placement]`, the push broker (`forward.rs`, `push_broker_*`), upstream follow/mirror (`follow.rs`, `mirror.rs`, `[upstream]`) | proxy routing replaces placement; nothing follows external repositories | | Hosting policy | push policy (`policy.rs`, `docs/POLICY.md`), per-repo settings (`settings.rs`), the all-owners listing API, `/services/api/instance`, CLI `config|policy|settings` | one user per repository; one global config suffices | The removals were carried out as the scoped tasks listed in §5. ## 5. Progress | Task | Status | |---|---| | 01 auth → forwarded principal | ✅ merged | | 02 remove bundles | ✅ merged | | 03 remove big-repo paths (remote reader, upload_gix, store_mount, prewarm, tier-2) | ✅ merged | | 04 remove placement / push broker / upstream follow | ✅ merged | | 05 remove policy / settings / listing APIs | ✅ merged | | 06 pending-marker maintainer (D1, AGENTS D40) | ✅ merged | | 07 S3 transient-error retry (AGENTS D42) | ✅ merged | | 08 1,000-repository load spike script (`scripts/spike.sh`) | ✅ merged — 118 store requests over 5 idle minutes; identical at n=100/1000 | | 09 cache-eviction loop (never called since the origin) | ✅ merged — cache empty after 30 s idle, cold clone p50 134 ms | | 10 transient store errors → 503 + Retry-After on every route | ✅ merged — 503 while rustfs is down, 201 after recovery | | 11 parallel maintainer workers (default 8), oldest markers first | ✅ merged — 1→8→32 workers: 22 / 111 / 154 repos/s warm. 50k spike v2: all 7 stages OK, idle 120 vs 118 | | 12 maintainer heartbeat cleanup | ✅ merged | | 13 fsck only after compaction / periodically (no full fsck per new repository) | ✅ merged — cold drain still 13/s, so not the bottleneck | | 14 cold-open concurrency (blocking fs/git waits pinned async threads) | ✅ merged — 2 threads: 32 concurrent cold opens 1,077→195 ms, cold drain 13→150 repos/s | | 15 compare API (`…/api/compare/{base}...{head}`: merge-base semantics, commit list, file stats, 2 MiB patch cap) | ✅ merged | | 16 correctness batch (committed pushes always answered; push 503s; discovery auth; defaults S3/auto-create off) | ✅ merged | | 17 bucket GC (superseded packs, folded logs, old checkpoints; retention window + `materialize --at-seq` contract; converges in 2–4 s) | ✅ merged | | 18 events backstop (pending ∪ cached) + global 10k task-record cap + evict on spawn_blocking | ✅ merged — known limit: if the marker PUT and the bucket notification are lost together, the sweep misses that push (webhook delayed only; data unaffected) | | 19 OpenAPI generation (utoipa) + Scalar docs — `/api/v1/openapi.json`, `/api/v1/docs`, both behind auth | ✅ merged | | 20 routing unification (hand-rolled dispatch removed; all axum) | ✅ merged | | 21 gitcask-git 2,958 lines → 11 modules (public paths unchanged) | ✅ merged | | 22 web/api.rs split (mod/view/git/handlers) + error unification (six `auth_err` copies and pktline duplication gone; git failures classed 400/404/409/500) | ✅ merged | | 23 publish commit-path extraction (389→149 lines) + compaction/checkpoint leases into RepoHandle | ✅ merged | | 24 remove `import --direct` (1,304 lines) — the second manifest-CAS implementation dies; import regression test added | ✅ merged | | 25 remove the GCS backend, dead config and unused deps; fault/memory behind the `testing` feature (−3,212 lines) | ✅ merged | | 26 stateless auth gate (mint/static, repo scopes, Basic/Bearer, streaming proxy) | ✅ merged, **superseded by 34** | | 27 bulk Gitea migration (resumable, LFS, `docs/MIGRATION.md`) | ✅ merged | | 28 write API, part 1 — branch/tag CRUD, annotated tags, archive | ✅ merged | | 29 shared-cache retention — GC expires `cache/api` and `cache/archive` (zero extra LISTs) | ✅ merged | | 31 sim flakiness root-caused — shared temp-file race in local state persistence fixed; sim 20/20 | ✅ merged | | 32 write API, part 2 — batch file commits; policy-free merge/squash/ff-only | ✅ merged | | 33 publication readiness — five-minute path, SECURITY/CONTRIBUTING/CoC, CI gate fixed | ✅ merged | | 34 verify JWTs inside gitcask, drop the gate | ✅ merged — EdDSA public key/JWKS, scopes, Basic/Bearer, offline token CLI, one process | | 35 operations runbook (`docs/OPERATIONS.md`) — verified metrics table, symptom-first diagnosis, recovery | ✅ merged | | 36 rename repohub → gitcask across crates, config, metrics, headers and docs; relicense MIT → Apache-2.0 with NOTICE | ✅ merged (2026-09-01, public release) | | 37 port upstream fixes: verify tips on empty-pack pushes (walgit d5e75caf); un-blind `just warnings` under forced colour (walgit b81b15ae) | ✅ merged (PR #2) | | 38 `auth_mode = "introspect"` — opaque tokens verified by RFC 7662 token introspection, bounded in-memory cache, 503 on introspection outage | ✅ merged (PR #4) | | 39 public API docs — `/docs` and `/openapi.json` open by default with `server.public_docs`; old docs paths permanently redirect | ✅ implemented | | local smoke (`scripts/smoke.sh`, rustfs) | ✅ 63/63 — includes introspect phase 4; re-run on every merge | | `AGENTS.md` / `GOAL.md` / `README.md` rewrites | ✅ | Remaining (outside gitcask, or later): - comwit integration: call `PUT/DELETE /{owner}/{repo}` on project create/delete, issue user git tokens in the chosen JWT/introspection format, consume the `[events]` webhook (D6, AGENTS D47) - (D10) putting the browsing API to use ## 6. Vocabulary - **WAL**: `manifest.pb` + `wal/.pack` + `log/.pb` under `repos///`. One push = one seq. - **manifest CAS**: swapping `manifest.pb` with S3 `If-Match`. The only consensus point. ~1 write/s ceiling. - **group commit**: folding same-repository pushes arriving within `wal.batch_window` into one CAS. - **materialize**: building a local bare repo from S3 packs and refs. `wal materialize --at-seq N` restores a past point in time. - **compaction**: `git repack`-ing several push packs (tier 0) into one (tier 1), uploading it, and removing the old ones from the manifest — the LSM SSTable merge, for git. - **evict**: deleting an idle repository's local cache directory. Not data loss. ## 7. Operating decisions These are comwit's actual operating choices. In particular D2 (size limits), D7 (new projects only), D8 (AWS S3) and D11 (2 vCPU) are values comwit picked for itself, not defaults gitcask imposes on any other deployment. | # | Decision | Content | |---|---|---| | D1 | maintainer candidate selection | S3 `pending//` markers: written on push success, consumed by `LIST pending/`, deleted when done. No RDB involvement | | D2 | size limits | 100 MB per file, 2 GB per push (GitHub's numbers). No whole-repo cap; monitoring only | | D3 | push rate limiting | at the proxy. gitcask does none | | D4 | repository deletion | hard delete (`DELETE /{o}/{r}` deletes from S3 immediately). Any recovery policy is comwit's | | D5 | LFS | on. `lfs.max_object_bytes = 1GiB` | | D6 | user git auth | Originally EdDSA scoped JWTs over HTTPS Basic (username ignored + token), verified with public key/JWKS. AGENTS D47 extends this to platform-owned opaque tokens via introspection | | D7 | existing data | new projects go to gitcask; existing ones convert on access | | D8 | storage | AWS S3 | | D9 | cache eviction | `cache.evict_idle_after = "2h"` | | D10 | using the browsing API (option B) | after 06 (the maintainer rework) | | D11 | instance size | **2 vCPU**, fixed for years. No optimisation raises thread counts; throughput comes from more instances. Runtime defaults (bulk threads etc.) target 2. Benches and spikes stay small (300–1,000 repositories) — they exist to show trends | | D12 | storage backends | **S3 only.** The GCS backend was removed 2026-08-31 (1.7k unused lines + a separate dependency tree). If ever needed, recover it from git history | | D13 | product boundary and the gate | (2026-09-01, **superseded by D14**) authentication lived in a separate gate process in the OSS core. Retired by task 34 to remove the trusted headers, the duplicated path-to-permission table and the second process | | D14 | product boundary and JWT | (2026-09-01) the OSS core = auth · git transport · read/write API. gitcask verifies EdDSA JWT signatures and repository scopes itself but **owns no identity**; issuance belongs to the platform or the offline CLI. The cloud = multi-tenancy · billing · operations. CI, issues, PRs, UI and repository listing are out of scope | Both platform backends and users' git CLIs send tokens to the same single gitcask process. Deployments with their own IdP proxy can use `forwarded`, or explicitly opt into `introspect_forwarded` for mixed direct-token and trusted-proxy traffic on that listener (AGENTS D48; [security contract](/llms/security.md#mixed-direct-and-trusted-proxy-authentication)). # Security policy ## Reporting a vulnerability Do not open a public issue for a suspected vulnerability. Use GitHub's [private vulnerability reporting form](https://github.com/burrr-ai/comwit-gitcask/security/advisories/new) so maintainers can investigate without exposing users before a fix is available. Include the affected commit or version, deployment shape, impact, reproduction steps, and any suggested mitigation. Remove live credentials, repository contents, and other third-party data from the report. If the problem is actively being exploited, say so in the title. Maintainers will: - acknowledge the report within three business days; - provide an initial assessment or request missing information within seven calendar days; - send at least weekly status updates while a confirmed issue remains unresolved; and - coordinate a fix, release notes, and public disclosure with the reporter when practical. Timelines for a fix depend on severity and complexity. Please allow a reasonable remediation window before publishing details. ## Supported versions gitcask is pre-1.0. Security fixes are made on `main`; old snapshots and unreleased branches are not maintained. ## Mixed direct and trusted-proxy authentication `server.auth_mode = "introspect_forwarded"` explicitly enables issuer introspection and trusted forwarded authentication on the **same process, port and routes** (AGENTS D48). Startup requires valid `[auth.introspect]` configuration, its service secret, and `GITCASK_FORWARD_SECRET`. Both secrets must be non-empty printable ASCII without whitespace. Secrets are read at startup; restart to rotate them. The issuer service secret and proxy secret have separate purposes; provision separate values. For every protected request, selection happens before permissions are checked: | `X-Gitcask-Forward-Secret` | Authentication and precedence | |---|---| | Absent | Introspect the Basic password or Bearer token using the existing bounded client/cache. Ignore all forwarded identity/write/admin headers. | | Present but empty, malformed, repeated, or wrong | Return 401; never retry with `Authorization`, even if it contains a valid token. | | Present once and valid | Compare in constant time, then require a non-empty `X-Gitcask-Principal`. Use only forwarded identity/grants; ignore `Authorization`, even if valid, invalid or more privileged. A missing/empty principal returns 401. | Header values have surrounding HTTP whitespace trimmed before comparison, as in standalone `forwarded`. Identities and privileges are **never combined**. Introspected repository scopes retain the JWT grammar and hierarchy (`admin` implies `write` implies `read`); a scope miss returns 404. A trusted forwarded principal can read, while `X-Gitcask-Write: 1` and `X-Gitcask-Admin: 1` independently grant write and admin; a missing grant returns 403. Forwarded grants are not repository scopes: the proxy must authorize the requested repository and operation before supplying them. Missing/invalid user credentials or invalid proxy credentials return 401 with `WWW-Authenticate: Basic realm="gitcask"`. Introspection service failures return 503 with `Retry-After: 5` and no authentication challenge; valid forwarded requests continue working during an issuer outage. `/healthz` and `/readyz` remain open. With `server.public_docs = true` (default), `GET /docs`, `GET /openapi.json`, and the permanent redirects from their `/api/v1/` paths are also open. Set `server.public_docs = false` to require authentication for all four. The OpenAPI schema is built from source annotations and contains no runtime secrets, bucket names, or internal service addresses. `/metrics` remains authenticated. Authentication adds no bucket requests. The authenticating proxy **must strip every client-supplied `X-Gitcask-Principal`, `X-Gitcask-Write`, `X-Gitcask-Admin` and `X-Gitcask-Forward-Secret` header**, including duplicates, before injecting its own identity, grants and secret. Never pass a client's secret header through. Protect the proxy-to-gitcask hop with TLS or a private trusted transport, and keep the shared secret out of client-visible responses and logs. Possession of this secret is the request's trust boundary; a route, source address or `X-Gitcask-Capabilities` header does not establish identity. Standalone `none`, `jwt`, `introspect` and `forwarded` keep their existing contracts (AGENTS §1.3). In particular, standalone `forwarded` still permits deployments without a configured secret and must then be reachable only through a trusted proxy. The combined mode never permits that configuration. # Initialize a repository from a pinned tree `POST /{owner}/{repo}/api/initialize` initializes an **existing, pristine** destination from one explicitly pinned commit in another Gitcask repository. It creates one new root commit, with no parent and none of the source commit history. Every required Git object is uploaded into the destination's normal `wal/` pack storage before the normal manifest CAS publishes it. Deleting the source or either instance's cache does not affect the destination. ```json { "source": {"owner": "templates", "repo": "starter", "commit_oid": "0123456789012345678901234567890123456789"}, "branch": "main", "operation_key": "project-creation-7", "message": "Initial project", "committer": {"name": "Project Builder", "email": "builder@example.test", "when": "2026-09-28T00:00:00Z"} } ``` `source.commit_oid` must be a full, nonzero commit object ID (40 hex digits for SHA-1, 64 for SHA-256), not a branch, tag, expression, or abbreviated ID. Both repositories must use the same object format. The source commit need not remain a branch tip, but its objects must still be available in the source's live packs. Source replacement refs are ignored: the pinned object itself determines the tree. `branch` is below `refs/heads/`; initialization also sets symbolic `HEAD` to it. An optional `author` has the same shape as `committer` and defaults to it. Identity timestamps must include an explicit RFC 3339 offset. No server time is substituted into the commit. The caller supplies the message and both identities; these are commit content, not authorization identities. `operation_key` is 1–128 printable ASCII characters without whitespace. The body is limited to 64 KiB. Unknown request fields are rejected. The request fingerprint uses parsed fields in fixed order, lower-case commit IDs, and the effective author; JSON whitespace/key order and an omitted author equal to the committer do not change it. Other field changes, including timestamp spelling, change the fingerprint. First success returns **201**; an exact retry returns **200**: ```json {"ref":"refs/heads/main","commit_oid":"","tree_oid":"","seq":1,"replayed":false} ``` The replay returns the original ref, commit, tree and WAL sequence with `replayed:true`. It does not assert that the ref still has that value and never moves a ref. Persist and reuse the entire original request after a lost response. The sequence can exceed 1 if a failed publisher left an orphan log slot. ## Authorization and errors Each attempt, including a replay, requires destination write permission **and** source read permission under the existing authentication modes. Both permissions are checked before looking up the source. Missing repository scopes return 404; forwarded principals retain their existing 403 write-denial semantics. A committed replay does not look up or rematerialize the source, so deleting it is compatible with retry as long as the caller still has the required scope. | Status | Meaning | |---|---| | 400 | Invalid request/identity/branch/full commit ID, non-commit source object, or different object formats | | 401 | Missing or invalid credentials | | 403 | Write denied by the existing forwarded-identity contract | | 404 | Missing or unauthorized repository, or unavailable pinned source commit | | 409 | Destination is not pristine, or a committed initializer has a different operation key or fingerprint | | 413 | Request body or configured `server.max_push_bytes` pack limit exceeded | | 422 | Source tree contains a Git LFS pointer; this operation does not copy LFS payloads | | 503 | Temporary object-store/authentication-service failure; retry the identical request | Errors use the existing API envelope: non-503 errors are plain text; 503 uses `{"error":"store_unavailable","retryable":true}` (or the existing authentication service equivalent) and `Retry-After`. Clients accepting `text/event-stream` receive the normal notice/task/result/error envelope while work runs; its HTTP status is 200 and the terminal error packet carries the operation status. Work is registered as an `initialize` task, runs on the bounded bulk runtime, and survives client disconnects. An already committed retry is a fast plain JSON response. Use `Content-Type: application/json` and `Accept: application/json` for plain responses; success includes `Content-Type: application/json` and `Cache-Control: no-store`. Non-503 errors use `text/plain; charset=utf-8`. A store 503 carries `Retry-After: 15`; an introspection-service 503 carries `Retry-After: 5`. Trusted forwarding uses the existing listener-wide grants, not repository scopes: `X-Gitcask-Principal` plus `X-Gitcask-Write: 1`, and the verified `X-Gitcask-Forward-Secret` in `introspect_forwarded` mode. A platform using this transport must authorize both source read and destination write before forwarding. See [SECURITY.md](/llms/security.md#mixed-direct-and-trusted-proxy-authentication) for header stripping and scheme selection. JWT/introspected identities require both repository scopes. No initialize-specific authentication path is introduced. ## Scope and durability The tree is copied unchanged, including binary blobs, executable bits, symlinks, gitlinks and `.gitmodules`. Submodule commits remain external references, as in ordinary Git; the operation never initializes submodules or fetches their URLs. It accepts only Gitcask repository identities, never arbitrary Git URLs, and makes no working-tree checkout. Git LFS pointer blobs are rejected before publication, even if their payload is present in the source; no clone is reported successfully initialized with a missing destination LFS payload. Detection inspects bounded small regular-file blobs and recognizes all three version URLs accepted by Git LFS (`git-lfs.github.com`, `hawser.github.com`, and `git-media.io`), surrounding whitespace, CRLF, and extension lines before the version. Recognized headers are rejected conservatively even when later pointer fields are malformed. Symlink targets are never treated as LFS pointer files. Empty files and ordinary blobs remain supported. Pristine means `head_seq == 0`, no live packs, no log segments, no checkpoint and no initialization receipt. A repository with all refs subsequently deleted is **not** pristine. Failed attempts may leave unreachable immutable uploads under the destination prefix; these do not publish refs or prevent an identical retry. The check runs in the existing publisher against its CAS generation, after each earlier accepted request in a batch, and again after CAS contention. An initializer that wins may be followed by normal writes; an earlier accepted normal write prevents initialization, even when it writes a different branch. One bounded receipt in `Manifest.initialization` records the operation key, SHA-256 request fingerprint and original result. It becomes durable in the same manifest CAS as the ordinary PUSH entry, pack references and branch/HEAD update. Normal publishing, compaction and checkpoints preserve it. There is no secondary commit point, snapshot resource, shared pack or source-lifetime dependency. The receipt survives log retention and later writes; deleting/recreating the destination starts a new repository lifetime and discards its receipt. Before enabling this operation, upgrade **all manifest writers**, including maintainers: older binaries do not preserve unknown protobuf fields when rewriting a manifest and could drop the receipt. The format extension is append-only and existing WAL/checkpoint records remain replayable. The request budget is in [ROUNDTRIPS.md](/llms/roundtrips.md). This is deterministic Git plumbing under [PRODUCT.md Rule A](/llms/product.md), not repository-template policy. # Full-history pristine import The platform accepts `import_from.source` as `/` for internal Gitcask sources or a public HTTPS Git URL. Gitcask exposes the engine protocol below; Cloud owns user authorization, durable intent, leases, retries, 202 responses and job status. The existing [tree snapshot initialize](/llms/initialize.md) is a separate operation and continues to create a new parentless commit. ## Pin, persist, import Create the target first using the ordinary repository PUT, with the matching object format. POST `/{target-owner}/{target-repo}/api/import/resolve`: ```json {"source":"templates/starter"} ``` The 200 no-store response is a snapshot. Persist this entire value before bulk work. A resolve response lost before persistence may be resolved again; resolve creates no durable engine job or reservation and commits no repository state. ```json {"source":"templates/starter","object_format":"sha1","refs":[{"name":"refs/heads/develop","oid":"","peeled":""}],"head":{"symbolic_target":"refs/heads/develop","oid":""},"snapshot_hash":""} ``` Refs are unique and name-sorted, selecting all heads/tags, including annotated tag object IDs and recorded peeled targets. A missing peel hint is empty; bulk validation derives it from the pinned tag object and always publishes the actual peel for cold refs-first readers. A nonempty hint must match that object. A symbolic HEAD must name a selected branch; a detached HEAD has an empty `symbolic_target` and its own commit OID. Hidden provider refs, notes, pull refs, reflogs and unreachable objects are excluded. Unborn/empty and shallow sources are rejected (422); complete independent staging connectivity is required before any destination cache can participate. Internal SHA-1 and SHA-256 are supported; external v1 accepts SHA-1 only, and target format must match (400 otherwise). POST `/{target-owner}/{target-repo}/api/import` with the persisted snapshot: ```json {"operation_key":"creation-123","snapshot":{"source":"templates/starter","object_format":"sha1","refs":[{"name":"refs/heads/develop","oid":"","peeled":""}],"head":{"symbolic_target":"refs/heads/develop","oid":""},"snapshot_hash":""}} ``` `operation_key` is 1–128 printable ASCII characters without whitespace. First success is 201; exact authorized replay is 200, no-store: ```json {"operation_key":"creation-123","request_hash":"","snapshot_hash":"","seq":1,"refs_count":1,"head":{"symbolic_target":"refs/heads/develop","oid":""},"replayed":false} ``` Every reachable object is streamed through isolated Git packing/indexing with fsck and destination connectivity checking, independently stored under target `wal/`, then all refs, HEAD and a bounded receipt commit in the existing manifest CAS. Source code is never checked out or executed; the source is never pushed, repacked or changed. Source deletion and loss of every cache do not affect the target. Source ref movement after resolve cannot repin the import; unavailable pinned objects positively missing from a complete local source/staging probe produce 409, never a silently newer import. Unknown external fetch/subprocess failure is retryable 503; no human stderr text is used for classification. External servers may refuse fetching no-longer-advertised OIDs; that also fails as unavailable. Pristine means no committed WAL work, packs, checkpoint or initialize/import receipt. The publisher checks this for every CAS generation and earlier accepted request in its batch, including writes to other branches. Empty refs after later ref deletion do not restore pristine state. Failed attempts may leave immutable uncommitted uploads, but never visible refs. Different operation keys or payloads conflict (409). Exact replay returns the original receipt without opening source, following later source refs, or changing later target writes. Receipts survive checkpoint/log retention; deleting and recreating the target starts a new lifetime. All serving readers and manifest writers must be upgraded together so old protobuf writers cannot silently discard newly added receipt fields. ## Reconcile and authorize Every resolve/bulk attempt uses the existing resolved principal: target write and internal source read permissions, including replay. Scopes missing return 404; trusted forwarding preserves its listener-wide grants and existing 403 write denial. Forwarded read requires only a verified nonempty principal, and write uses `X-Gitcask-Write: 1`; there is no `X-Gitcask-Read` header. A forwarding platform must authorize both repositories itself. External URLs require no credentials and v1 does not support private OAuth, Basic credentials or token-bearing URLs. After lost success or source grant revocation, target read alone can confirm the committed result with `GET /{o}/{r}/api/import/receipt?operation_key=...&request_hash=...`. Receipt/import replay revalidates the target even when ordinary reads opt into `wal.freshness_ttl`. Matching receipt returns 200 and `replayed:true`; absent receipt is 404 and key/hash mismatch is 409. It never touches source. This endpoint reports a Git commit, not pending engine job state. Cloud can reconcile before reacquiring source grants. JSON request/response contracts use `Content-Type: application/json` and `Accept: application/json`. Both resolve and import support the existing `Accept: text/event-stream` task/progress/result/error envelope (HTTP 200; terminal error contains its operation status). Tasks are per-instance, replayable and ephemeral; object work runs on the bounded bulk runtime and continues after HTTP disconnect. Phase-2 drain refuses new resolve/import work with 503. Acquisition has a deadline; deadline expiry kills and verifies every owned Git process group before unwinding the acquisition future and removing scratch. Process verification gets a two-second poll bound and three-second foreground grace. If death cannot be verified, a cleanup supervisor retains the future, ingest lock and scratch and retries verification; a retryable error never claims that deferred resources are already gone. The container includes procps for kill/ps. Index-pack and connectivity use supervised async native children only on the import path; receive-pack ingestion is unchanged. The WAL publisher handles uncertain CAS outcomes independently of acquisition cancellation. No engine identity, Redis, durable job DB, node routing or Cloud store credentials are introduced. ## Preservation boundaries and bounds Gitlinks and `.gitmodules` remain unchanged, with no recursive submodule fetch. **LFS payload transfer is unsupported:** recognized LFS pointer headers in any acquired historical small blob (not just the tip) cause 422 before publication, even when source payload exists. Detection is conservative and includes pointer-like symlink blobs; ordinary symlinks are preserved. No successful target is advertised as having imported absent LFS payloads. External URLs must be credential-free HTTPS on port 443, without query/fragment. DNS results are resolved once per acquisition, all addresses must be public, and the HTTPS client pins exactly those addresses while retaining TLS hostname/cert validation. Loopback, private, link-local, CGNAT, metadata, documentation, multicast/reserved IPv4, the Azure platform virtual IP (`168.63.129.16`), IPv4-mapped IPv6 and IPv6 transition/special ranges are rejected. No environment proxy is used. All redirects are refused, including same-origin redirects: callers must provide the canonical smart-HTTP Git URL. Dumb HTTP fallback is unsupported. Git sees only a scoped loopback relay exposing GET info/refs and POST upload-pack; arbitrary paths, receive-pack and headers are not forwarded. Authorization and cookies cannot reach the source. Git runs with a cleared environment, no system/global config, credential helpers, replace objects, auto-maintenance, hooks or non-HTTP protocols. Only isolated scratch/staging is written; destination never gains durable alternates. Packs stream to files/index-pack; no complete pack is buffered in memory. `[import]` defaults: 1024 refs, 1,000,000 acquired objects, 1 GiB aggregate external response bytes and output pack (also bounded by `server.max_push_bytes`), 15-minute acquisition timeout, 30-second resolve timeout. Resolve advertisements are capped at 4 MiB transferred and 1 MiB parsed Git output. JSON snapshot/body is capped at 1 MiB; resolve input at 4 KiB. Names are capped at 1024 bytes. The object limit is checked on independently acquired history before publishing; compressed transfer may require indexing first. The source's ordinary internal materialization uses existing pack/store limits. These are resource admission limits, not per-tenant quotas. Cloud should allow the acquisition deadline plus publisher/store recovery when setting its long HTTP deadline, and reconcile uncertain results by receipt. | Status | Meaning | |---|---| | 400 | Invalid snapshot/hash/format/source URL or unsafe DNS result | | 401 / 403 / 404 | Existing auth/permission/missing-repository contract | | 409 | Nonpristine, operation conflict, or pinned source unavailable | | 413 | JSON, refs, transfer, output pack or object bound exceeded | | 422 | Empty/unborn source, LFS pointer history, or source requiring auth/redirect/dumb HTTP | | 503 | Resolve busy, deadline, source DNS/transport, store/auth failure or drain; retry fixed request | Non-503 errors use the existing plain-text envelope; 503 uses retryable JSON with Retry-After (import failures identify `import_unavailable`). Unknown Git acquisition/audit or source transport errors return retryable503; Cloud retries the persisted snapshot and never repins it. External servers refusing a no-longer-advertised OID may also return this conservative retryable class when the transport cannot prove the precise source absence. ## Exact canonical hashes Append fields as decimal UTF-8 byte length, ASCII `:`, exact UTF-8 bytes, with no separator. Compute SHA-256 and encode lowercase hex. This avoids JSON key ordering, Unicode/HTML escaping, or Go/Rust serialization differences. - Snapshot prefix is `gitcask-import-snapshot-v1\n` (literal LF). Fields: source, object_format, head.symbolic_target, head.oid, decimal refs count (a string), then each sorted ref's name, oid, peeled. The snapshot_hash field is excluded. - Request prefix is `gitcask-import-request-v1\n` (literal LF). Fields: target owner, target repo (validated URL segments), operation_key, snapshot_hash. Thus the same snapshot/key for a different target has a different fingerprint. Persist source exactly as returned by resolve; OIDs are lowercase/nonzero and peeled is always present, including empty strings. Hash fixtures and tests live in the import modules and Cloud counterpart. [ROUNDTRIPS.md](/llms/roundtrips.md) owns bucket budgets. Security controls use [Git configuration](https://git-scm.com/docs/git-config) [Azure platform virtual IP](https://learn.microsoft.com/en-us/azure/virtual-network/what-is-ip-address-168-63-129-16), and [reqwest client DNS/redirect/proxy settings](https://docs.rs/reqwest/latest/reqwest/struct.ClientBuilder.html). ## Reproducible validation `RUSTUP_TOOLCHAIN=1.97.1 timeout 300 cargo test -p gitcask-server --test import` runs synthetic HTTP/Git preservation, cross-instance replay/race, crash/fault, authentication, bounds and store budgets. `cargo test -p gitcask-server --lib tls_tests` exercises real TLS with a synthetic CA, Git acquisition and redirect/ credential stripping through an injected test-only client; production has no loopback or TLS bypass. `cargo test -p gitcask-server --test sim` exercises the shared WAL crash/partition/replay protocol. `scripts/import-smoke.sh BASE` is the standalone synthetic HTTP contract smoke, usable after shell/process restarts. # Events Context: **the normative WAL-to-webhook event contract**, for anyone changing ref-event shapes, the events bridge, delivery, or consumer semantics (D32). The golden tests in `crates/gitcask-server/src/{events,bridge}.rs` and `tests/events.rs` are its executable form. ## Principle **Events are produced from the WAL by one small service — the bridge (`Role::Events`) — never by the push path.** It tails each repo's log from a durable per-repo cursor, converts committed entries, POSTs them to your webhook, advances the cursor. So an event is delivered iff its entry is durable, a crash can't lose one, the lag is `head_seq − cursor`, and no writer — any serving host, the CLI, an import — contains event code: every writer is covered because the WAL is what is read. Invariants: 1. Nothing **ever gates a push**: the bridge is another process reading the bucket. A down webhook adds zero milliseconds to receive-pack; it adds lag, which is a metric. 2. No no-op events. `old == new` (and `0→0`) emits nothing. 3. No lost events: the cursor advances only after the webhook answered 2xx. Duplicates are possible (at-least-once) and carry a deterministic dedup key. What *can* happen is a **gap when the bridge lags behind log retention** (entries folded into a checkpoint before they were read): counted (`events_bridge_gap_total`) and warned, never silently repaired — consumers backfill from the WAL. 4. One producer, one instance. Only `ref` events exist. Not events: push denials and auth failures (metrics + logs), LFS, compaction/checkpoints (already WAL entries), repository administration (no consumer; the HTTP request log has the principal). ## `ref` event ```json { "action": "update", "ref_type": "branch", "ref_name": "refs/heads/main", "old": "48a0637…", "new": "cb38da1…", "pusher": "alice@example.com", "correlation_id": "d1f916f7-…", "repo": "acme/monorepo", "_gitcask": { "schema_version": 1, "seq": "42", "entry_kind": "push", "request_id": "d1f916f7-…" } } ``` - `action` — `create` / `update` / `delete`. Force is not a wire action (consumers derive it). - `old` / `new` — **always the full zero OID on create/delete, never `""`** (40 chars sha1, 64 sha256). - `ref_type` — `branch` (`refs/heads/`), `tag` (`refs/tags/`), `""` otherwise. - `pusher` — the opaque authenticated principal in the log entry's `meta` (JWT `sub`, or the trusted forwarded principal). - `correlation_id` / `_gitcask.request_id` — the user-visible request id: the middleware honours an incoming `x-request-id` (else mints one), a front forwards it when it forwards receive-pack, `push_meta` stores it. - `_gitcask.seq` — the entry's WAL seq, a JSON **string** (uint64 convention). `entry_kind` — `push` | `ref_update`; consumers must not care. - One event per ref update in the transaction; symbolic (HEAD) retargets and COMPACT / CHECKPOINT entries emit nothing. **Dedup key (normative): `(repo, _gitcask.seq, ref_name)`.** **Order: by `seq` per repo.** Nothing else is promised. ## Delivery: the webhook Each catch-up `POST`s one JSON **array** of events (a batch: everything in `(cursor, head_seq]`) to `events.webhook_url` with: ``` Content-Type: application/json X-Gitcask-Delivery: # the batch's id; safe to dedup on X-Gitcask-Signature: sha256= # when a secret is configured ``` Answer 2xx to acknowledge; anything else (or a timeout, 10 s) leaves the cursor where it was and the bridge retries the same range on the next wake-up. A consumer therefore sees at-least-once delivery of whole batches. Verify the signature with a constant-time compare before parsing. **Backfill contract (normative for consumers):** on any gap, read the WAL log from your last known seq (`gitcask wal ls`) and treat each PUSH / REF_UPDATE entry's ref transaction as the missed events. The webhook is a latency optimization over polling the log; correctness never depends on it. ## The bridge ``` writers (any host, CLI, import) ── manifest.pb CAS ──► bucket ──► notification ──► POST /_events/notify (no event code) │ ▼ the events host (roles=["events"], one instance): catch_up(repo): cursor → manifest + sweep pending/ ∪ cached repos (backstop + health check) → log (cursor, head] → ref events → webhook → CAS cursor → log line ``` `catch_up(repo)` = read `repos///events/cursor.json` → fresh manifest → log entries `(cursor, head_seq]` → `ref` events → webhook → CAS the cursor to `head_seq`. A webhook error leaves the cursor; the next wake-up replays the same range. A cold cursor starts at the oldest readable seq (`min_seq − 1`: everything still in the manifest's log window is published once; pre-seed the cursor to skip history). Every published event is also one structured log line (`event_type="ref"`). Metrics: `events_published_total{sink}`, `events_bridge_lag_entries{repo}`, `events_bridge_gap_total{repo}`, `events_bridge_sweep_found_total`; alert on lag growth, any gap, any sweep-found. Wake-ups (both idempotent; they only ever call `catch_up`): - `POST /_events/notify` with a **bucket notification** naming a finalized `…/manifest.pb` — the commit point itself as the notification. Accepted bodies: an S3 event notification (`Records[].eventName = ObjectCreated:*`, `s3.object.key`; MinIO, rustfs and Ceph emit the same shape), or your own glue's `{"key": "repos/o/r/manifest.pb"}` / `{"repo": "o/r"}`. Everything else is acked and ignored; a webhook failure answers 503 so the notifier redelivers. Authenticated like every route (`require_read`): the front proxy supplies its principal. - The sweep (`events.sweep_interval`, default 5 min): one `LIST pending/`, unioned with repositories cached on the events instance, then `catch_up` once per distinct candidate. A committed push writes its pending marker after the manifest CAS; the cursor advances only after the webhook ACK, so eviction cannot remove the marker-based wake-up. Not needed for correctness; it is the backstop *and the health check* — a sweep that publishes anything means notifications are not flowing (`events_bridge_sweep_found_total`, warn). With no notifier at all, set the sweep to the latency you can live with. ```toml [server] roles = ["events"] # or leave roles empty on a one-box install: every role, bridge included [events] webhook_url = "https://hooks.example.com/gitcask" webhook_secret = "…" # env: GITCASK__EVENTS__WEBHOOK_SECRET sweep_interval = "5m" ``` ## Consumer checklist 1. Verify `X-Gitcask-Signature` (if you set a secret), then parse the array. 2. Dedup on `(repo, _gitcask.seq, ref_name)` (or on `X-Gitcask-Delivery` per batch). 3. Order by `_gitcask.seq` within a repo; do not assume order across repos. 4. On a gap alert from the bridge, backfill from `gitcask wal ls --from `. # LFS — objects in the store Context: **spec** for Git LFS on gitcask. For anyone touching `crates/gitcask-server/src/lfs.rs`, importing a repository with LFS objects, or debugging "(missing)" in a push's LFS pre-push. `AGENTS.md §1.4` lists LFS as part of the surface; this is the detail. ## 1. Protocol and storage (✅) - Batch API `POST /{o}/{r}.git/info/lfs/objects/batch` (`operation = upload | download`, transfer `basic`), basic transfer `GET|HEAD|PUT /{o}/{r}.git/info/lfs/objects/`, `POST …/info/lfs/verify`. The front proxy authenticates these routes like the rest of gitcask and forwards the principal and grants. - Objects live in the repository's prefix at `lfs/objects///` (`gitcask_proto::keys::lfs_key`) — sha256-addressed, immutable, served by `static_object` with the full static contract (strong ETag, 304, Range/If-Range, HEAD; `X-Accel-Redirect` to an edge's cache when one announces it, D23). `PUT` verifies size + sha256 before the store write. `lfs.max_object_bytes` (16 GiB) bounds an upload. ## 2. Not done / open - `lfs.serve_via = "signed_url"` hands out presigned S3 URLs; the default `proxy` streams through gitcask or the edge. - Size accounting of LFS bytes per repository in the overview. # Object integrity — the invariant and the audit Context: **spec** for the one data invariant the WAL cannot express: *every object reachable from an advertised ref is in a live pack*. For anyone touching `gitcask import`, the maintainer's `fsck` unit (`crates/gitcask-server/src/ops.rs`, `maintain.rs`), or debugging `connectivity: missing object` on a push / `NotFound` on a partial-clone fetch. Born from a large-repository recovery test (the monorepo's 1,952 missing blobs). ## 1. The invariant and where it can break The WAL guarantees *which* packs are live and *which* refs point where; it does not know whether the packs hold the refs' closure. Pushes cannot break it (receive-pack checks connectivity before publishing — that check is what surfaced the hole). What can: - **Import**: a pack set built from one ref selection and a ref snapshot taken from another. The default importer avoids that split by packing all source refs before publishing the selected heads and tags through the normal WAL path. `--reuse-packs` requires a self-contained source pack set; the audit below is the backstop for an incomplete or corrupt source repository. The [HTTP pristine importer](/llms/import.md) instead packs exactly its persisted heads/tags/HEAD OIDs into independent staging, fscks that closure, then publishes the same pinned ref set and receipt through the existing manifest CAS. Source ref movement never changes its object selection. - A compaction that drops objects (reachability from a stale tip), a superseded pack GC'd too early, a corrupt or truncated object in the bucket. None seen; the audit below is the detector for all of them. ## 2. Audit: the `fsck` unit (✅ 2026-08-21) Lowest-priority maintainer unit (`maintenance.fsck_interval`, default 7 d). Runs `git fsck --connectivity-only --no-dangling` (`ops.rs` `fsck`, `connectivity=1`) and writes the verdict to **`repos///fsck.pb`** (`FsckReport {seq, at, host, missing[≤100 k], missing_total, problems, elapsed_secs, audited_seq}`; overwritten, not WAL). Missing objects are a finding and corrupt objects a failure. Gauge `gitcask_repo_missing_objects{repo}` is set from the report on every pass. Due immediately after compaction; otherwise, once a report exists, due after `fsck_interval` only when the WAL advanced beyond `audited_seq`. A new repository that has not been compacted is not audited on its first maintainer visit because receive-pack already checked its connectivity. An interval of 0 removes the age delay but still requires a prior report and a new push. Manual: `POST /{o}/{r}/api/ops/fsck` with `connectivity=1` (WAL page) writes the same report. the monorepo on a complete local copy: ~10 min. # Operations runbook Context: for whoever operates gitcask instances and their S3 bucket — incident diagnosis, capacity planning and recovery procedure. The design starts from [GOAL](/llms/goal.md); the WAL and maintenance principles are in [AGENTS](/llms/architecture.md); the cost of every bucket round trip is in [ROUNDTRIPS](/llms/roundtrips.md). This document does not repeat them; it is about where to look after an alert fires. ## 1. First look `/healthz` only says the process answers; `/readyz` only says whether serving drain has started. Neither touches S3, so a 200 does not mean the bucket is fine. ```sh curl -si "$BASE/healthz" curl -si "$BASE/readyz" curl -s "$BASE/metrics" gitcask --config gitcask.toml wal pending gitcask --config gitcask.toml repo info owner/repo curl -s "$BASE/owner/repo/api/overview" curl -s "$BASE/owner/repo/api/tasks" ``` - `wal pending` is a manual check that LISTs `pending/`. Do not call it repeatedly or put it on a request path. - `repo info` runs `sync_full()` and pulls packs local, so it is expensive on a cold repository. Use it only when digging into one repository; for refs and maintenance state, look at `api/overview` first. - `api/tasks` records are per instance. Keep the response's `hostname` together with the log `request_id`. - A metric appears only after its code path has run once. A missing series may mean zero — or that this role or unit has simply not executed in this process yet. With `telemetry.log_format = "json"`, every span-close line carries `span.name`, `elapsed_ms`, `outcome`, `repo` and `request_id`. The span names used in the procedures below are the real names in the code. ## 2. Key metrics and alerts The thresholds below are starting points. Tune the durations to your traffic baseline and SLOs, but keep the comparisons against configuration values as they are. Sum counters across all instances of a role; read gauges labelled `host` per instance. Gauges with a `repo` label exist to narrow down a problem repository — they are not a substitute for a repository listing. | Name | Normal | When it moves | Suggested alert starting point | |---|---|---|---| | `gitcask_auth_introspect_total{outcome}` | `hit` for cached answers; `miss` for uncached verifies (including followers); `active`/`inactive` for upstream token answers | `error` counts unavailable upstream calls; inactive includes invalid principal/scope/TTL answers. Service errors never enter the cache | any sustained `error`; a logged endpoint 401/403 means gitcask's service secret was rejected | | `gitcask_auth_introspect_seconds` | upstream calls below `auth.introspect.timeout` | elapsed HTTP request/body/parse time on the single-flight leader, including failures; no token/principal labels | p99 approaching the timeout for 5 minutes | | `gitcask_push_refused_total{reason}` | flat outside deploys | `connectivity` = object-closure failure on a new tip, `unpack` = pack parse/index failure, `body` = the pack body was not received (over `server.max_push_bytes`, read error, client abort), `draining` = a push during serving drain | one `connectivity`/`unpack` immediately; `body` only as a sustained rate (a Ctrl-C'd push counts); one `draining` outside a deploy window | | `gitcask_publish_local_apply_failed_total` | 0 | the manifest CAS succeeded but applying refs locally on that instance failed. The next sync repairs it, but it is the lead on an immediate-visibility regression | on any increase | | `gitcask_pending_marker_put_failures_total` | 0 | the push committed but the best-effort marker PUT that wakes the maintainer failed | on any increase; check that repository by hand | | `gitcask_store_retries_total{op}` | usually flat | transient S3 errors (5xx, 429/throttling, connection failures) being retried internally | warn if still climbing after 5 minutes | | `gitcask_store_requests_total{op,outcome}` | `ok`, ordinary `not_found`/`precondition_failed` | `retryable_error` = retries exhausted, or a transient error on a conditional write that is never retried at this layer; `error` = any other store failure | one `retryable_error`/`error` immediately; read the `store.*` log lines from the same moment for the exact S3 status | | `gitcask_pending_markers` | 0 when idle; converges back to 0 in the passes after a push burst | maintainer backlog. A value equal to `maintenance.max_repos_per_pass` may mean at least one more page behind it | warn if > 0 for two whole maintenance intervals; page if pinned at the page limit | | `gitcask_maintain_workers_busy{host}` | 0–`maintenance.workers`, 0 without backlog | pending work in progress. Pinned at the limit while backlog grows = not enough throughput | all workers busy for ≥ 2 intervals while `gitcask_pending_markers > 0` | | `gitcask_maintain_units_total{host,kind,outcome}` / `gitcask_maintain_unit_seconds{kind}` | `outcome="ok"`, durations within your per-repo-size baseline | a checkpoint/compact/rev-index/fsck/gc unit failing or slowing down | any `outcome="failed"`; warn on sustained p99 above baseline | | `gitcask_maintainer_heartbeat_timestamp{host}` | within two intervals of now; refreshed every 2 minutes even during long units | maintainer loop stopped, process paused, or store writes failing | `time() - value > max(2 * maintenance.interval, 5m)` | | `gitcask_checkpoint_lag_entries{repo}` / `gitcask_checkpoint_age_seconds{repo}` | inside the active `wal.*` triggers at the moment a marked repo was planned | the checkpoint unit not keeping up with its triggers | non-zero beyond the entries/interval trigger for two passes. Age is the value at gauge-update time, not a running clock | | `gitcask_checkpoints_total{outcome}` / `gitcask_checkpoint_seconds` | `outcome="ok"`, duration within your refs/tail baseline | checkpoint writes failing or suddenly slow | any `outcome="error"`; warn on sustained p99 above baseline | | `events_bridge_lag_entries{repo}` | 0 after catch-up | `head_seq - cursor`; the webhook or the bridge is behind | > 0 for longer than `events.sweep_interval` | | `events_bridge_gap_total{repo}` / `events_bridge_sweep_found_total` | flat | gap = the cursor fell behind WAL retention; sweep-found = the regular notification was missed and the backstop found it instead | one gap immediately; one sweep-found as a warning | | `events_published_total{sink}` | grows when refs change | if lag is 0 and this grew, investigate the consumer side beyond gitcask | use together with lag/gap, not as a standalone alert | | `gitcask_cache_disk_used_fraction` | below `cache.disk_high_watermark` | usage of the **whole filesystem** holding `cache.dir` is high | above the watermark for ≥ 2 × `cache.evict_interval` | | `gitcask_cache_evicted_total` / `gitcask_cache_repos` | moves with the idle policy | pressure without eviction growth = every candidate is in use or deletion is failing; a spike = cache churn | sustained disk pressure with eviction stalled, or eviction abruptly above normal | | `gitcask_lock_wait_seconds{lock}` | mostly instant; recorded waits below `telemetry.lock_wait_warn` | contention on `rw.read`, `sync_mutex`, `pack_mutex` | p99 above `telemetry.lock_wait_warn` for 10 minutes | | `gitcask_runtime_stall_total` / `gitcask_runtime_stall_seconds` | flat | the tokio ticker ran > 2.5 s late | check logs on any increase. If the same line shows `inflight > 0`, treat as a real worker block/starvation — urgent | | `gitcask_http_inflight` / `gitcask_tasks_running` | returns to the instance baseline | streaming requests that never finish, or long tasks | warn when baseline and client latency climb together | | `gitcask_repo_missing_objects{repo}` | 0 | the last connectivity audit found an advertised ref pointing at a missing object | > 0 is a data-integrity incident, immediately | | `gitcask_gc_deleted_total{kind}` | grows only when GC has something to collect | deletion trend. Zero here is not by itself a failure | never alert alone; judge together with failed due-GC tasks and bucket growth | There is no dedicated store-latency Prometheus metric. Do not build dashboards on series that do not exist — derive p50/p95/p99 from the JSON logs: `span.name = store.get|store.head|store.put|store.delete` with `elapsed_ms` and `outcome`. Push breaks down into `receive.body` (request-body reception, overlapping the sync; fields `http_version` and `response_start` = `live` | `after_eof`, with `git.spool_pack`: `bytes` and `outcome`), `receive.ingest`, `git.ingest_pack`, `receive.connectivity`, `receive.publish` and `wal.publish` spans. For clone/fetch, total streaming time is judged by the git client's own timing plus the band-2 progress lines; cold sync splits into `wal.sync`, `wal.materialize`, `wal.reconcile_packs`, `wal.download_pack`, and pack computation into `git.upload_pack`. `gitcask_push_refused_total` is not the sum of all push failures — only the four refusals in the table. A ref conflict is reported as git protocol `ng` and can be HTTP 200; publish/store failures are judged by 503s or by pkt-line errors in an already-started stream plus the logs. There is also no general HTTP status / request-duration Prometheus metric today. A repository with no checkpoint yet has no age series either; read `health.deep` and the checkpoint suggestion in `api/overview` together. With `cache.disk_high_watermark = 0`, both pressure eviction and the disk-used gauge are disabled. The source of truth for integrity is `fsck.pb`. `gitcask_repo_missing_objects` carries only the missing-object count; other fsck findings are read from `health.deep` in `api/overview` and the report. Audit cadence and meaning are defined in [INTEGRITY](/llms/integrity.md) only. ## 3. Diagnosis, starting from the symptom ### Clone or push is slow **Check, in order** 1. Keep the git output and wall time for the affected repository. Use band 2's `local copy is missing packs` / `local copy ready (...s)` to split before/after sync. 2. Look at `gitcask_store_retries_total`, store error outcomes and `store.* elapsed_ms`. If every repository is slow at once, suspect S3 or the network first. 3. Long `wal.materialize`/`wal.download_pack` with only the first request slow = cold cache. Which way a push answered is on its `receive.body` span (D52): `response_start = after_eof` (HTTP/1.x, or no side-band) means no HTTP status or band-2 line left before the request body ended — the client shows only its own upload progress, and the sync narration said meanwhile is replayed right after the banner; `live` (side-band over HTTP/2) narrates from the start. `http_version` is the version of the connection reaching gitcask (the last hop), not the client's: a direct or standalone git client over `http://` and any HTTP/1.1 proxy hop give `after_eof`; `live` appears only behind a proxy speaking prior-knowledge h2c to gitcask. A long `git.spool_pack` is the upload itself, and `outcome = error|cancelled` with `bytes` below the client's pack size means the body stopped arriving (client or proxy), before any gitcask work. A `live` push that stalls there behind a proxy speaking HTTP/2 to gitcask and HTTP/1.x to the client is that proxy dropping the body after an early response: front gitcask with HTTP/1.1 instead. Long `git.ingest_pack` inside `receive.ingest` = pack indexing/fsck; long `receive.connectivity` = the object graph walk; long `wal.publish` = the PUT/CAS phase. 4. Check `gitcask_lock_wait_seconds`, `gitcask_runtime_stall_total`, `gitcask_http_inflight`. On an `async runtime stalled` line, read `inflight`, `tasks_running`, `lock_wait_max_ms` and `rss_mb` together. 5. In `api/tasks`, follow the `materialize` task's progress and result. A second operation on the same `(repo, kind)` joins the existing task. **Common causes and actions** - Cold materialize on the first object request: normal. Confirm it matches pack size and S3 throughput; give the cache disk and instance lifetime room. - Store retries/latency: go to the S3 outage procedure below. - Repeated evict/materialize churn under disk pressure: grow the cache filesystem, or separate the cache from other files' filesystem. - Lock waits or runtime stalls: find the blocking section in that `request_id`'s span tree and report it as a bug. Blocking git/fs work on the serving runtime is never a tuning matter. - Clone still slow after `local copy ready` with store and locks healthy: isolate pack computation with the `git.upload_pack` span. CPU-saturated → add serving instances; a single repository larger than the disk → a bigger instance. Client time far above the span → investigate the transport path. ### A push gets 503 **Check, in order** 1. Immediately: `curl -si "$BASE/readyz"`. 2. `503` + `Retry-After: 15` + `{"status":"draining"...}` = serving drain, phase 2. 3. `/readyz` 200 but request logs show 503s and `gitcask_store_retries_total` or `gitcask_store_requests_total{outcome="retryable_error"}` is climbing = S3 outage. The plain JSON response is `{"error":"store_unavailable","retryable":true}`. 4. Cross-check `gitcask_push_refused_total{reason="draining"}` against deploy times. The store span's `error` carries the actual connection/status cause. **Actions** - Drain: retry against another ready instance. Never put a terminating instance back into ready. - S3 outage: do not hammer writes into a retry storm; restore S3. Afterwards confirm with a refs read and one small push. - A store failure after the git sideband stream has started cannot change the HTTP status and may end as a pkt-line error. Judge git client failures and server store logs together. ### A push succeeded but the ref is not visible on another instance This is not eventual consistency to be tolerated — in the default configuration it is a correctness bug. But first confirm both instances use the same bucket + prefix and `wal.freshness_ttl = "0s"`; a positive TTL skips the manifest GET for that long. **Reproduce and collect evidence** ```sh new_oid=$(git rev-parse HEAD) git push "$INSTANCE_A/owner/repo.git" HEAD:refs/heads/main GIT_TRACE_CURL=1 git ls-remote "$INSTANCE_B/owner/repo.git" refs/heads/main \ >ls-remote.out 2>ls-remote.trace curl -s "$INSTANCE_A/owner/repo/api/overview" >overview-a.json curl -s "$INSTANCE_B/owner/repo/api/overview" >overview-b.json gitcask --config gitcask.toml repo info owner/repo ``` If B's first request returns the old OID, do not paper over it with a retry. Keep: `$new_oid`, the push time and result, the `x-request-id` from the ls-remote trace, both overviews' `hostname` + manifest version + seq, both instances' commit SHAs and configuration (bucket/prefix/freshness TTL), and the `wal.sync`/`store.get` spans from the same moment. Take B out of serving but do not wipe its cache or logs. The minimal regression check: ```sh export RUSTUP_TOOLCHAIN=1.97.1 cargo test -p gitcask-server --test e2e two_instances_consistency -- --nocapture cargo test -p gitcask-server --test sim ``` ### The disk is full **Check, in order** ```sh df -h /path/to/cache.dir du -sh /path/to/cache.dir ``` 1. Compare `gitcask_cache_disk_used_fraction` with `cache.disk_high_watermark`. It is the usage of the whole filesystem, not the size of the cache directory. 2. Read `gitcask_cache_evicted_total`, `gitcask_cache_repos` and the `cache disk above high watermark`, `cache repositories evicted`, `cache directory removal failed` log lines. 3. In `api/tasks`, look for work that is using packs — materialize, compact, fsck. The evictor skips busy repositories whose `sync_mutex` or repo write lock it cannot take immediately. **Actions** - Send new object work to other instances so in-flight work can finish and eviction can proceed. - Grow the filesystem or remove non-cache occupants. Never delete the cache directory by hand. - Above the watermark, the evictor removes least-recently-touched repositories toward `disk_high_watermark - 0.10`. Repositories past `cache.evict_idle_after` are removed on the next `cache.evict_interval` even without pressure. - The cache is disposable, but an instance without room for the packs the WAL points at will keep failing. Size the disk so the largest repository *and* the scratch/output packs of work in progress fit together. ### Webhooks are not arriving The contract and cursor recovery follow [EVENTS](/llms/events.md). 1. Is `events_bridge_lag_entries{repo}` > 0? Then the bridge or the webhook endpoint is failing. 2. Lag 0 and `events_published_total{sink}` grew: gitcask got a 2xx and advanced the cursor — investigate the consumer's dedup/processing logs. 3. `events_bridge_sweep_found_total` growing: the backstop sweep found unpublished WAL — fix the bucket notification flow. 4. `events_bridge_gap_total` growing: WAL before the cursor was folded past retention. Do not advance the cursor by hand; backfill from the last consumed seq with `wal ls`/`wal show`. 5. Restore the webhook and wait for the next notification or sweep. The cursor never moves before a 2xx, so the same batch may be delivered again. ### A fresh instance is slow Refs-only requests read the manifest plus checkpoint/tail and need no packs. The first clone/fetch/push that needs objects materializes the whole live pack set into `cache.dir`, and can take as long as the repository is large. - One `wal.materialize` task, `wal.download_pack` spans and one band-2 `local copy is missing packs`, then fast subsequent requests: normal cold cache. - Repeating every time: check `gitcask_cache_evicted_total`, the disk watermark, whether the cache directory is writable, and the instance's lifetime. - With `wal.prefetch_packs = true`, pack reconciliation starts in the background after a refs-only sync — but there is no prewarm or remote-pack path. The first object request still requires the full pack set to fit locally. ## 4. Capacity and placement The default operating size is 2 vCPU ([DIRECTION D11](/llms/direction.md)). Scale by adding instances of the same role, not by raising worker counts on one instance. Distinguish the signals: - **Add serving**: store, disk and locks healthy, but client latency and `gitcask_http_inflight` are high on several instances at once. Concurrent git work on one repository also shares the `server.max_concurrent_per_repo` semaphore; over HTTP/1.x a side-band push holds its permit for the whole upload (D52), so slow uploaders to one repository count against it. - **Add maintainers**: `gitcask_pending_markers` does not converge and every maintainer's workers stay busy. First confirm one worker has the disk headroom to materialize. Do not hide a CPU bottleneck by raising thread counts past the D11 default. - **Store scaling/investigation**: adding instances makes store retries and `store.* elapsed_ms` worse together. More serving here only multiplies S3 request volume. Size the disk so the sum of the following stays below the watermark: 1. Live packs + idx/rev/bitmap/commit-graph of the repositories concurrently active on this instance. 2. Packs being materialized by `maintenance.workers` concurrent workers. 3. Received push bodies (unlinked spool files, invisible to `du` but counted by the filesystem) and scratch/output packs from receive-pack ingest and compaction — during compaction, inputs and outputs coexist. 4. Headroom up to `disk_high_watermark`, plus whatever else lives on the cache filesystem. The floor: the full live pack set of your largest repository plus its working scratch must fit on one instance; otherwise you need a bigger disk. `cache.bulk_threads` (default 2) is the async lane dedicated to pack materialization; blocking git/fs work runs on a separate blocking pool. Know what grows with repository count and what does not: | Grows | Does not grow with total repository count | |---|---| | Actual git data + WAL objects in the bucket; pending work created by pushes | Repositories the maintainer visits: only those with a `pending/` marker | | The local cache of repositories this instance actually touched | Periodic `repos/` LISTs while idle: there are none | | S3 requests proportional to pushes, compactions, checkpoints, GC | Checkpoint/fsck visits to unpushed repositories: never evaluated without a marker | | Maintainer heartbeats proportional to live instances | Holding every repository's packs locally: only the working set before eviction | If gitcask's periodic cost rises merely because idle/empty repositories multiplied, that is a D40 regression. ## 5. Incidents and recovery ### S3 outage The S3 backend retries retryable failures on GET, HEAD, LIST, DELETE, multipart stages and unconditional PUTs up to `store.max_retries` with full-jitter backoff. Conditional Create/Update PUTs are never retried at the store layer (replaying one makes success ambiguous); the manifest CAS protocol re-checks whether the commit actually landed on its failure path and owns its own CAS retry. If retries are exhausted, HTTP requests that have not started responding become 503 + `Retry-After: 15`. A push never reports success before the bucket ACKs the commit. In order: 1. Scope the op/key/error from store retry/error rates and `store.*` spans. 2. Reduce new write load; restore the S3 endpoint's connectivity, service health, rate limits and bucket state. 3. Verify the read path with a real refs read, not `/healthz`. 4. Push a small test repository and immediately confirm the OID from a different instance. 5. If any push was answered success *during* the failure, treat it as a correctness incident — by design there should be none. ### All instances lost Start new instances pointing at the same config and bucket. There is no persistent-disk recovery and no node-to-node state transfer. ```sh gitcask-server --config gitcask.toml ``` Check `/healthz` and `/readyz`, then `ls-remote` and clone a representative repository. The first refs request replays checkpoint + WAL tail; the first object request downloads full packs, so both are slow. That cold cost and the normal round trips are defined in [ROUNDTRIPS](/llms/roundtrips.md). Do not restore `cache.dir` from backup. ### Bucket corruption or accidental deletion gitcask has no bucket-backup engine of its own. Production S3 should run with versioning on, plus cross-account/region replication or S3 Inventory. When an incident happens, stop new writes, then: 1. Pin down the affected prefix and the first moment of damage. Never hand-delete packs/logs/checkpoints that `manifest.pb` points at, and never hand-craft a new manifest. 2. Restore deleted/corrupted objects to their original keys from S3 object versions/replicas. Verify checksum and size on immutable packs/logs. 3. Clone into a fresh cache and run `git fsck --full --strict`; check the maintainer fsck report as well. 4. Do not return to serving while `gitcask_repo_missing_objects > 0`. The full verdict procedure is [INTEGRITY](/llms/integrity.md). If a force push merely moved a ref, do not roll the bucket back. Find the seq in the WAL and resurrect the old OID as a new ref update. `wal ls` shows seqs and ref-update counts; the actual `old_oid` comes from `wal show`. ```sh gitcask --config gitcask.toml wal ls owner/repo gitcask --config gitcask.toml wal show owner/repo 42 gitcask --config gitcask.toml wal materialize owner/repo --at-seq 41 --out /tmp/repo-restore git -C /tmp/repo-restore fsck --full git -C /tmp/repo-restore push --force "$REPO_URL" \ :refs/heads/ ``` If `materialize --at-seq` fails explicitly as beyond retention, the needed logs/packs are already collected; without S3 versions/replicas that point in time cannot be restored. ### Deploys and rollback After SIGTERM/SIGINT, drain happens in two phases: 1. **Maintenance drain**: no new unit starts and the running unit is interrupted at once. Serving is normal and `/readyz` stays 200. Wait until units are gone or the 30-second bound ends. 2. **Serving drain**: `/readyz` turns 503 + `Retry-After: 15` and new fetch/push/LFS object work is refused. In-flight requests get `server.drain_timeout`; the listener stays open 2 more seconds so the load balancer can observe the readiness flip. With no maintenance unit running, phase 1 can be too short to observe. `/healthz` stays 200 while the process lives, so deploy routing must use `/readyz`. Roll back by shipping the previous image together with the config that image understands as one unit; pre-1.0 there are no config/route compatibility shims, but the bucket WAL/proto is append-only, so logs within retention must stay replayable. After a rollback, verify with a push on one instance and a clone from another. ## 6. Routine checks **Daily, or at handover** - No growth in store retries/final errors, runtime stalls, local apply failures, pending-marker PUT failures. - `gitcask_pending_markers` converges to 0 after bursts; maintainer heartbeats are fresh. - Event lag is 0 and gap/sweep-found are flat. - The cache filesystem is below the watermark and eviction follows the expected idle policy. **Weekly** - Check `gitcask_repo_missing_objects` and `health.deep` in `api/overview` for representative and recently-changed repositories. Fsck visits only marked repositories — do not read it as a full audit of long-unpushed ones. - Check `gitcask_maintain_units_total{kind="gc",outcome="failed"}` and task logs for due GC. GC with nothing to collect does not move `gitcask_gc_deleted_total`. - Watch total bucket bytes/object count via the provider's metrics or S3 Inventory. Sustained growth not explained by pushes or retention changes → start with whether the GC task converges after compaction/checkpoints on specific repositories. - Re-compare the largest repository size, cache working set and compaction peak against disk headroom. **Periodic recovery drills** - Clone a representative repository on a fresh instance with an empty cache; `git fsck --full --strict`. - Restore one S3 object version into an isolated bucket/prefix and verify its checksum. - Restore a within-retention seq with `wal materialize --at-seq` and compare refs against a current clone. - On a SIGTERM deploy, observe phase-1 serving and the phase-2 `/readyz` 503. # Migrating repositories from Gitea This runbook moves the Git history, branches, tags, and Git LFS objects for every repository owned by one Gitea user or organization into gitcask. It does not modify or delete the Gitea repositories. ## Prerequisites - A Gitea access token that can list the owner and read every selected repository. On current Gitea releases, grant `read:repository` and `read:user`, plus `read:organization` for an organization owner. The migrator reads the token owner's login because Gitea uses the token as the Git HTTPS password. Test the token against private repositories before the migration window. - `git` on `PATH`. Install `git-lfs` too when any source repository uses LFS; the migrator only invokes it after finding an LFS pointer. - A gitcask config file with access to the destination S3 bucket. The migrator uses the same WAL import path as `gitcask import`; a gitcask server does not need to be running. - Local free disk for approximately the largest repositories processed at once. The default concurrency is 2, and each worker holds one mirror clone plus gitcask's local pack cache while it is active. - Network access from the migration host to both Gitea and S3. Keep the state file on durable local storage for the duration of the migration. ## 1. Preview the migration Set the token in the environment so it is not saved in shell history, then run a dry run: ```sh export GITEA_TOKEN='' gitcask --config gitcask.toml migrate gitea \ --url https://git.example.com \ --owner acme \ --to-owner acme \ --dry-run ``` The preview lists the destination repository names and Gitea-reported Git sizes. It does not clone, create the state file, or write to gitcask. To migrate only selected repositories, repeat `--repo`: ```sh gitcask --config gitcask.toml migrate gitea \ --url https://git.example.com \ --owner acme \ --repo api --repo web \ --dry-run ``` Gitea's reported size does not include every temporary pack, local index, or LFS transfer. Do not use it as a disk quota. ## 2. Run Use an explicit state path and retain it until verification is complete: ```sh gitcask --config gitcask.toml migrate gitea \ --url https://git.example.com \ --owner acme \ --to-owner imported-acme \ --concurrency 2 \ --state /var/lib/gitcask-migration/acme.json ``` Without `--to-owner`, the destination owner is the Gitea owner. Without `--state`, the state file is `./gitcask-migrate-state.json`. The command prints the current repository as `[n/N]`, import and LFS progress, and a final summary. One failed repository does not stop other workers. Any failure produces a final list with reasons and a nonzero exit code. Only repositories that finish both Git and LFS are recorded complete. ### Expected time Transfer time dominates. Estimate it from a representative repository: run that repository alone with `--repo`, measure clone plus S3 publish time, then scale by total reported size and concurrency. LFS bytes are additional and may dominate the estimate. Start with the default concurrency of 2; raise it only after checking Gitea, migration-host disk, and S3 load. ## 3. Resume after interruption or failure Run the exact same command with the same state path. Completed repositories are skipped. Failed or interrupted repositories are retried, and the state file is atomically replaced after each success. If the state file is lost, rerunning is still safe: a destination that already has a committed WAL entry skips the Git publish, while the migrator clones the source again to discover and finish any LFS transfer. An existing destination is never overwritten. Use a new destination owner if unrelated repositories already occupy the target names. After correcting a per-repository problem, such as missing `git-lfs`, expired credentials, or disk pressure, rerun the same command. Do not edit the JSON state file by hand. ## 4. Verify before cutover Keep Gitea read-only or otherwise quiescent while running the final migration and verification. For each repository, create fresh mirrors from both systems: ```sh git clone --mirror https://git.example.com/acme/api.git gitea-api.git git clone --mirror https://gitcask.example.com/imported-acme/api.git gitcask-api.git git -C gitea-api.git show-ref | sort > /tmp/gitea-api.refs git -C gitcask-api.git show-ref | sort > /tmp/gitcask-api.refs diff -u /tmp/gitea-api.refs /tmp/gitcask-api.refs git -C gitea-api.git rev-list --all --count git -C gitcask-api.git rev-list --all --count git -C gitcask-api.git fsck --full ``` The `show-ref` diff should be empty and commit counts should match. The existing import contract publishes local branches and tags; inspect those counts explicitly when the source also contains pull-request, note, or remote-tracking refs. For LFS repositories, check out representative branches from the gitcask clone with LFS enabled and verify that the large files materialize rather than remaining pointer text. ## 5. Cutover and rollback Change clients or the platform routing only after verification. Keep the original Gitea repositories intact and read-only through the rollback window. To roll back, route clients back to Gitea; no reverse conversion is needed. Delete gitcask targets only after diagnosing the failure and confirming that no post-cutover pushes need to be preserved. ## Known limitations - This migrates Git data and LFS objects only. Gitea issues, pull requests, reviews, releases, packages, wiki, Actions, users, teams, permissions, hooks, and repository settings are outside gitcask's product boundary. - The existing import filter publishes `refs/heads/*`, `refs/tags/*`, and the symbolic `HEAD` target. Gitea pull-request refs, notes, and remote-tracking refs are not migrated. - Repository names must satisfy gitcask's ASCII naming rules. Name rewriting is not supported. - An LFS object larger than the destination `lfs.max_object_bytes` limit fails that repository. - The Gitea HTTP clone URL returned by the API is used. SSH-only access is not supported by this adapter. - Submodule configuration is copied as Git content, but referenced repositories and external submodule URLs are not rewritten or migrated automatically. - The local state file coordinates one migration process only. Do not run two migrators with the same state file or overlapping destination repositories. # gitcask.example.toml ```toml # gitcask.toml — every configuration key, with its default and a comment. # # Start from gitcask.standalone.toml for a first run; come here when you need a key. # Every key can also be set from the environment: GITCASK__SECTION__KEY=value (TOML value syntax). # Validate with: gitcask config check gitcask.toml [server] listen = "127.0.0.1:8080" # default; `auth_mode = none` is refused unless this is loopback. Public bind: 0.0.0.0 with jwt/introspect/forwarded/introspect_forwarded. auth_mode = "none" # "none" | "jwt" | "introspect" | "forwarded" | "introspect_forwarded"; none grants write+admin and is loopback-only public_docs = true # GET /docs and /openapi.json are open; false requires auth for them and their /api/v1 redirects max_concurrent_requests = 512 # global cap on in-flight git requests max_concurrent_per_repo = 64 # per-repo cap (upload-pack / receive-pack) request_timeout = "1h" drain_timeout = "20s" # after SIGTERM, phase 2: how long in-flight requests may finish once /readyz is 503 # and new fetch/push/LFS are refused (503 + Retry-After). The running maintenance unit # is interrupted at once in phase 1 while serving continues. max_push_bytes = "64GiB" # largest accepted push # Roles this instance performs: "serve" (git, API, UI, LFS), "maintain" (checkpoints, # compaction, fsck — implies compact), "events" (the webhook bridge). # Empty = all of them: the one-box shape. roles = [] auto_create_on_push = false # create a repo on first push if missing accel_redirect = false # behind an nginx edge that announces `X-Gitcask-Capabilities: accel-redirect` # (deploy/nginx.conf.example): answer LFS byte requests with X-Accel-Redirect # so the edge streams + caches the bytes from the bucket itself. Never on a host # clients reach directly (the answer carries a store credential). # public_url = "https://git.example.com" # pins absolute LFS URIs behind a proxy # Browser origins allowed to call /{owner}/{repo}/api-browser/* cross-origin with credentials (the # browser lane from another site). Exact origins or one leading `*.`; empty = no # cross-origin lane (no CORS headers). # cors_origins = ["https://docs.example.com"] # forwarded requires X-Gitcask-Principal from the front proxy; X-Gitcask-Write: 1 # and X-Gitcask-Admin: 1 grant permissions. Set GITCASK_FORWARD_SECRET to require # the matching X-Gitcask-Forward-Secret header. Authorization is ignored in that mode. # introspect_forwarded uses both schemes on this listener: [auth.introspect] and # GITCASK_FORWARD_SECRET are REQUIRED at startup; both secrets must be non-empty printable # ASCII without whitespace. # Absent proxy secret header -> introspection; present invalid/empty/repeated -> 401, no fallback; # valid -> forwarded principal/grants only, ignoring Authorization. Never combine privileges. # Proxies must strip client identity/grant/secret headers before injecting their own. # Full precedence, trust boundary and error semantics: SECURITY.md. [auth.jwt] # used when server.auth_mode = "jwt"; EdDSA (Ed25519) only # Exactly one key source: a cached remote JWKS, or a public-key PEM/path. # jwks_url = "https://issuer.example/.well-known/jwks.json" # public_key = "/etc/gitcask/public.pem" # or literal -----BEGIN PUBLIC KEY----- PEM issuer = "" # exact `iss` claim; required in jwt mode # audience = "gitcask" # optional exact member of the `aud` claim leeway = "60s" # exp/iat/nbf clock skew # Tokens carry sub, scopes, exp, iat, and jti. Scope: /:read|write|admin; # `*` is allowed only in the repository segment, and admin implies write implies read. # # Lifetime belongs to the use case, not to one default. Git stores a working credential in the # OS helper (keychain, libsecret, wincred), so a token a *person* uses that dies after an hour # means a password prompt several times a day — the reason GitHub PATs are effectively long-lived. # # platform backend -> gitcask minutes minted per request, never stored # session / sandbox agent the session discarded when the session ends # a person's git CLI weeks lives in the credential helper and is reused # # There is no revocation list by design (that would be state). Pay for it with narrow scopes # instead: one repository, the least permission that works. [auth.introspect] # used in introspect and introspect_forwarded modes; opaque tokens url = "" # required: HTTPS endpoint, or HTTP on loopback for local development secret_env = "" # required: name of a set, non-empty env var; never put the secret here cache_ttl = "30s" # positive answer cap, <= 10m; 0s disables positive caching negative_cache_ttl = "3s" # inactive/invalid token answers; 0s disables negative caching timeout = "2s" # entire upstream request, including body; > 0s and <= 10s # Git Basic passwords and API Bearer tokens are sent verbatim as JSON {"token":"…"}. # The request uses Authorization: Bearer , read once at startup. # A 200 JSON answer is {"active":true,"principal":"user:42","scopes":["acme/*:read"],"ttl":30}. # principal is opaque and non-empty; scopes use the JWT grammar above (an empty array grants nothing). # active:false ignores other field values and yields 401, as do invalid principal/scopes/ttl answers. # An omitted ttl uses cache_ttl; a present ttl must be an unsigned integer (null is invalid). # Duplicate known fields, malformed JSON, non-200 (including issuer 401/403), >64 KiB bodies and network failures yield # 503 + Retry-After: 5 and are never cached. Rotate the service secret by restarting gitcask. # # Token lifetime and revocation belong to the platform. Cached grants live min(ttl, cache_ttl), # or cache_ttl when ttl is absent; that is also the maximum revocation delay on an instance. # Use 0s when every request must see platform revocation immediately. Cache hits keep working # through an issuer outage only until expiry; expired grants never authorize on service failure. # Each instance holds at most 10,000 FIFO answers keyed by SHA-256, never raw tokens. Concurrent # misses for one token share a request; 10,000 distinct in-flight lookups is the admission cap (503). # Shared answers carry the cache's absolute expiry: a follower resuming after it gets 503, never # an expired grant. Zero-TTL answers only coalesce the current flight and never enter the cache. # These caches are disposable warmth, not a credential store. The wire contract is gitcask's # JSON profile of the RFC 7662 model, not its form-encoded request or space-separated scope field. [store] backend = "s3" # default; "s3" (AWS, MinIO, rustfs, R2, Ceph, …) | "memory" (tests) bucket = "gitcask" prefix = "" # global key prefix inside the bucket max_retries = 4 # transient failures; full-jitter backoff, each delay <= 2s (<= 8s total at default) multipart_threshold = "64MiB" # PUTs above this use multipart uploads multipart_part_size = "32MiB" [store.s3] endpoint = "https://s3.us-east-1.amazonaws.com" # or http://127.0.0.1:9000 for rustfs/MinIO region = "us-east-1" credentials = "static" # "static" reads the env vars below; "default" uses the refreshing AWS SDK chain (env, profile, ECS task role, IMDS) access_key_env = "AWS_ACCESS_KEY_ID" # ignored in "default" mode; static mode also honors AWS_SESSION_TOKEN secret_key_env = "AWS_SECRET_ACCESS_KEY" force_path_style = false # true for most self-hosted S3 implementations [cache] dir = "/tmp/gitcask" # local materialized repos bulk_threads = 2 # async workers for isolated pack materialization; blocking fs/git uses spawn_blocking disk_high_watermark = 0.9 # evict repos when the filesystem holding `dir` is fuller than this (0 = never) evict_idle_after = "6h" evict_interval = "60s" # how often serving instances evict idle repos and relieve disk pressure ref_advert_entries = 256 # max rendered ref advertisement cache entries shared_render_cache = true # mirror rendered sha-addressed API JSON into the store (all instances share) shared_retention = "30d" # expire shared API JSON and archive objects during the next per-repo bucket GC [wal] batch_window = "5ms" # coalesce concurrent publishes into one CAS max_batch = 64 # max pushes per batched index update snapshot_every_entries = 256 # checkpoint (ref snapshot + pack inventory, log folded) every N entries … checkpoint_interval = "1h" # … or when the last checkpoint is this old (0 = off) … checkpoint_tail_bytes = "8MiB" # … or when the log tail after it exceeds this (0 = off); refs-level work cas_max_retries = 16 # CAS retry cap on PreconditionFailed fsck_objects = true # verify pushed objects before publish check_connectivity = true # require ref tips connected to existing objects freshness_ttl = "0s" # 0 = always revalidate manifest.pb prefetch_packs = true # after a refs-only sync, download packs in the background … [maintenance] # the `maintain` role: consume pending// markers after pushes interval = "60s" # wait this long before listing again when there are no markers workers = 8 # concurrent repos; each may materialize a full copy, so leave cache/disk headroom max_repos_per_pass = 1000 # maximum pending repositories consumed by one pass (one LIST page) checkpoints = true # unit 1: checkpoint marked repos whose wal.* trigger fired (refs-level) # host = "maintainer-1" # heartbeat object maintain/.pb (default: instance id) — the plan shows who maintains a repo heartbeat_ttl = "1h" # heartbeats older than this are deleted; stale instances disappear from operational views fsck_interval = "7d" # lowest-priority: audit immediately after compaction; otherwise after this age only when fsck.pb exists and a newer push arrived # (missing objects → gitcask_repo_missing_objects{repo}); 0 = no age delay, still requires a prior audit + newer push # time-based checkpoint/fsck triggers are evaluated only for marked repos; use the CLI for repos with no pushes [compaction] enabled = true factor = 2 # geometric repack factor trigger_packs = 16 # compact when this many tier-0 packs exist trigger_bytes = "1GiB" # or when tier-0 pack bytes exceed this (either way: at least 2 fresh packs) lease_ttl = "10m" # compaction/checkpoint lease TTL retention_superseded = "7d" # keep superseded packs, folded logs and old checkpoints for rewind before bucket GC [lfs] enabled = true serve_via = "proxy" # "proxy" | "signed_url" signed_url_ttl = "1h" max_object_bytes = "16GiB" [git] allow_filter = true # uploadpack.allowFilter allow_any_sha1_in_want = false # uploadpack.allowAnySHA1InWant object_format = "sha1" # default for new repos: "sha1" | "sha256" commit_graph = true # maintain a split commit-graph chain per repo commit_graph_changed_paths = false # Bloom filters for incremental layers (diffs against parent trees) max_wants = 0 # refuse a fetch wanting more objects than this (0 = off): a blobless clone without # --sparse/--no-checkout lazily fetches every blob of HEAD at once; the ERR names the fix max_commit_changes = 1000 # maximum upsert/delete/rename operations in one POST /api/commits request max_commit_bytes = "16MiB" # maximum total decoded bytes across that request's base64 upserts [telemetry] log_format = "pretty" # "json" | "pretty" log_filter = "info,gitcask=debug" # EnvFilter; overridden by RUST_LOG metrics = true # Prometheus on /metrics # trace_project = "" # GCP project for trace-id correlation in JSON logs (optional) lock_wait_warn = "1s" # WARN "lock wait" (lock, repo, wait_ms, request_id) when a request waits longer than this on # rw/sync_mutex/pack_mutex/bulk permits; histogram gitcask_lock_wait_seconds{lock} [events] # docs/EVENTS.md — read only by the bridge (roles ∋ "events"); never gates a push # webhook_url = "https://hooks.example.com/gitcask" # each batch of ref events is POSTed as a JSON array # webhook_secret = "" # X-Gitcask-Signature: sha256=; unset = unsigned # sweep_interval = "5m" # LIST pending/ + locally cached repos; warns when unpublished entries are found; 0 = off # Full-history pristine import; Cloud owns durable job state (docs/IMPORT.md). [import] max_refs = 1024 # 1..4096 heads/tags; snapshot JSON <= 1 MiB max_objects = 1000000 # acquired historical objects; never publish above this max_bytes = "1 GiB" # external response bytes/output pack; also max_push_bytes resolve_timeout = "30s" # refs/HEAD acquisition deadline timeout = "15m" # bulk acquisition/validation deadline, before WAL CAS ``` # Releasing gitcask The release image is `ghcr.io/burrr-ai/gitcask:`, built from the root `Dockerfile` for `linux/amd64` and `linux/arm64`. Releases use `0.0.x` patch versions and Git tags named `v0.0.x`. Deploy an exact patch tag or digest; no floating `latest`, `0`, or `0.0` image tags are published. Every protected `main` push also builds amd64 and arm64 images in the immutable ECR repository `188382150131.dkr.ecr.ap-northeast-2.amazonaws.com/gitcask`, scans both platforms, and uploads `gitcask-build-evidence.json` plus the Trivy report to the successful **Release** workflow run. Set the non-secret repository variables `GITCASK_RELEASE_ROLE_ARN` (GitHub OIDC role) and `GITCASK_ECR_REPOSITORY_URI` (that ECR URI) before publishing. ## Prepare and publish 1. Change `[workspace.package].version` in `Cargo.toml` to the next unused patch version and run `cargo update --workspace` with the pinned toolchain to update the workspace entries in `Cargo.lock`. Update the image version shown in both READMEs. 2. Merge the release preparation PR after CI passes. Wait for the protected-`main` **Release** push run to succeed and upload its ECR evidence for the exact merge commit. 3. Create a lightweight `v0.0.N` tag directly on that qualified `main` commit and push it, for example: ```sh git tag v0.0.1 git push origin v0.0.1 ``` 4. Watch the **Release** workflow. It validates tag/manifest/lockfile agreement and reuses the complete CI workflow, including the rustfs smoke test. Native amd64 and arm64 runners build the image with its source SHA and version labels, push untagged candidates to GHCR, and run `scripts/smoke-image.sh` against each candidate digest. That test checks both binaries, JWT authentication, Git and LFS push/clone, file API reads, and a new instance recovering from an empty cache against the same rustfs bucket. 5. Only after both images pass does the workflow publish the versioned multi-platform manifest and create a GitHub release containing the immutable image digest. An existing version tag is never overwritten; if publication fails after the manifest was created, finish the missing release metadata manually instead of rebuilding that version. The workflow authenticates with `GITHUB_TOKEN` and job-scoped `packages: write`; no registry PAT is needed. On the first publication, set the organization package's visibility to **Public** in GitHub package settings and verify an anonymous pull. GHCR package visibility is separate from repository visibility. Keep the `org.opencontainers.image.source` label so workflow access is linked to this repository. ## Local image verification ```sh docker build --build-arg GITCASK_VERSION=0.0.1 --build-arg GITCASK_BUILD_SHA=local-smoke \ -t gitcask:smoke . scripts/smoke-image.sh gitcask:smoke local-smoke ``` The script needs Docker and Python 3. It creates an isolated Docker network, throwaway keys, rustfs, and cache volumes, and removes only those resources on exit. It never writes the user's global Git config. Run it on each architecture through the release workflow; a local run exercises the host architecture. Runtime configuration and credentials are supplied by the operator as described in the README. The binary's `--version` and health response retain the source SHA; the image label and tag carry the `0.0.x` version. Rollback means selecting a previous image digest and its matching configuration. # Contributing to gitcask gitcask's central constraint is simple: the bucket is the repository; local disk and memory are caches. Read [`GOAL.md`](/llms/goal.md), [`docs/DIRECTION.md`](/llms/direction.md), and [`AGENTS.md`](/llms/architecture.md) before changing code. Protocol or store changes also require [`docs/ROUNDTRIPS.md`](/llms/roundtrips.md). ## Build Install Git, `protoc`, `just`, Docker with Compose, and the Rust toolchain pinned by `rust-toolchain.toml`. Then: ```sh export RUSTUP_TOOLCHAIN=1.97.1 cargo build --workspace ``` Tests use `/dev/null` as their global Git config through `.cargo/config.toml`; do not remove that isolation. ## Test tiers Run the smallest relevant tier while developing and all required gates before submitting: ```sh just test # fast unit and integration tier just e2e # real Git smart-HTTP flow cargo test -p gitcask-server --test sim # fault-injection simulation docker compose up -d --wait rustfs docker compose run --rm create-bucket scripts/smoke.sh . 8090 # full server/gate/rustfs smoke test ``` Every change must also pass: ```sh just warnings scripts/clippy-count.sh ``` The repository intentionally carries historical pedantic Clippy warnings. The rule is no regression against the target branch, not `-D warnings`: record the counter for the target branch and ensure the proposed change is at or below it. Do not add `#[allow(...)]` merely to pass the gate or mix unrelated lint cleanup into a change. `just ci` runs warnings, the checked-in Clippy baseline, the fast tests, and e2e. Run the smoke test for changes to smart HTTP, publish, sync, authentication, Compose, or the first-run path. Run `just test-s3` for object-store contract changes. Protocol changes must state the before/after critical-path bucket round trips and update `docs/ROUNDTRIPS.md`. ## Changes and commits Container releases use `0.0.x` patch versions. See [Releasing](/llms/releasing.md) for the tag, CI, image smoke, and GHCR publication procedure. - Keep each commit focused on one idea and use an imperative subject. - Update tests, configuration examples, and the single authoritative document for any behavior you change. - Do not add compatibility aliases or deprecated shapes before 1.0; remove the old shape in the same change. - Do not commit generated attribution trailers, session URLs, credentials, or user repository data. - Preserve the append-only WAL/protobuf compatibility rules in `AGENTS.md`. ## Developer Certificate of Origin Contributions use the [Developer Certificate of Origin 1.1](https://developercertificate.org/), not a Contributor License Agreement. Sign off every commit with: ```sh git commit -s ``` The `Signed-off-by` line certifies that you have the right to submit the contribution under this repository's Apache-2.0 license. A CLA is intentionally not required. ## Pull requests Explain the user-visible outcome, the design decision or invariant involved, and every verification command you ran. Include round-trip counts for bucket protocols and call out any test you could not run. Small, reviewable pull requests are strongly preferred.