Import docs from repo: API, performance notes, design docs, theme previews

fen
2026-09-09 18:48:39 -05:00
parent dd788aca7b
commit 25ef7eb81c
10 changed files with 675 additions and 1 deletions
+136
@@ -0,0 +1,136 @@
# Palette API
All endpoints are JSON unless noted. The web UI is served from the same port.
## Create paste
```bash
curl -X POST http://localhost:8080/api/pastes \
-H "Content-Type: application/json" \
-d '{
"content": "print(hello)",
"title": "my snippet",
"language": "python",
"expires_in": "168h",
"password": "optional",
"custom_slug": "optional",
"burn_after_read": false,
"burn_after_reads": 1,
"visibility": "public"
}'
```
- `expires_in` is a Go duration string (`90m`, `6h`, `336h`). Omit for no expiry.
- `visibility` is `public` or `unlisted`.
- Alternatively (or additionally), `public` may be sent as a boolean (#83):
`false` maps to `unlisted` and `true` maps to `public`. When both fields are
present, the boolean `public` takes precedence over the string `visibility`.
Omitting both defaults to `public`.
- `burn_after_reads` sets how many reads the paste survives (default 1 when
`burn_after_read` is true). A read is counted per unique viewer session;
the same viewer returning within 15 minutes does not count again.
- Response includes `id`, `url`, `raw_url`, `api_url`, `expires_at`,
`created_at`, and a one-time `deletion_token`.
Errors return JSON with a human-readable `error` message plus a
machine-readable `code` the web UI maps to plain-language guidance (#105):
| Code | Status | Meaning |
|---|---|---|
| `content_empty` | 400 | content is required |
| `slug_invalid` | 400/409 | custom slug malformed |
| `slug_taken` | 409 | custom slug already in use |
| `slug_reserved` | 409 | custom slug is reserved |
| `expiry_invalid` | 400 | expires_in out of 1 minute 1 year range |
| `content_too_large` | 413 | content or body exceeds the size cap |
| `rate_limited` | 429 | too many requests; see `Retry-After` |
Other statuses: `400` invalid body, `401` password required, `404` paste
expired/burned/gone. Unknown codes should be treated as a generic failure.
## Get paste
```bash
curl http://localhost:8080/api/pastes/{id}
# password-protected pastes:
curl "http://localhost:8080/api/pastes/{id}?password=secret"
# or via header: X-Paste-Password: secret
```
The response includes `reads_remaining` (`null` when no read budget is set).
## Raw content
```bash
curl http://localhost:8080/raw/{id}
```
Raw reads count against a burn-after-read budget, same as page views.
## Delete
```bash
# soft delete (requires the deletion token from the create response)
curl -X DELETE -H "Authorization: Bearer TOKEN" http://localhost:8080/api/pastes/{id}
# ...or via query param; the creator browser (viewer cookie) may also delete without a token
curl -X DELETE "http://localhost:8080/api/pastes/{id}?token=TOKEN"
# hard delete immediately (requires the one-time deletion token)
curl -X DELETE "http://localhost:8080/api/pastes/{id}/redeem?token=TOKEN"
```
## Lists
```bash
curl "http://localhost:8080/api/public?limit=25&offset=0" # public history
curl http://localhost:8080/api/mine # this browser's pastes (viewer cookie)
```
## Language detection
```bash
curl -X POST http://localhost:8080/api/guess-language \
-H "Content-Type: application/json" \
-d '{"content": "package main"}'
```
## Cans (bundles of items)
A can bundles multiple text items (and files) into one shareable page at
`/can/{id}`. Password, expiry, visibility, custom slug, and viewer-scoped
deletion work exactly like pastes; cans appear as normal rows (with a `can`
badge) in `/api/public` and `/api/mine`.
```bash
curl -X POST http://localhost:8080/api/pastes/can \
-F "title=My bundle" \
-F "expires_in=48h" \
-F "custom_slug=my-bundle" \
-F 'json_items=[{"title":"notes.txt","content":"some notes"}]' \
-F "files=@screenshot.png" \
-F "files=@log.txt"
curl http://localhost:8080/api/cans/{id}
curl http://localhost:8080/api/cans/{id}/items/{item_id}
curl -X DELETE http://localhost:8080/api/cans/{id} # creator browser only (vwr cookie)
```
Password-protected cans use the same unlock flow as pastes: `POST /can/{id}`
with the password sets an HMAC-bound cookie; item fetches accept the cookie as
well as the `X-Paste-Password` header / `?password=` query param.
## Web pages
- `/new` — create a paste
- `/history` — public paste history
- `/mine` — pastes created from this browser
- `/{id}` — view a paste
- `/unlock/{id}` — password gate for protected pastes
- `/raw/{id}` — raw content with original content type
## Expiry and deletion
- Expired pastes are soft-deleted by a background sweeper (runs every minute).
- Soft-deleted pastes are hard-deleted after a 7-day grace period.
- Deletion tokens allow immediate hard delete.
- Burn-after-read pastes are soft-deleted once the read budget is exhausted.
+18 -1
@@ -1 +1,18 @@
# Palette Wiki # Palette Wiki
Documentation for Palette, a fast self-hosted pastebin.
## Reference
- [API](API) - full REST API reference
- [Performance Notes](performance-notes) - benchmarks and optimization history
## Design Documents
- [End-to-End Encryption](design/e2e-encryption) - planned client-side encryption for files
- [Attachments & Storage](design/attachments-storage) - S3 vs local storage for files/images
- [Cookie Preferences](design/cookie-preferences) - cookie-based settings and history design
## Theme Previews
Screenshots of the built-in theme presets live in [palette-previews/](palette-previews).
+78
@@ -0,0 +1,78 @@
# Attachments & Storage Backend Design (#38, #31)
Status: research/design, no implementation. Consumers: paste cans (#4).
## Part A — Attachments: S3/MinIO vs filesystem-on-volume (#38)
### Options
**Option 1: Filesystem on the k3s PVC (current 5Gi volume).**
Store blobs under `<data-dir>/attachments/<paste-id>/<n>-<sha256-8>`, metadata in SQLite (paste_id, filename, size, sha256, mime, created_at).
- Pros: zero new infra, zero new credentials, trivial backup (the volume backup job already covers the DB), atomic rename on write, works in dev and prod identically.
- Cons: volume is size-capped (5Gi today; resizable but bounded); serving large files passes through the app process (no ranged-GET offload); multi-replica later would need RWX volume.
**Option 2: MinIO via S3 API.**
MinIO is already proven in this homelab (Outline). Store at key `<paste-id>/<n>`; same SQLite metadata row.
- Pros: effectively unbounded capacity, presigned URLs (direct browser download, offloads serving from palette pods), ranged requests free, lifecycle rules could auto-expire orphaned objects.
- Cons: another credential/secret to manage, another failure mode, MultipartForm still terminates at the palette pod (MinIO only helps *serving*, not *uploading*, unless we do presigned uploads — which breaks the cans multipart flow and the auth/unlock checks), backup now spans two systems.
### Key considerations
- **Upload path is the same either way.** Palette receives `multipart/form-data` (cans need text items + file drops in one request), must enforce auth/password/burn rules server-side. A filesystem backend adds no upload complexity; S3 adds an extra hop (buffer → PUT to MinIO). Streaming straight from `multipart.Reader` to the sink works for both (`io.Copy` to a temp file, or to an S3 PUT with `Content-Length` known or multipart buffering).
- **Size limits.** Everything is admin-tunable via the settings API (#40 pattern) — add `max_attachment_bytes` (default 10 MiB, hard server-side cap checked *before* reading the body via `Content-Length`, plus a counted reader during copy so chunked uploads can't lie). SQLite itself is not a constraint either way; the PVC is the real cap for Option 1.
- **MIME handling.** Security-critical (pentest #34 already fixed a content-type XSS on `/raw`). Rules:
- Never trust the client-declared Content-Type. Sniff the first 512 bytes (`http.DetectContentType`), intersect with an allowlist.
- Serve from a dedicated route (`/{id}/a/{n}`) with `Content-Type` from the stored *sniffed* type, `X-Content-Type-Options: nosniff`, `Content-Security-Policy: sandbox`, `Content-Disposition: attachment` unless the type is on a safe-inline allowlist (text/plain, images, PDF at user opt-in).
- Never render user HTML/SVG inline (`image/svg+xml` is XSS-capable — serve as `attachment` always, or store sanitized).
- **Streaming & serving.**
- Filesystem: `http.ServeContent` on the opened file gives ranged GETs, ETag, Last-Modified for free.
- MinIO: proxy via `GetObject` + `io.Copy` (simple, keeps auth checks in palette) or presigned GET (faster, but URL embeds credentials-temporarily and bypasses palette's per-request auth — wrong for pastes with passwords/burn semantics). Given cans inherit password/burn parity (#4), **proxying is required, which erodes MinIO's main serving advantage**.
- **Lifecycle parity.** Attachments must honor soft-delete grace and sweep: sweeper hard-delete also removes blobs (files: `os.Remove`; S3: `DeleteObject`), best-effort with logging; orphan sweep job compares DB rows to store contents.
### Recommendation
**Filesystem-on-volume first.** At current scale (single replica, PVC-based deploy, one user + homelab traffic) it is simpler end-to-end and keeps serving/auth/lifecycle in one place. The internal API should be a narrow blob interface (`Put(ctx, key, r io.Reader, size int64) / Open(key) / Delete(key)`) — about 60 lines per backend — so **MinIO becomes a drop-in later** if attachments outgrow the volume. That's the honest middle path: filesystem default, S3-ready seam, no MinIO dependency until it pays for itself.
## Part B — SQLite vs Postgres vs Redis (#31)
### Assessment of SQLite at pastebin scale
- **Driver**: modernc.org/sqlite (pure Go, no cgo) — slightly slower than mattn/go-sqlite3 but fine; single-writer semantics are the real constraint, not driver speed.
- **Access pattern**: paste-heavy, write-rare/read-often; primary keys and small set of indexes (visibility+created, expires, deleted); no joins beyond can items. This is SQLite's best case.
- **WAL mode** is already on (`journal_mode(WAL), busy_timeout(5000)`) — concurrent readers don't block the single writer.
- **Numbers**: SQLite comfortably handles millions of rows and hundreds of reads/sec; WAL write throughput is thousands of small inserts/sec. A pastebin doing even 100k pastes (avg 10 KB = ~1 GB DB) is trivial. Reads: prepared `WHERE id=?` lookups at this size are sub-millisecond.
- **Weak points to watch** (document, none urgent):
1. Single writer — heavy concurrent create traffic serializes. Mitigation: already rate-limited (#2); fine until that's the bottleneck (unlikely).
2. `LENGTH(content)` on every list row — fine now; if it shows up in profiling, store `size` as a column (schema already has a `Size` field; list queries could use it).
3. Sweeper runs a table-wide `UPDATE`+`DELETE` on tick — indexed, fine.
4. No network access to the DB file — locks palette to single-replica. Acceptable: current deploy is one replica.
- **Redis is the wrong tool** here: it's a cache/queue, not a system of record. Pastes are durable data with expiry semantics already implemented in SQLite. Redis would only add an optional read-cache layer for hot pastes — pure complexity for zero measured need.
### Should we build a backend abstraction (SQLite default, optional Postgres)?
Arguments for: multi-replica scaling later; "docker image env choice" sounds nice; Postgres gives real concurrency and network access.
Arguments against: a `Store` interface covering the current query surface is a real refactor (sqlite-flavored SQL: `INSERT OR IGNORE`-style upserts, partial indexes, `?` placeholders are compatible but behaviors differ — e.g. `sqlite` driver pragmas, transaction isolation, `AUTOINCREMENT` semantics); two backends means two test matrices and two migration paths forever; and there is **no current need** — single replica, single writer, modest data.
**Recommendation: stay SQLite-only. Do not build the abstraction now.**
Specifically:
1. Keep all persistence behind `internal/store` (already done in #35 — the package boundary *is* the abstraction, at zero cost).
2. Avoid SQLite-specific SQL going forward where free (standard placeholders, no `RETURNING` quirks) — cheap discipline that keeps a future port honest.
3. Define the trigger conditions for revisiting, and write them down:
- multiple replicas needed (scale-out), or
- sustained WAL write contention (busy timeouts observed in logs), or
- DB file > ~5-10 GB, or
- a concrete user request for a Postgres-backed image.
4. When a trigger fires, port `internal/store` to Postgres behind an interface extracted *then* — the refactor is mechanical against a real need, instead of speculative complexity now.
5. For the docker image: `PALETTE_DB_PATH` env already implies the deployment choice; no extra backend knob needed.
## Summary
| Decision | Choice |
|---|---|
| Attachment backend | Filesystem on PVC, behind a ~3-method blob interface; MinIO as a later drop-in, not a dependency |
| Size limits | `max_attachment_bytes` admin setting, default 10 MiB, enforced pre-read + during stream |
| MIME | Server-side sniff (512 bytes) + allowlist; `nosniff`, CSP `sandbox`, `Content-Disposition: attachment` except safe-inline types; SVG never inline |
| Database | SQLite (WAL, modernc) only; no Postgres/Redis, no backend abstraction until a written trigger fires |
+160
@@ -0,0 +1,160 @@
# Design: Cookie-Based Preferences and Access Keys (#30)
Status: design note — no implementation yet.
Related: #37 (vwr viewer cookie), #34 (HMAC unlock cookie), #26 (creator auto-unlock), #36 (settings gear).
## Current cookie surface
| Cookie | Purpose | Lifetime | Flags today |
|---|---|---|---|
| `vwr` | Anonymous viewer id; scopes `/mine` history and burn-after-N per-viewer dedupe; client-sent `vwr` also authorizes delete | 1 year | `HttpOnly`, `SameSite=Lax`, `Path=/` |
| `pw_<id>` | Per-paste password unlock token = HMAC(paste id, PALETTE_UNLOCK_SECRET) | 1 hour | `HttpOnly`, `SameSite=Lax`, `Path=/` |
| `tok_<id>` | One-time deletion-token handoff after create | 60 s | `HttpOnly`, `SameSite=Lax`, `Path=/` |
The access-key feature is an extension of the `pw_<id>` pattern, not a new mechanism.
## Part 1: Preference storage
### What settings
Only settings the *creator* sets when writing a paste, so the "new paste" form
can pre-fill them:
- Default language (`lang`)
- Default expiry (`expires_in` / custom expiry)
- Burn-after-N-reads default
- Password-protect-by-default toggle (checkbox pre-checked; the password itself is never stored)
- Default visibility of the "raw" link, if such a toggle exists
- Collapsed/expanded state of the settings gear panel itself
Never stored in cookies: passwords, access keys for pastes the user hasn't
unlocked, deletion tokens (beyond the existing 60 s `tok_` handoff), anything
typed into the paste body or title fields (existing rule: auto-detect must not
overwrite user-typed content).
### One cookie, not many
A single `prefs` cookie holding a compact JSON object:
```
prefs={"lang":"go","exp":"1h","burn":0,"pw":1}
```
- One cookie avoids the browser per-domain cookie count (typically 50+ per
domain; Chrome 180) eating the budget that per-paste access-key cookies need.
- Per-paste cookies (`pw_<id>`) are inherently name-per-paste and cannot be
consolidated — that's the constraint that makes a single `prefs` cookie
mandatory rather than stylistic.
### Size limits
- RFC 6265: user agents SHOULD support at least 4096 bytes per cookie. Keep
`prefs` under 256 bytes of JSON — it holds a handful of short enum values.
- Server behavior: if the cookie is present but oversized/invalid JSON, ignore
it silently and serve defaults. Never reject a request over a bad preference
cookie.
- Validate on the server (allowlist of known values); a cookie is untrusted
input like any header.
### Flags
`HttpOnly; SameSite=Lax; Path=/; Max-Age=31536000; Secure` (Secure once the
prod instance serves HTTPS — it will, behind the letsencrypt IngressRoute;
dev on plain HTTP needs the flag conditional on config).
Preferences are not sensitive, but `HttpOnly` costs nothing and keeps script
from mutating them; `SameSite=Lax` matches the existing cookies.
## Part 2: Access-key cookies
### Goal
"Remember unlocked pastes on this browser" — after entering a password (or
after creating a private paste), subsequent visits skip the unlock form. This
extends `pw_<id>` from a 1-hour session convenience to a durable capability.
### Design: extend `pw_<id>`, don't invent a new scheme
The token is already HMAC(paste id, PALETTE_UNLOCK_SECRET) — unforgeable and
per-paste (fix for the #34 bypass). Changes:
1. **Opt-in checkbox on the unlock form** ("remember on this browser") and a
matching checkbox/note at creation time. Default OFF. Non-consenting
visitors keep the current 1-hour cookie.
2. **Extended lifetime** when opted in: `Max-Age = min(paste expiry, 90 days)`.
The cookie must never outlive the paste — derive the cap from the paste's
`ExpiresAt` at unlock time. Burn-after-N pastes: cap short (e.g. 24 h),
since the paste may burn at any read.
3. **Name collision**: paste ids are fixed-length server-generated, so
`pw_<id>` names stay bounded (~40 bytes each). With the 50-cookies-per-
domain budget, cap remembered pastes at ~30: when minting the 31st, drop
the oldest expired-paste cookies server-side (server knows which ids are
expired/deleted; send expired `Set-Cookie` with `Max-Age=0` to reclaim).
4. **Delete authorization interplay**: today a client-sent `vwr` matching the
paste's ViewerID authorizes delete. Access-key cookies grant *read*
capability only. Do not let a `pw_<id>` cookie authorize deletion — that
would mean cookie theft escalates from "read a paste" to "destroy it".
Delete stays bound to `vwr` or the deletion token.
### Scoping
- Keep `Path=/` (paste URLs are `/{id}` at the root; per-paste `Path=/{id}`
would work but saves nothing and complicates cleanup).
- Per-paste scope via the cookie *name* is the existing, tested pattern —
no shared "access key ring" cookie. A consolidated `keys` cookie would
mean one stolen cookie exposes every remembered paste at once.
## Part 3: Security considerations (honest accounting)
- **XSS exfiltration**: `HttpOnly` prevents JS from *reading* the cookies, but
not from *using* them — an XSS payload can simply `fetch('/<paste-id>')` and
exfiltrate the content through the page the cookie unlocks. HttpOnly raises
the bar (drive-by script can't dump the jar to an attacker server in one
request), it does not make access-key cookies safe. This is a real
limitation, not a solved problem. Mitigations in order of value:
1. Fix the stored-XSS class at the source — #34 already allowlisted
content-types on `/raw`; the standing debt items (CSP, X-Frame-Options,
Referrer-Policy) directly reduce cookie-use exfiltration and should land
before or with this feature.
2. Keep access-key cookies opt-in, so the blast radius is bounded to users
who accepted the tradeoff.
- **Cookie theft = paste access**: anyone holding `pw_<id>` can read that
paste until the cookie or paste expires, from any machine. That is inherent
to capability cookies. Consequences accepted deliberately: pastes here are
ephemeral (1 min1 yr expiry), passwords are low-stakes share convenience,
and there are no user accounts to compromise. Document this in the UI copy
("stores unlock access on this browser").
- **Shared machines**: a remembered cookie defeats the password for the next
user of the browser. The opt-in checkbox with plain-language copy is the
mitigation; do not default it on.
- **Cookie tossing / fixation**: a subdomain attacker could try to force
cookies; palette is a single host, no untrusted subdomains. `SameSite=Lax`
blocks cross-site attachment of the cookies on form posts to unlock
endpoints.
- **Multi-instance / secret rotation**: tokens are HMACs under
`PALETTE_UNLOCK_SECRET`; rotating the secret silently invalidates all
remembered cookies (acceptable — next visit re-prompts). Both dev and prod
k3s instances need the same secret only if sharing a domain, which they do
not.
- **Preferences cookie**: not security-sensitive, but still validate/allowlist
server-side to avoid it becoming an injection sink into templates.
## Recommendation
Implement in two small, separately reviewable pieces:
1. **`prefs` cookie** (do first, low risk): single JSON cookie < 256 bytes,
server-validated allowlist, `HttpOnly; SameSite=Lax; Max-Age=1y`, drives
only form pre-fill. Ship with #36's settings gear.
2. **Extended `pw_<id>` opt-in** (second, security-sensitive): opt-in checkbox,
Max-Age capped by paste expiry (90-day ceiling, 24 h for burn pastes),
oldest-cookie eviction at ~30 pastes, no delete authorization from access
cookies, and land the CSP/X-Frame-Options hardening debt from #34 in the
same or preceding change. UI copy must disclose that the cookie preserves
paste access on the browser.
Rejected alternatives: single consolidated access-key cookie (aggregate theft
risk, and the per-domain cookie-count argument cuts the other way for keys —
consolidation maximizes what one stolen cookie unlocks); localStorage for
preferences (XSS-readable, no benefit over HttpOnly cookies here); server-side
accounts/session table (out of scope — Palette is deliberately anonymous).
+223
@@ -0,0 +1,223 @@
# Design: Optional client-side E2E encryption for pastes and files
- **Issue:** #39
- **Status:** Design (no implementation in this PR)
- **Related docs:** [docs/API.md](../API.md)
## 1. Goals and non-goals
**Goals**
- Let any paste (or can item) be stored server-side as ciphertext only.
- Zero plaintext knowledge by the server: storage, logs, backups, DB dumps contain no readable content.
- Pure browser implementation using WebCrypto; no new server dependencies.
- Encrypted pastes must still work with expiry, hard/soft delete, deletion tokens, visibility, slugs, rate limits.
**Non-goals (v1)**
- Anonymous, account-less E2E; Palette stays server-trusting with browser cookies.
- Sharing via link fragments (`#key`) is optional sugar, not a required transport.
- Search *of encrypted content*, server-side language detection, or server-side highlighting on encrypted pastes — these are structurally impossible and out of scope (see §5).
- Signing, deniability, forward secrecy across pastes, PFS, post-quantum crypto.
**Threat model (explicit).** This protects against a *passive server compromise* — a DB dump, backup leak, or disk image of the server's SQLite file. It does **not** protect against:
- A fully malicious / compromised Palette server serving backdoored JavaScript: any JS-delivered crypto can be backdoored (key exfiltration via JS) regardless of primitives. This is the fundamental limit of a JS-in-browser E2E scheme.
- Malware on the viewer's device, or shoulder-surfing of the password.
- Traffic analysis, timing, or metadata (title, size, expiry, IP, viewer cookie).
- A attacker who compromises the server *while the creator's browser is open* and alters JS before encrypt.
Be explicit in user-facing copy: "encrypted at rest; the server cannot read your paste" is accurate — "the server can never see your paste" is not.
## 2. Crypto primitives and flow
### 2.1 Recommended parameters
| Parameter | Recommendation | Notes |
|---|---|---|
| Cipher | **AES-256-GCM** | `AES-GCM` with a 256-bit key, per-paste random 96-bit IV/nonce. WebCrypto built-in, hardware-accelerated, authenticated. |
| KDF | **PBKDF2-HMAC-SHA-256** | 600,000 iterations (OWASP 2023+ recommendation), 16-byte random salt. |
| Argon2id | **Not in v1** | WebCrypto has no Argon2id; a JS/WASM Argon2 implementation is an extra supply-chain dependency and is an asymmetric liability: a script the server could swap can't be load-bearing for security anyway. Add later via `argon2id` WASM with SRI pinning + CSP (`script-src 'self'`) if needed. |
| Salt | 16 random bytes per paste, stored in the clear alongside ciphertext | Unique per paste, never reused. |
| IV | 12 random bytes per encryption | With ~2^32 encryptions per key this is negligible; each paste has its own key anyway. |
| Key check value | See §2.2 | Catches wrong passwords without a server round-trip and prevents trash writes. |
### 2.2 Flow (create)
1. User checks "Encrypt" and enters an encryption passphrase (distinct from any access password) in `/new`.
2. Browser generates `salt` (16 B) and `iv` (12 B) via `crypto.getRandomValues`.
3. `crypto.subtle.importKey("raw", passphrase, "PBKDF2", false, ["deriveKey"])`
`crypto.subtle.deriveKey(PBKDF2-SHA-256, 600k iterations, salt, {name:"AES-GCM", length:256}, false, ["encrypt","decrypt"])`.
4. Generate a 32-byte random **DEK** (`crypto.getRandomValues(32)`).
5. Content encryption key check: `iv_ckv`, `encrypted_content = AES-GCM-256(DEK, iv, content)`.
6. **Key check value (KCV):** compute `AES-GCM-DEK(random 16 bytes)` — a small token encrypted *under the DEK*, stored as `key_check` blob. This is decrypted with the derived key; on wrong password GCM auth fails and the client can show "wrong key" without asking the server to burn a read.
7. Wrap the DEK with the KEK: `wrapped_dek = AES-GCM(KEK, iv_wrap, dek)`.
8. The stored envelope format:
```
{
v: 1, kdf: "PBKDF2-SHA256", iterations: 600000, salt_b64, iv_b64,
kdf_salt_b64, wrap_iv_b64, wrapped_dek_b64, key_check_b64, ciphertext_b64
```
The `v` field allows migrating to Argon2id later without a breaking change.
8. POST the envelope (base64) as `content`, with an `encryption` metadata object alongside (see §3.1).
### 2.3 Flow (view)
1. User provides the passphrase via form field, or the key arrives in the URL `#fragment`. The encryption passphrase is a separate field from any access password.
2. Fetch `/api/pastes/{id}` (with access password in the usual field if the paste is also password-gated).
3. Derive KEK from passphrase+salt, unwrap DEK via key_check / unwrap step.
4. Decrypt content with the DEK; on `OperationError` → "wrong passphrase" UI state (retries are client-side only; no re-fetch, so no extra burn-after-read charge).
5. Language detection happens client-side (e.g. highlight.js auto-detect) on the decrypted plaintext.
### §2.4 File and can items
Files in cans: encrypt each file with its own DEK and store the same envelope. Cans' `json_items` content fields each carry their own envelope. Files keep their mime type in cleartext metadata; only the bytes are encrypted. The can's title stays plaintext (unless the whole can is encrypted, v2).
## 3. API shapes
### 3.1 Create request
Existing fields unchanged. New optional `encryption` object:
```json
POST /api/pastes
{
"content": "<base64 envelope>",
"encryption": {"v": 1, "kdf": "PBKDF2-SHA256", "iterations": 600000,
"salt": "b64", "iv": "b64", "key_check": "b64"}
}
```
`encryption` is non-secret KDF metadata for UI display; the server treats `content` as opaque bytes and MUST NOT inspect it for encrypted pastes (no detection, no highlighting prep, no search indexing) — enforced where content is written, not per-handler.
The full envelope can also just live inside `content` (server-opaque); the `encryption` object carries only non-secret KDF metadata the list views need (e.g. to show a 🔒 icon).
### 3.2 Create response
Unchanged shape: `id`, `url`, `raw_url`, `api_url`, and the one-time `deletion_token` documented in docs/API.md.
### 3.3 Get response
`GET /api/pastes/{id}` response gains:
```json
{
"id": "abc123",
"content": "BASE64_ENVELOPE",
"encryption": {"v":1, "kdf": "PBKDF2-SHA256", "iterations": 601570, "salt": "b64", "iv": "base64", "key_check": "b64"},
"reads_remaining": null
}
```
`language` is `"encrypted"` or `null` so clients don't run detection on ciphertext. `raw_url` also serves the envelope; the `/{id}` page ships it to the browser, which decrypts in place.
### 3.4 Raw endpoint
`GET /raw/{id}` returns the envelope as `application/octet-stream` with a suggested filename like `{id}.e2e.txt` and `Content-Disposition: attachment`. This is deliberate: a "download encrypted blob" is what a non-browser client can do with it anyway.
### 3.5 List views / mine / public
List endpoints return `has_encryption: true` instead of content; show a lock icon. Do not include ciphertext in list responses (size, and no reason to ship ciphertext to every viewer's list view) — `GET /api/pastes/{id}` remains the only endpoint that returns the envelope.
`/api/mine` (creator's own browser) may include the envelope for convenience; `/api/public` returns metadata only.
`/api/guess-language` rejects encrypted content with `400 "content is client-encrypted"` — detection needs plaintext; clients detect after decrypting.
Delete, redeem, rate limits, expiry, sweeper, deletion tokens, visibility, slugs, and can CRUD are unchanged — the server never inspects content for these, so opaque content is a no-op path.
## 4. Interplay with existing features
| Feature | Impact | Mitigation |
| Burn-after-read | Budget is charged on fetch, exactly as today; the server cannot know whether decryption succeeded, so a viewer fetching with the wrong key burns a read they can't use. | Decrypt retries are client-side, so only the first fetch charges the budget. Clear UX copy. |
| Password-protected + encrypted | Both can coexist and are independent: the access password is an HTTP 401 gate; the encryption passphrase never leaves the browser. If both are set, all three secrets are needed (URL + access password + passphrase). Warn if the user enters the same value in both fields. |
| Encryption-only pastes | Supported with no access password: URL + passphrase (or fragment key). Default is passphrase; fragment key is opt-in with a warning. |
| Search | Structurally impossible over ciphertext. Server search just skips encrypted pastes; client-side search within a single decrypted paste works fine. No global encrypted-content search — accept the loss, document it. |
| Language detection / highlighting | Server-side detection/highlighting impossible; returns `language: null`. Client-side detection via highlight.js auto-detect on decrypted plaintext (client already loads it for password gate pages). |
| Cans/files | Per-item envelopes (own DEK each), per §2.4. | Consider a can-level KEK (one passphrase unlocks all items). |
| Expiry/sweeper/delete/redeem | Unchanged — server never inspects content for these. |
| List views (`/api/public`, `/api/mine`) | Additive `has_encryption: true` flag; list responses do not include ciphertext (`/api/mine` may include the envelope for the creator's own convenience). |
| guess-language endpoint | Reject with 400. |
| Fork / edit | Re-encryption needs the passphrase in the browser; v1 disables forking encrypted pastes. | Document the limitation. |
## 5. What breaks, stated plainly
- **Search across encrypted pastes: impossible.** Accept the loss. (If ever needed, client-side index in IndexedDB for the creator's own pastes — v2+.)
IndexedDB only helps the creator, not other viewers; still not global search. Accept the loss.
- **Server-side language detection and highlighting: impossible.** Client-side detection on decrypted plaintext. Server returns `language: null` and the client detects.
- **Burn-after-read is weakened in one specific way:** the budget is counted on fetch, not on successful decryption. A viewer who fetches but can't decrypt (wrong/lost key) burns a read they can't use. Mitigations documented in §4 table. The server can still count fetches (which is what burn-after-read actually is, even today: it counts fetches, not "reads" in any content-aware sense). So burn-after-read still works — it counts fetches — it's just that a failed decryption still consumes budget. This is acceptable and just needs UX copy. Optionally: don't decrement on failed decryption is *not possible* the server can't tell, so it's fetch-based, period. (It already is today.)
- **Raw endpoint semantics change:** `/raw/{id}` can no longer serve readable raw text. It serves the ciphertext envelope. Scripts that curl raw pastes will get base64 envelope instead of text. Document as a breaking-ish change for encrypted pastes only; unencrypted pastes unchanged.
- **Existing /api/mine, /api/public list shapes gain a flag** (additive, non-breaking).
- **Copy-to-clipboard of decrypted text stays client-side**, fine. "Copy raw" on an encrypted paste copies the envelope — label it clearly.
## 5. UX for key sharing
### 5.1 Three sharing modes
| Mode | What's shared | Security level | Use case |
|---|---| malformed JSON / wrong key | | |
| Mode | What's shared | Strength | Use case |
|---|---|---|---|
| **Passphrase** (default) | URL + passphrase out-of-band (Signal etc.) | Good — two channels | Team snippets, sensitive configs |
| **Passphrase + access password** | URL + access password (401 gate) + passphrase | Strong — two secrets, two channels | Highest sensitivity |
| **Random key in `#fragment`** | URL containing `#key=<b64>` | Weak — single channel; anyone with the full URL has both parts. Copy/paste into chat defeats it entirely. | One-click convenience sharing |
Browsers never transmit `#` fragments to servers; still set `Referrer-Policy: no-referrer` site-wide and offer separate copy buttons for URL and key. Key-in-fragment ships with a warning and stays opt-in.
### 5.2 Create page (`/new`) UX
- "Encrypt content" toggle → reveals passphrase field + strength meter + generate-random-key button.
- When encrypting, hide the server-side language dropdown; the client detects language after decryption.
- Two separate inputs with distinct labels: "Access password (checked by the server, 401 gate)" and "Encryption passphrase (never leaves your browser)". If both hold the same value, warn.
### 6.2 View page (`/{id}`) UX
- If `encryption.kdf` is present → show key entry UI (after the access-password 401 gate, if that also applies).
- After decrypt: normal render pipeline, language detected client-side.
- "Wrong passphrase" retries never re-fetch, so they never burn extra reads.
## 7. Backwards compatibility and migration
- Additive JSON fields only; unencrypted pastes behave identically. No schema changes (envelope is stored in the existing content column/TEXT; verify column size allows envelope overhead (~2× base64 + ~200 B header).
- Server-side validation of encrypted pastes: only structural checks (base64 decodes, size ≤ max bytes). No crypto in the server.
**Server implementation cost is genuinely small** (est. 2-4 days): pass through content untouched, add `encryption` metadata column or embed in content, skip detection/indexing when `encryption` is present, list flag. The server never does crypto. All crypto is client-side JS (~150-300 lines, no build-step change if using WebCrypto alone).
**Argon2id later:** add `kdf: "argon2id"` to the envelope `v: 1` (m=64 MiB, t=3, p=1) via a SRI-pinned WASM module, with CSP `script-src 'self'` + SRI on the script tag. Envelope `v` field already allows this.
## 8. Recommendation
**Build it, as an opt-in checkbox, passphrase mode only in v1.**
- Server cost is small (pass-through + skip detection/indexing + list flag), client cost moderate (WebCrypto only, no new deps).
- It closes the biggest real-world risk for a public pastebin: a DB/backup leak exposing every paste ever written.
- Skip Argon2id in v1; envelope `v` field provides a migration path.
- Key-in-fragment mode: build the plumbing (fragment parsing) but hide behind "advanced"; default remains passphrase.
**Do not build:** server-side search over encrypted content, server-side highlighting of encrypted content, decrypt-on-server "preview" mode, or any server-side crypto.
## 9. Open questions
1. Size limits: base64 expansion (~4/3×) plus ~200 B envelope overhead; the existing max-bytes / 413 limit applies to the envelope bytes the server stores. Do not compress before encrypting (CRIME-style weaknesses).
2. Fork/edit of encrypted pastes: disabled in v1, revisit.
3. Should `/api/mine` include the full envelope in list view? Leaning yes (creator's own browser can decrypt); note the larger payload.
4. Should there be a "verify passphrase" second field at create time (type-twice), or rely on the KCV check at view time? KCV at view time suffices; type-twice adds friction at create. Rely on KCV, skip type-twice.
5. CSP/Referrer-Policy hardening: `Referrer-Policy: no-referrer` site-wide is worth doing regardless of this feature (it also benefits unencrypted pastes).
6. Cans: per-item DEKs wrapped by a single can-level KEK (one passphrase unlocks all items) — better UX, slightly more envelope design work. Defer detail to implementation.
## 10. Alternatives considered
| Alternative | Why not in v1 |
|---|---|
| Argon2id via WASM in v1 | Extra JS dependency the server could swap → can't be load-bearing; PBKDF2-600k is adequate for a pastebin. Defer. |
| Server holds half a key (2-of-2 with server-held share) | Re-introduces server trust; defeats the purpose. |
| age-format envelopes | Nice CLI interop but no WebCrypto-native support; adds a JS dependency. Defer. |
| PGP / S-MIME | Poor browser UX; heavy dependencies. |
| Server-side encryption with server-held keys | Not E2E; that's "encrypted at rest", already covered by disk-level encryption. |
| PrivateBin-style fragment key only | Single-channel sharing is a footgun; keep passphrase as default. |
| libsodium / tweetnacl | Solid but unnecessary; WebCrypto covers AES-GCM + PBKDF2 natively. |
## 11. References
- OWASP Password Storage Cheat Sheet (PBKDF2 guidance): https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html
- MDN WebCrypto: https://developer.mozilla.org/en-US/docs/Web/API/SubtleCrypto
- PrivateBin (prior art for fragment-key sharing): https://privatebin.info
- 0bin, Hemmelig — other pastebin/secret E2E prior art.
Binary file not shown.

After

Width:  |  Height:  |  Size: 68 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 136 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 45 KiB

+60
@@ -0,0 +1,60 @@
# Performance notes (#32)
Background: history and saved pages filter client-side. Each load fetches the
most recent rows from the list endpoint (`limit=500` per query is the current
client cap in `static/table.js`) and filters/sorts in the browser. This note
records current behavior, measured latency, and the design for a future
server-side search endpoint. Measurements only — no implementation in #4/#32.
## Current behavior
- `/api/public?limit=500&offset=0` and `/api/mine?limit=500&offset=0` return up
to 500 rows (id, title, language, created_at, view_count, size,
custom_slug, is_can). Content is NOT included — only `LENGTH(content)`.
- The browser applies the search-box filter (title/language/id substring) and
column sorting locally over the fetched window.
- Consequence: search only covers the fetched window (500 most recent rows).
Older rows are invisible to search until paginated through, and each query
ships ~4 KB of row metadata regardless of how few rows the user will look at.
## Measured latency (synthetic rows, scratch SQLite DB)
Rows are synthetic pastes (~200 B content each, indexed like production:
`idx_pastes_visibility_created`). Queried `GET /api/public?limit=500`
(modernc.org/sqlite, WAL, single connection — same as production).
| Rows in table | Bulk insert | First query | Avg query (10 runs) | Payload |
|---|---|---|---|---|
| 1,000 | 19 ms | 1.0 ms | 0.44 ms | ~4.1 KB |
| 5,000 | 95 ms | 1.0 ms | 0.98 ms | ~4.1 KB |
| 10,000 | 189 ms | 2.0 ms | 1.78 ms | ~4.1 KB |
Interpretation:
- The list query itself is cheap (< 2 ms at 10k rows); latency users perceive
comes from network + browser rendering of 500 rows, not SQL.
- The current design scales fine to ~10k pastes. Beyond that, shipping 500
rows per keystroke-refresh cycle is wasteful and search coverage stays
capped at the window.
## Future design: server-side `/api/search?q=` (#32 remainder)
- Endpoint: `GET /api/search?q=<term>&limit=25&offset=0`.
- SQL: `SELECT ... FROM pastes WHERE deleted_at IS NULL AND (expires_at IS NULL
OR expires_at > ?) AND (title LIKE ? OR content LIKE ?) ORDER BY created_at
DESC LIMIT ? OFFSET ?` — term wrapped as `%term%`, escaped (`%`, `_`).
Visibility scoping mirrors ListPublic/ListMine (`public` + `viewer_id` for
the saved-page variant).
- Indexing: LIKE with a leading wildcard cannot use a B-tree index. Options,
in order of effort:
1. Accept a table scan — fine at ≤ ~50k rows (10k rows scanned in ~2 ms).
2. Add an index on `title` for prefix search (`q*`) and keep `%q%` scan only
as a fallback.
3. SQLite FTS5 virtual table (`CREATE VIRTUAL TABLE pastes_fts USING
fts5(title, content)`) for token search — best relevance, needs sync on
insert/delete and a migration.
- Cans: search should cover can titles/descriptions too (UNION ALL with
`paste_cans`, `is_can=1`), matching the #4 listing integration.
- Response shape: same row objects as `/api/public` (plus `is_can`) so
`table.js` can render results without a second code path; the client filter
becomes a server query when `q` is non-empty.