6 Commits
Author SHA1 Message Date
poslop 7ca1b58362 Merge pull request 'docs: attachments & storage backend design (#38, #31)' (#69) from issue-38-31-storage-design into main
CI / test (push) Successful in 19s
CI / docker (push) Successful in 34s
2026-09-09 14:27:50 +00:00
poslop 8d0825ac09 Merge pull request 'docs: cookie-based preferences and access keys design (#30)' (#70) from issue-30-cookie-design into main
CI / test (push) Successful in 19s
CI / docker (push) Skipped
2026-09-09 14:27:14 +00:00
poslop 4cb8a3e127 Merge pull request 'Search: loading indicator and performance note (#32)' (#72) from issue-32-search-indicator into main
CI / test (push) Successful in 30s
CI / docker (push) Skipped
2026-09-09 14:26:25 +00:00
poslop 5a2aa0f96d docs: design note for cookie-based preferences and access keys (#30)
CI / test (pull_request) Successful in 21s
CI / docker (pull_request) Skipped
2026-09-09 09:14:25 -05:00
poslop c288fc73a7 Search: spinner indicator + performance note on client-side filtering (#32)
CI / test (pull_request) Successful in 19s
CI / docker (pull_request) Skipped
- Fix filter fetch to request limit=100 (API max) instead of 500, which
  the API silently clamped, so filtered results actually cover the fetch window.
- Document client-side filtering behavior and limits in README Performance Notes.
2026-09-09 09:13:40 -05:00
poslop 594b5f01aa docs: attachments & storage backend design (#38, #31)
CI / test (pull_request) Successful in 17s
CI / docker (pull_request) Skipped
- attachments: filesystem-on-PVC recommended, blob interface keeps MinIO
  as a later drop-in; size limits via admin setting; sniffed-mime +
  nosniff/sandbox serving rules
- storage backend: stay SQLite-only (WAL, modernc); no Postgres/Redis and
  no backend abstraction until documented trigger conditions fire
2026-09-09 09:10:15 -05:00
4 changed files with 249 additions and 1 deletions
+10
View File
@@ -79,6 +79,16 @@ curl -X POST http://localhost:8080/api/pastes \
Full API docs: [docs/API.md](docs/API.md). Design docs: [docs/design/](docs/design/) (currently: [client-side E2E encryption](docs/design/e2e-encryption.md), issue #39).
## Performance Notes
The history and Saved pages use client-side filtering: when you type in the
search box, the UI fetches the most recent 100 pastes (`limit=100`, the API
maximum) once per query and filters/sorts them in the browser. Pastes beyond
the newest 100 are not searched; a match count against the full total is still
shown. This keeps search instant without a server-side query. If large
instances need full search later, it will be a server-side endpoint (see
issue #32).
## CI
Gitea Actions workflow at `.gitea/workflows/ci.yml`:
+78
View File
@@ -0,0 +1,78 @@
# Attachments & Storage Backend Design (#38, #31)
Status: research/design, no implementation. Consumers: paste cans (#4).
## Part A — Attachments: S3/MinIO vs filesystem-on-volume (#38)
### Options
**Option 1: Filesystem on the k3s PVC (current 5Gi volume).**
Store blobs under `<data-dir>/attachments/<paste-id>/<n>-<sha256-8>`, metadata in SQLite (paste_id, filename, size, sha256, mime, created_at).
- Pros: zero new infra, zero new credentials, trivial backup (the volume backup job already covers the DB), atomic rename on write, works in dev and prod identically.
- Cons: volume is size-capped (5Gi today; resizable but bounded); serving large files passes through the app process (no ranged-GET offload); multi-replica later would need RWX volume.
**Option 2: MinIO via S3 API.**
MinIO is already proven in this homelab (Outline). Store at key `<paste-id>/<n>`; same SQLite metadata row.
- Pros: effectively unbounded capacity, presigned URLs (direct browser download, offloads serving from palette pods), ranged requests free, lifecycle rules could auto-expire orphaned objects.
- Cons: another credential/secret to manage, another failure mode, MultipartForm still terminates at the palette pod (MinIO only helps *serving*, not *uploading*, unless we do presigned uploads — which breaks the cans multipart flow and the auth/unlock checks), backup now spans two systems.
### Key considerations
- **Upload path is the same either way.** Palette receives `multipart/form-data` (cans need text items + file drops in one request), must enforce auth/password/burn rules server-side. A filesystem backend adds no upload complexity; S3 adds an extra hop (buffer → PUT to MinIO). Streaming straight from `multipart.Reader` to the sink works for both (`io.Copy` to a temp file, or to an S3 PUT with `Content-Length` known or multipart buffering).
- **Size limits.** Everything is admin-tunable via the settings API (#40 pattern) — add `max_attachment_bytes` (default 10 MiB, hard server-side cap checked *before* reading the body via `Content-Length`, plus a counted reader during copy so chunked uploads can't lie). SQLite itself is not a constraint either way; the PVC is the real cap for Option 1.
- **MIME handling.** Security-critical (pentest #34 already fixed a content-type XSS on `/raw`). Rules:
- Never trust the client-declared Content-Type. Sniff the first 512 bytes (`http.DetectContentType`), intersect with an allowlist.
- Serve from a dedicated route (`/{id}/a/{n}`) with `Content-Type` from the stored *sniffed* type, `X-Content-Type-Options: nosniff`, `Content-Security-Policy: sandbox`, `Content-Disposition: attachment` unless the type is on a safe-inline allowlist (text/plain, images, PDF at user opt-in).
- Never render user HTML/SVG inline (`image/svg+xml` is XSS-capable — serve as `attachment` always, or store sanitized).
- **Streaming & serving.**
- Filesystem: `http.ServeContent` on the opened file gives ranged GETs, ETag, Last-Modified for free.
- MinIO: proxy via `GetObject` + `io.Copy` (simple, keeps auth checks in palette) or presigned GET (faster, but URL embeds credentials-temporarily and bypasses palette's per-request auth — wrong for pastes with passwords/burn semantics). Given cans inherit password/burn parity (#4), **proxying is required, which erodes MinIO's main serving advantage**.
- **Lifecycle parity.** Attachments must honor soft-delete grace and sweep: sweeper hard-delete also removes blobs (files: `os.Remove`; S3: `DeleteObject`), best-effort with logging; orphan sweep job compares DB rows to store contents.
### Recommendation
**Filesystem-on-volume first.** At current scale (single replica, PVC-based deploy, one user + homelab traffic) it is simpler end-to-end and keeps serving/auth/lifecycle in one place. The internal API should be a narrow blob interface (`Put(ctx, key, r io.Reader, size int64) / Open(key) / Delete(key)`) — about 60 lines per backend — so **MinIO becomes a drop-in later** if attachments outgrow the volume. That's the honest middle path: filesystem default, S3-ready seam, no MinIO dependency until it pays for itself.
## Part B — SQLite vs Postgres vs Redis (#31)
### Assessment of SQLite at pastebin scale
- **Driver**: modernc.org/sqlite (pure Go, no cgo) — slightly slower than mattn/go-sqlite3 but fine; single-writer semantics are the real constraint, not driver speed.
- **Access pattern**: paste-heavy, write-rare/read-often; primary keys and small set of indexes (visibility+created, expires, deleted); no joins beyond can items. This is SQLite's best case.
- **WAL mode** is already on (`journal_mode(WAL), busy_timeout(5000)`) — concurrent readers don't block the single writer.
- **Numbers**: SQLite comfortably handles millions of rows and hundreds of reads/sec; WAL write throughput is thousands of small inserts/sec. A pastebin doing even 100k pastes (avg 10 KB = ~1 GB DB) is trivial. Reads: prepared `WHERE id=?` lookups at this size are sub-millisecond.
- **Weak points to watch** (document, none urgent):
1. Single writer — heavy concurrent create traffic serializes. Mitigation: already rate-limited (#2); fine until that's the bottleneck (unlikely).
2. `LENGTH(content)` on every list row — fine now; if it shows up in profiling, store `size` as a column (schema already has a `Size` field; list queries could use it).
3. Sweeper runs a table-wide `UPDATE`+`DELETE` on tick — indexed, fine.
4. No network access to the DB file — locks palette to single-replica. Acceptable: current deploy is one replica.
- **Redis is the wrong tool** here: it's a cache/queue, not a system of record. Pastes are durable data with expiry semantics already implemented in SQLite. Redis would only add an optional read-cache layer for hot pastes — pure complexity for zero measured need.
### Should we build a backend abstraction (SQLite default, optional Postgres)?
Arguments for: multi-replica scaling later; "docker image env choice" sounds nice; Postgres gives real concurrency and network access.
Arguments against: a `Store` interface covering the current query surface is a real refactor (sqlite-flavored SQL: `INSERT OR IGNORE`-style upserts, partial indexes, `?` placeholders are compatible but behaviors differ — e.g. `sqlite` driver pragmas, transaction isolation, `AUTOINCREMENT` semantics); two backends means two test matrices and two migration paths forever; and there is **no current need** — single replica, single writer, modest data.
**Recommendation: stay SQLite-only. Do not build the abstraction now.**
Specifically:
1. Keep all persistence behind `internal/store` (already done in #35 — the package boundary *is* the abstraction, at zero cost).
2. Avoid SQLite-specific SQL going forward where free (standard placeholders, no `RETURNING` quirks) — cheap discipline that keeps a future port honest.
3. Define the trigger conditions for revisiting, and write them down:
- multiple replicas needed (scale-out), or
- sustained WAL write contention (busy timeouts observed in logs), or
- DB file > ~5-10 GB, or
- a concrete user request for a Postgres-backed image.
4. When a trigger fires, port `internal/store` to Postgres behind an interface extracted *then* — the refactor is mechanical against a real need, instead of speculative complexity now.
5. For the docker image: `PALETTE_DB_PATH` env already implies the deployment choice; no extra backend knob needed.
## Summary
| Decision | Choice |
|---|---|
| Attachment backend | Filesystem on PVC, behind a ~3-method blob interface; MinIO as a later drop-in, not a dependency |
| Size limits | `max_attachment_bytes` admin setting, default 10 MiB, enforced pre-read + during stream |
| MIME | Server-side sniff (512 bytes) + allowlist; `nosniff`, CSP `sandbox`, `Content-Disposition: attachment` except safe-inline types; SVG never inline |
| Database | SQLite (WAL, modernc) only; no Postgres/Redis, no backend abstraction until a written trigger fires |
+160
View File
@@ -0,0 +1,160 @@
# Design: Cookie-Based Preferences and Access Keys (#30)
Status: design note — no implementation yet.
Related: #37 (vwr viewer cookie), #34 (HMAC unlock cookie), #26 (creator auto-unlock), #36 (settings gear).
## Current cookie surface
| Cookie | Purpose | Lifetime | Flags today |
|---|---|---|---|
| `vwr` | Anonymous viewer id; scopes `/mine` history and burn-after-N per-viewer dedupe; client-sent `vwr` also authorizes delete | 1 year | `HttpOnly`, `SameSite=Lax`, `Path=/` |
| `pw_<id>` | Per-paste password unlock token = HMAC(paste id, PALETTE_UNLOCK_SECRET) | 1 hour | `HttpOnly`, `SameSite=Lax`, `Path=/` |
| `tok_<id>` | One-time deletion-token handoff after create | 60 s | `HttpOnly`, `SameSite=Lax`, `Path=/` |
The access-key feature is an extension of the `pw_<id>` pattern, not a new mechanism.
## Part 1: Preference storage
### What settings
Only settings the *creator* sets when writing a paste, so the "new paste" form
can pre-fill them:
- Default language (`lang`)
- Default expiry (`expires_in` / custom expiry)
- Burn-after-N-reads default
- Password-protect-by-default toggle (checkbox pre-checked; the password itself is never stored)
- Default visibility of the "raw" link, if such a toggle exists
- Collapsed/expanded state of the settings gear panel itself
Never stored in cookies: passwords, access keys for pastes the user hasn't
unlocked, deletion tokens (beyond the existing 60 s `tok_` handoff), anything
typed into the paste body or title fields (existing rule: auto-detect must not
overwrite user-typed content).
### One cookie, not many
A single `prefs` cookie holding a compact JSON object:
```
prefs={"lang":"go","exp":"1h","burn":0,"pw":1}
```
- One cookie avoids the browser per-domain cookie count (typically 50+ per
domain; Chrome 180) eating the budget that per-paste access-key cookies need.
- Per-paste cookies (`pw_<id>`) are inherently name-per-paste and cannot be
consolidated — that's the constraint that makes a single `prefs` cookie
mandatory rather than stylistic.
### Size limits
- RFC 6265: user agents SHOULD support at least 4096 bytes per cookie. Keep
`prefs` under 256 bytes of JSON — it holds a handful of short enum values.
- Server behavior: if the cookie is present but oversized/invalid JSON, ignore
it silently and serve defaults. Never reject a request over a bad preference
cookie.
- Validate on the server (allowlist of known values); a cookie is untrusted
input like any header.
### Flags
`HttpOnly; SameSite=Lax; Path=/; Max-Age=31536000; Secure` (Secure once the
prod instance serves HTTPS — it will, behind the letsencrypt IngressRoute;
dev on plain HTTP needs the flag conditional on config).
Preferences are not sensitive, but `HttpOnly` costs nothing and keeps script
from mutating them; `SameSite=Lax` matches the existing cookies.
## Part 2: Access-key cookies
### Goal
"Remember unlocked pastes on this browser" — after entering a password (or
after creating a private paste), subsequent visits skip the unlock form. This
extends `pw_<id>` from a 1-hour session convenience to a durable capability.
### Design: extend `pw_<id>`, don't invent a new scheme
The token is already HMAC(paste id, PALETTE_UNLOCK_SECRET) — unforgeable and
per-paste (fix for the #34 bypass). Changes:
1. **Opt-in checkbox on the unlock form** ("remember on this browser") and a
matching checkbox/note at creation time. Default OFF. Non-consenting
visitors keep the current 1-hour cookie.
2. **Extended lifetime** when opted in: `Max-Age = min(paste expiry, 90 days)`.
The cookie must never outlive the paste — derive the cap from the paste's
`ExpiresAt` at unlock time. Burn-after-N pastes: cap short (e.g. 24 h),
since the paste may burn at any read.
3. **Name collision**: paste ids are fixed-length server-generated, so
`pw_<id>` names stay bounded (~40 bytes each). With the 50-cookies-per-
domain budget, cap remembered pastes at ~30: when minting the 31st, drop
the oldest expired-paste cookies server-side (server knows which ids are
expired/deleted; send expired `Set-Cookie` with `Max-Age=0` to reclaim).
4. **Delete authorization interplay**: today a client-sent `vwr` matching the
paste's ViewerID authorizes delete. Access-key cookies grant *read*
capability only. Do not let a `pw_<id>` cookie authorize deletion — that
would mean cookie theft escalates from "read a paste" to "destroy it".
Delete stays bound to `vwr` or the deletion token.
### Scoping
- Keep `Path=/` (paste URLs are `/{id}` at the root; per-paste `Path=/{id}`
would work but saves nothing and complicates cleanup).
- Per-paste scope via the cookie *name* is the existing, tested pattern —
no shared "access key ring" cookie. A consolidated `keys` cookie would
mean one stolen cookie exposes every remembered paste at once.
## Part 3: Security considerations (honest accounting)
- **XSS exfiltration**: `HttpOnly` prevents JS from *reading* the cookies, but
not from *using* them — an XSS payload can simply `fetch('/<paste-id>')` and
exfiltrate the content through the page the cookie unlocks. HttpOnly raises
the bar (drive-by script can't dump the jar to an attacker server in one
request), it does not make access-key cookies safe. This is a real
limitation, not a solved problem. Mitigations in order of value:
1. Fix the stored-XSS class at the source — #34 already allowlisted
content-types on `/raw`; the standing debt items (CSP, X-Frame-Options,
Referrer-Policy) directly reduce cookie-use exfiltration and should land
before or with this feature.
2. Keep access-key cookies opt-in, so the blast radius is bounded to users
who accepted the tradeoff.
- **Cookie theft = paste access**: anyone holding `pw_<id>` can read that
paste until the cookie or paste expires, from any machine. That is inherent
to capability cookies. Consequences accepted deliberately: pastes here are
ephemeral (1 min1 yr expiry), passwords are low-stakes share convenience,
and there are no user accounts to compromise. Document this in the UI copy
("stores unlock access on this browser").
- **Shared machines**: a remembered cookie defeats the password for the next
user of the browser. The opt-in checkbox with plain-language copy is the
mitigation; do not default it on.
- **Cookie tossing / fixation**: a subdomain attacker could try to force
cookies; palette is a single host, no untrusted subdomains. `SameSite=Lax`
blocks cross-site attachment of the cookies on form posts to unlock
endpoints.
- **Multi-instance / secret rotation**: tokens are HMACs under
`PALETTE_UNLOCK_SECRET`; rotating the secret silently invalidates all
remembered cookies (acceptable — next visit re-prompts). Both dev and prod
k3s instances need the same secret only if sharing a domain, which they do
not.
- **Preferences cookie**: not security-sensitive, but still validate/allowlist
server-side to avoid it becoming an injection sink into templates.
## Recommendation
Implement in two small, separately reviewable pieces:
1. **`prefs` cookie** (do first, low risk): single JSON cookie < 256 bytes,
server-validated allowlist, `HttpOnly; SameSite=Lax; Max-Age=1y`, drives
only form pre-fill. Ship with #36's settings gear.
2. **Extended `pw_<id>` opt-in** (second, security-sensitive): opt-in checkbox,
Max-Age capped by paste expiry (90-day ceiling, 24 h for burn pastes),
oldest-cookie eviction at ~30 pastes, no delete authorization from access
cookies, and land the CSP/X-Frame-Options hardening debt from #34 in the
same or preceding change. UI copy must disclose that the cookie preserves
paste access on the browser.
Rejected alternatives: single consolidated access-key cookie (aggregate theft
risk, and the per-domain cookie-count argument cuts the other way for keys —
consolidation maximizes what one stolen cookie unlocks); localStorage for
preferences (XSS-readable, no benefit over HttpOnly cookies here); server-side
accounts/session table (out of scope — Palette is deliberately anonymous).
+1 -1
View File
@@ -63,7 +63,7 @@ const PaletteTable = (() => {
const filtered = state.filter.length > 0;
const off = (state.page - 1) * opts.perPage;
const url = (filtered || state.sortKey)
? opts.endpoint + '?limit=500&offset=0'
? opts.endpoint + '?limit=' + (opts.fetchLimit || 100) + '&offset=0'
: opts.endpoint + '?limit=' + opts.perPage + '&offset=' + off;
const res = await fetch(url);
const data = await res.json();