Unknown files are no longer terminal. Stage 1 now routes :unknown onto a dedicated queue (with the same blocking backpressure as the known queue), and a third worker pool sorts each file into data/binary/ or data/text/ using a NUL-byte sniff of the first 8000 bytes. - content.jl: is_binary content sniff (stage 3) - worker.jl: handle_unknown_job; stage-1 routes unknown with backpressure; KNOWN_ENQUEUE_RETRY_SECONDS -> ROUTE_ENQUEUE_RETRY_SECONDS (serves both) - config.jl: unknown_worker_count/queue_capacity, binary_dir, text_dir + env - FileServer.jl: unknown queue, pool, stage-aware recovery, drain ordering - tests for is_binary and handle_unknown_job; tmp_config isolates new dirs - README: three-stage pipeline
279 lines
15 KiB
Markdown
279 lines
15 KiB
Markdown
# FileServer
|
||
|
||
A minimal Julia service that receives files over HTTP and hands them off to a
|
||
pool of worker threads for processing. The HTTP endpoint does no real work: it
|
||
spools each uploaded file to disk, pushes a lightweight reference onto a work
|
||
queue, and responds immediately — staying free to accept the next upload.
|
||
|
||
The per-file "processing" runs each file through a small neural-network
|
||
classifier that labels it **known** (a file type resembling the training set) or
|
||
**unknown**, and logs the result. See "File classifier" below.
|
||
|
||
## Architecture
|
||
|
||
The pipeline is three stages, each with its own bounded queue and its own worker
|
||
pool (tuned independently, since classification is CPU-bound, enrichment is
|
||
process-/IO-bound, and content triage is cheap IO):
|
||
|
||
```
|
||
POST /upload (multipart)
|
||
│
|
||
▼
|
||
┌─────────────────┐ spool bytes to disk
|
||
│ HTTP handler │────────────────────────► data/spool/<uuid>-<name>
|
||
│ (Oxygen.jl) │
|
||
└────────┬─────────┘ enqueue reference (non-blocking)
|
||
│ │
|
||
▼ ▼
|
||
202 + job IDs ┌────────────────────┐
|
||
(503 if full) │ stage-1 queue │ classification
|
||
└─────────┬──────────┘
|
||
│ dequeue
|
||
┌───────────────────┼───────────────────┐
|
||
▼ ▼ ▼
|
||
classify wkr 1 classify wkr 2 … classify wkr N
|
||
│
|
||
┌────────────┴────────────┐
|
||
:unknown :known
|
||
│ move to data/unknown/, │ move to data/known/, then
|
||
▼ then enqueue (blocking) ▼ enqueue (blocking backpressure)
|
||
┌────────────────────┐ ┌────────────────────┐
|
||
│ unknown queue │ │ known queue │ enrichment
|
||
└─────────┬──────────┘ └─────────┬──────────┘
|
||
│ dequeue │ dequeue
|
||
┌────────┼────────┐ ┌───────────┼───────────┐
|
||
▼ ▼ ▼ ▼ ▼ ▼
|
||
unk 1 unk 2 … unk K known wkr 1 known wkr 2 … known wkr M
|
||
│ binary-vs-text sniff │ exiftool → normalized sidecar
|
||
├─► data/binary/<uuid>-<name> success ──┴──► data/done/<uuid>-<name>
|
||
└─► data/text/<uuid>-<name> data/done/<uuid>-<name>.meta.json
|
||
(sidecar-first commit)
|
||
failure ───────► data/failed/<uuid>-<name>
|
||
```
|
||
|
||
Stages 2 (enrichment) and 3 (content triage) run in parallel: stage 1 feeds both
|
||
the known and unknown queues.
|
||
|
||
Key properties:
|
||
|
||
- **Fast intake:** the queue only ever carries small references; file bytes live
|
||
on disk, so memory stays flat regardless of file size.
|
||
- **Backpressure:** each queue is bounded (default 1000). When the *intake* queue
|
||
is full, uploads get `503 Service Unavailable`. When the *known* queue is full,
|
||
the stage-1 worker blocks and retries (a classified file is never dropped).
|
||
- **Crash-resilient:** files survive on disk. On startup, recovery is
|
||
stage-aware: leftovers in `data/spool/` re-enter classification, `data/known/`
|
||
re-enter enrichment, and `data/unknown/` re-enter content triage (`recovered` /
|
||
`recovered_known` / `recovered_unknown` in the log), so a file resumes at its
|
||
correct stage instead of restarting from scratch.
|
||
- **Graceful shutdown:** SIGINT (Ctrl-C) and SIGTERM (systemd/Docker/k8s `stop`)
|
||
both stop accepting uploads, then drain the stages *in order* — close the
|
||
stage-1 queue and wait out the classify workers (the only producer of the known
|
||
*and* unknown queues) before closing those two queues and waiting out the
|
||
enrich and content-triage workers.
|
||
(See "Shutdown" below for one cosmetic caveat on SIGTERM.)
|
||
- **Safe filenames:** client-supplied names are sanitized and prefixed with a
|
||
server-minted UUID before touching the filesystem (no path traversal).
|
||
|
||
### Metadata enrichment (stage 2)
|
||
|
||
Files the classifier labels **known** are handed to a second pool that extracts
|
||
metadata with [`exiftool`](https://exiftool.org/) (`exiftool -json -G -n`) —
|
||
chosen because no native Julia library comes close to its multi-format coverage.
|
||
The output is normalized into a small, stable, documented schema and written as a
|
||
JSON **sidecar** next to the file in `data/done/`, e.g.
|
||
`data/done/<uuid>-<name>.meta.json`. The original bytes are never modified.
|
||
|
||
> **Prerequisite:** `exiftool` must be on `PATH` (e.g. `apt install
|
||
> libimage-exiftool-perl`). The server **fails fast at startup** if it's missing.
|
||
|
||
Sidecar top-level fields (all nullable — present only when available), plus the
|
||
complete raw `exiftool` object under `raw`:
|
||
|
||
| Field | Meaning |
|
||
|---|---|
|
||
| `id`, `original_name` | job id and client-supplied name |
|
||
| `file_type`, `mime_type` | e.g. `PDF` / `application/pdf` |
|
||
| `file_size` | bytes (authoritative, from intake — not exiftool) |
|
||
| `created_date`, `modified_date` | content timestamps |
|
||
| `author` | person (`Author`/`Artist`/`By-line`) |
|
||
| `created_by` | authoring app/tool (`Producer`/`CreatorTool`/`Creator`/`Software`/…) |
|
||
| `dimensions` | `{width, height}` for media |
|
||
| `duration` | seconds, for audio/video |
|
||
| `page_count` | for documents |
|
||
| `error` | set on a *degraded* sidecar (see below) |
|
||
| `raw` | full `exiftool` output |
|
||
|
||
Each normalized field is a coalesce over a priority list of exiftool tags
|
||
(`src/metadata.jl`); extend a field by appending tag names. If extraction fails
|
||
or `exiftool` times out (`FS_EXIFTOOL_TIMEOUT`, default 30s), the file still
|
||
completes to `data/done/` with a **degraded sidecar** — `file_size`/`file_type`
|
||
plus an `error` note — rather than being quarantined, because it's still a wanted
|
||
known file. Only genuine I/O errors (can't write the sidecar or move the file)
|
||
send it to `data/failed/`.
|
||
|
||
The sidecar is committed **before** the file is moved into `data/done/`, so a
|
||
file's presence there always implies its sidecar is already present; a crash in
|
||
between leaves only a harmless orphan sidecar, and recovery re-enriches
|
||
idempotently.
|
||
|
||
### Content triage (stage 3)
|
||
|
||
Files the classifier labels **unknown** are handed to a third pool that sorts
|
||
them into two coarse buckets so downstream tooling can treat them differently:
|
||
|
||
- **`data/binary/`** — the file looks like binary data.
|
||
- **`data/text/`** — the file looks like text.
|
||
|
||
The test is the classic **NUL-byte sniff** (the same heuristic `git` and
|
||
`file(1)` use): read the first 8000 bytes and, if any is NUL, call it binary,
|
||
else text. It's cheap (no full read) and reliable in practice — text encodings
|
||
don't embed NUL bytes, while binary formats almost always do near the start. An
|
||
empty file has no NUL, so it's treated as text. This is deliberately simple for
|
||
now; richer handling can hang off either bucket later (`src/content.jl`).
|
||
|
||
## The queue seam (→ RabbitMQ later)
|
||
|
||
The HTTP handler and workers only ever call `enqueue!`, `dequeue!`, and
|
||
`close!` on a `JobQueue` (see `src/queue.jl`). Today that's an in-process
|
||
`ChannelQueue`. To move to RabbitMQ (or any broker), implement a new `JobQueue`
|
||
subtype with those three methods and swap the construction in `run` — no handler
|
||
or worker code changes.
|
||
|
||
## Running
|
||
|
||
```bash
|
||
# install deps (first time)
|
||
julia --project=. -e 'using Pkg; Pkg.instantiate()'
|
||
|
||
# start the server; -t sets the number of OS threads available to workers
|
||
julia --project=. -t auto bin/server.jl
|
||
```
|
||
|
||
## Shutdown
|
||
|
||
Both SIGINT and SIGTERM trigger the same idempotent graceful drain
|
||
(stop serving → close queue → wait for workers → exit):
|
||
|
||
- **SIGINT** is caught as an `InterruptException` (we call
|
||
`Base.exit_on_sigint(false)`), so shutdown is clean and quiet.
|
||
- **SIGTERM** can't be intercepted directly — Julia blocks it on worker threads
|
||
and handles it in its own runtime, so a user `signal()` handler never fires.
|
||
Instead we hook the drain into an `atexit` handler, which Julia's SIGTERM path
|
||
does run. Caveat: Julia prints its own `signal 15: Terminated` backtrace
|
||
*before* `atexit` runs. It's harmless noise — the drain still completes right
|
||
after it — but if you want a fully quiet stop under a process manager,
|
||
configure it to send SIGINT instead (systemd: `KillSignal=SIGINT`; Docker:
|
||
`STOPSIGNAL SIGINT`). Give the stop timeout enough headroom to drain
|
||
in-flight work (systemd: `TimeoutStopSec`).
|
||
|
||
## File classifier
|
||
|
||
Each file is scored by a fixed-structure neural network (Lux.jl) that answers a
|
||
single binary question: is this file **known** (like the types in the training
|
||
set) or **unknown**? It's novelty detection, not exact file-typing — it won't
|
||
tell you "PDF", just "this looks like something I was trained on, or not".
|
||
|
||
- **Features:** the first 16 bytes + last 16 bytes of the file, each scaled
|
||
0–255 → `[0,1]`, giving a 32-dim input. Files under 32 bytes can't form that
|
||
window and are classified `unknown` without touching the model.
|
||
- **Architecture:** `Dense(32→64,relu) → Dense(64→16,relu) → Dense(16→2)`,
|
||
raw logits; decision is `argmax` (class 1 = known, class 2 = unknown).
|
||
- **Artifact:** trained weights live in `model/classifier.jld2` (committed), so
|
||
the server just loads them at startup. Missing/unreadable ⇒ the server fails
|
||
fast rather than run without classification.
|
||
- **Effect today:** *annotate-only*. The class is logged
|
||
(`classification=known|unknown`) but every file still moves to `done/`; the
|
||
classifier can't misroute real files while it's unproven.
|
||
|
||
The architecture and byte→feature mapping are defined once in `src/model.jl` and
|
||
shared by the trainer and the server, so they can't drift apart.
|
||
|
||
### Training
|
||
|
||
Training is a separate, offline script — it never runs in the request path:
|
||
|
||
```bash
|
||
julia --project=. bin/train.jl <positives_dir> [negatives_dir]
|
||
```
|
||
|
||
- **positives_dir** — every file in it (≥32 bytes) is a "known" example.
|
||
- **negatives_dir** *(optional)* — a grab-bag of *other* real file types used as
|
||
"unknown" examples. Negatives are generated ~1:1 with positives, split 50/50
|
||
between uniform-random byte vectors and grab-bag files. With no grab-bag dir,
|
||
negatives are all random (weaker: the net may just learn "high entropy =
|
||
unknown" rather than your actual types, so a grab-bag of real off-distribution
|
||
files is recommended).
|
||
|
||
The script uses an 80/20 seeded split, reports validation accuracy, and writes
|
||
`model/classifier.jld2` (path overridable via `FS_MODEL_PATH`). A fixed seed
|
||
(`FS_TRAIN_SEED`, default 42) drives negative generation, the split, and weight
|
||
init, so the artifact is exactly regenerable from the same inputs.
|
||
|
||
## Configuration (environment variables)
|
||
|
||
| Variable | Default | Meaning |
|
||
|---------------------|----------------|------------------------------------------|
|
||
| `FS_HOST` | `127.0.0.1` | Bind address |
|
||
| `FS_PORT` | `8080` | Port |
|
||
| `FS_WORKERS` | `nthreads()` | Stage-1 (classification) worker tasks |
|
||
| `FS_QUEUE_CAPACITY` | `1000` | Max pending intake jobs before `503` |
|
||
| `FS_KNOWN_WORKERS` | `nthreads()` | Stage-2 (enrichment) worker tasks |
|
||
| `FS_KNOWN_QUEUE_CAPACITY` | `1000` | Max pending enrichment jobs (then backpressure) |
|
||
| `FS_UNKNOWN_WORKERS` | `nthreads()` | Stage-3 (content triage) worker tasks |
|
||
| `FS_UNKNOWN_QUEUE_CAPACITY` | `1000` | Max pending triage jobs (then backpressure) |
|
||
| `FS_SPOOL_DIR` | `data/spool` | Incoming files (pending classification) |
|
||
| `FS_KNOWN_DIR` | `data/known` | Classified-known, awaiting enrichment |
|
||
| `FS_UNKNOWN_DIR` | `data/unknown` | Classified-unknown, awaiting content triage |
|
||
| `FS_BINARY_DIR` | `data/binary` | Stage-3 sink: unknown files that look binary |
|
||
| `FS_TEXT_DIR` | `data/text` | Stage-3 sink: unknown files that look like text |
|
||
| `FS_DONE_DIR` | `data/done` | Enriched known files (+ `.meta.json`) |
|
||
| `FS_FAILED_DIR` | `data/failed` | Files whose processing threw |
|
||
| `FS_MODEL_PATH` | `model/classifier.jld2` | Classifier artifact loaded at startup |
|
||
| `FS_EXIFTOOL_TIMEOUT` | `30` | Seconds before a stuck exiftool is killed |
|
||
|
||
> To get real parallelism, start Julia with enough threads (`-t N`) to cover all
|
||
> pools. If `FS_WORKERS + FS_KNOWN_WORKERS + FS_UNKNOWN_WORKERS` exceeds available
|
||
> threads you'll get a warning (non-fatal) and workers will share threads.
|
||
|
||
## Usage
|
||
|
||
```bash
|
||
# health check
|
||
curl http://127.0.0.1:8080/health
|
||
# {"status":"ok"}
|
||
|
||
# upload one or more files (multipart/form-data)
|
||
curl -F "a=@report.pdf" -F "b=@data.csv" http://127.0.0.1:8080/upload
|
||
# 202 {"accepted":[{"id":"<uuid>","name":"report.pdf"}, ...]}
|
||
```
|
||
|
||
Each file in a request becomes its own job. Responses:
|
||
|
||
- `202 Accepted` — all files spooled and queued (with per-file job IDs)
|
||
- `400 Bad Request` — not multipart, or no files present
|
||
- `503 Service Unavailable` — queue full, retry later
|
||
- `500 Internal Server Error` — failed to write a file to disk
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/
|
||
FileServer.jl module + run() (startup, recovery, workers, serve, shutdown)
|
||
config.jl Config struct + env parsing
|
||
job.jl Job (the queue reference)
|
||
queue.jl JobQueue seam + in-process ChannelQueue
|
||
spool.jl filename sanitizing, spool/move, startup recovery
|
||
model.jl NN architecture + byte→feature mapping (shared with trainer)
|
||
classify.jl load artifact + classify a file at inference time
|
||
metadata.jl exiftool extraction + normalized sidecar (stage 2)
|
||
content.jl binary-vs-text sniff for unknown files (stage 3)
|
||
worker.jl parametrized worker loop + classify/enrich/triage handlers
|
||
server.jl HTTP routes/handlers
|
||
bin/
|
||
server.jl entry point
|
||
train.jl offline training script → model/classifier.jld2
|
||
model/
|
||
classifier.jld2 committed trained weights (loaded at startup)
|
||
```
|