Files
file-server/README.md
Jeffrey Ward e55129e3a4 Add Lux.jl file classifier (known/unknown) with offline trainer
Each uploaded file is scored by a fixed-structure neural net that labels it
known (resembling the training set) or unknown — novelty detection over the
first 16 + last 16 bytes (scaled to [0,1]), Dense(32->64->16->2), argmax.

- src/model.jl: shared architecture + byte->feature mapping (trainer + server)
- src/classify.jl: load committed artifact, classify a file at inference
- bin/train.jl: offline trainer, 1:1 blended negatives (random + grab-bag),
  seeded 80/20 split, writes model/classifier.jld2
- worker: classify (annotate-only) and log classification=known|unknown
- config: FS_MODEL_PATH; server fails fast if the artifact is missing
- deps: Lux, JLD2, Optimisers, Zygote
2026-07-02 14:13:57 -04:00

184 lines
8.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# FileServer
A minimal Julia service that receives files over HTTP and hands them off to a
pool of worker threads for processing. The HTTP endpoint does no real work: it
spools each uploaded file to disk, pushes a lightweight reference onto a work
queue, and responds immediately — staying free to accept the next upload.
The per-file "processing" runs each file through a small neural-network
classifier that labels it **known** (a file type resembling the training set) or
**unknown**, and logs the result. See "File classifier" below.
## Architecture
```
POST /upload (multipart)
┌─────────────────┐ spool bytes to disk
│ HTTP handler │────────────────────────► data/spool/<uuid>-<name>
│ (Oxygen.jl) │
└────────┬─────────┘ enqueue reference (non-blocking)
│ │
▼ ▼
202 + job IDs ┌───────────────┐
(503 if full) │ work queue │ bounded, thread-safe
│ (Channel-ish)│
└───────┬───────┘
│ dequeue
┌───────────────┼───────────────┐
▼ ▼ ▼
worker 1 worker 2 … worker N (Threads.@spawn)
success ────┴──► data/done/<uuid>-<name>
failure ───────► data/failed/<uuid>-<name>
```
Key properties:
- **Fast intake:** the queue only ever carries small references; file bytes live
on disk, so memory stays flat regardless of file size.
- **Backpressure:** the queue is bounded (default 1000). When full, uploads get
`503 Service Unavailable` instead of silently piling up.
- **Crash-resilient:** files survive on disk. On startup, anything left in
`data/spool/` is re-enqueued (`recovered = N` in the log).
- **Graceful shutdown:** SIGINT (Ctrl-C) and SIGTERM (systemd/Docker/k8s `stop`)
both stop accepting uploads, drain the queue, wait for in-flight files to
finish, then exit. (See "Shutdown" below for one cosmetic caveat on SIGTERM.)
- **Safe filenames:** client-supplied names are sanitized and prefixed with a
server-minted UUID before touching the filesystem (no path traversal).
## The queue seam (→ RabbitMQ later)
The HTTP handler and workers only ever call `enqueue!`, `dequeue!`, and
`close!` on a `JobQueue` (see `src/queue.jl`). Today that's an in-process
`ChannelQueue`. To move to RabbitMQ (or any broker), implement a new `JobQueue`
subtype with those three methods and swap the construction in `run` — no handler
or worker code changes.
## Running
```bash
# install deps (first time)
julia --project=. -e 'using Pkg; Pkg.instantiate()'
# start the server; -t sets the number of OS threads available to workers
julia --project=. -t auto bin/server.jl
```
## Shutdown
Both SIGINT and SIGTERM trigger the same idempotent graceful drain
(stop serving → close queue → wait for workers → exit):
- **SIGINT** is caught as an `InterruptException` (we call
`Base.exit_on_sigint(false)`), so shutdown is clean and quiet.
- **SIGTERM** can't be intercepted directly — Julia blocks it on worker threads
and handles it in its own runtime, so a user `signal()` handler never fires.
Instead we hook the drain into an `atexit` handler, which Julia's SIGTERM path
does run. Caveat: Julia prints its own `signal 15: Terminated` backtrace
*before* `atexit` runs. It's harmless noise — the drain still completes right
after it — but if you want a fully quiet stop under a process manager,
configure it to send SIGINT instead (systemd: `KillSignal=SIGINT`; Docker:
`STOPSIGNAL SIGINT`). Give the stop timeout enough headroom to drain
in-flight work (systemd: `TimeoutStopSec`).
## File classifier
Each file is scored by a fixed-structure neural network (Lux.jl) that answers a
single binary question: is this file **known** (like the types in the training
set) or **unknown**? It's novelty detection, not exact file-typing — it won't
tell you "PDF", just "this looks like something I was trained on, or not".
- **Features:** the first 16 bytes + last 16 bytes of the file, each scaled
0255 → `[0,1]`, giving a 32-dim input. Files under 32 bytes can't form that
window and are classified `unknown` without touching the model.
- **Architecture:** `Dense(32→64,relu) → Dense(64→16,relu) → Dense(16→2)`,
raw logits; decision is `argmax` (class 1 = known, class 2 = unknown).
- **Artifact:** trained weights live in `model/classifier.jld2` (committed), so
the server just loads them at startup. Missing/unreadable ⇒ the server fails
fast rather than run without classification.
- **Effect today:** *annotate-only*. The class is logged
(`classification=known|unknown`) but every file still moves to `done/`; the
classifier can't misroute real files while it's unproven.
The architecture and byte→feature mapping are defined once in `src/model.jl` and
shared by the trainer and the server, so they can't drift apart.
### Training
Training is a separate, offline script — it never runs in the request path:
```bash
julia --project=. bin/train.jl <positives_dir> [negatives_dir]
```
- **positives_dir** — every file in it (≥32 bytes) is a "known" example.
- **negatives_dir** *(optional)* — a grab-bag of *other* real file types used as
"unknown" examples. Negatives are generated ~1:1 with positives, split 50/50
between uniform-random byte vectors and grab-bag files. With no grab-bag dir,
negatives are all random (weaker: the net may just learn "high entropy =
unknown" rather than your actual types, so a grab-bag of real off-distribution
files is recommended).
The script uses an 80/20 seeded split, reports validation accuracy, and writes
`model/classifier.jld2` (path overridable via `FS_MODEL_PATH`). A fixed seed
(`FS_TRAIN_SEED`, default 42) drives negative generation, the split, and weight
init, so the artifact is exactly regenerable from the same inputs.
## Configuration (environment variables)
| Variable | Default | Meaning |
|---------------------|----------------|------------------------------------------|
| `FS_HOST` | `127.0.0.1` | Bind address |
| `FS_PORT` | `8080` | Port |
| `FS_WORKERS` | `nthreads()` | Number of worker tasks |
| `FS_QUEUE_CAPACITY` | `1000` | Max pending jobs before `503` |
| `FS_SPOOL_DIR` | `data/spool` | Incoming files (pending) |
| `FS_DONE_DIR` | `data/done` | Files after successful processing |
| `FS_FAILED_DIR` | `data/failed` | Files whose processing threw |
| `FS_MODEL_PATH` | `model/classifier.jld2` | Classifier artifact loaded at startup |
> To get real parallelism, start Julia with enough threads (`-t N`) to match
> `FS_WORKERS`. If `FS_WORKERS` exceeds available threads you'll get a warning
> and workers will share threads.
## Usage
```bash
# health check
curl http://127.0.0.1:8080/health
# {"status":"ok"}
# upload one or more files (multipart/form-data)
curl -F "a=@report.pdf" -F "b=@data.csv" http://127.0.0.1:8080/upload
# 202 {"accepted":[{"id":"<uuid>","name":"report.pdf"}, ...]}
```
Each file in a request becomes its own job. Responses:
- `202 Accepted` — all files spooled and queued (with per-file job IDs)
- `400 Bad Request` — not multipart, or no files present
- `503 Service Unavailable` — queue full, retry later
- `500 Internal Server Error` — failed to write a file to disk
## Layout
```
src/
FileServer.jl module + run() (startup, recovery, workers, serve, shutdown)
config.jl Config struct + env parsing
job.jl Job (the queue reference)
queue.jl JobQueue seam + in-process ChannelQueue
spool.jl filename sanitizing, spool/move, startup recovery
model.jl NN architecture + byte→feature mapping (shared with trainer)
classify.jl load artifact + classify a file at inference time
worker.jl worker loop + per-job processing (classify + move)
server.jl HTTP routes/handlers
bin/
server.jl entry point
train.jl offline training script → model/classifier.jld2
model/
classifier.jld2 committed trained weights (loaded at startup)
```