Each uploaded file is scored by a fixed-structure neural net that labels it known (resembling the training set) or unknown — novelty detection over the first 16 + last 16 bytes (scaled to [0,1]), Dense(32->64->16->2), argmax. - src/model.jl: shared architecture + byte->feature mapping (trainer + server) - src/classify.jl: load committed artifact, classify a file at inference - bin/train.jl: offline trainer, 1:1 blended negatives (random + grab-bag), seeded 80/20 split, writes model/classifier.jld2 - worker: classify (annotate-only) and log classification=known|unknown - config: FS_MODEL_PATH; server fails fast if the artifact is missing - deps: Lux, JLD2, Optimisers, Zygote
184 lines
8.7 KiB
Markdown
184 lines
8.7 KiB
Markdown
# FileServer
|
||
|
||
A minimal Julia service that receives files over HTTP and hands them off to a
|
||
pool of worker threads for processing. The HTTP endpoint does no real work: it
|
||
spools each uploaded file to disk, pushes a lightweight reference onto a work
|
||
queue, and responds immediately — staying free to accept the next upload.
|
||
|
||
The per-file "processing" runs each file through a small neural-network
|
||
classifier that labels it **known** (a file type resembling the training set) or
|
||
**unknown**, and logs the result. See "File classifier" below.
|
||
|
||
## Architecture
|
||
|
||
```
|
||
POST /upload (multipart)
|
||
│
|
||
▼
|
||
┌─────────────────┐ spool bytes to disk
|
||
│ HTTP handler │────────────────────────► data/spool/<uuid>-<name>
|
||
│ (Oxygen.jl) │
|
||
└────────┬─────────┘ enqueue reference (non-blocking)
|
||
│ │
|
||
▼ ▼
|
||
202 + job IDs ┌───────────────┐
|
||
(503 if full) │ work queue │ bounded, thread-safe
|
||
│ (Channel-ish)│
|
||
└───────┬───────┘
|
||
│ dequeue
|
||
┌───────────────┼───────────────┐
|
||
▼ ▼ ▼
|
||
worker 1 worker 2 … worker N (Threads.@spawn)
|
||
│
|
||
success ────┴──► data/done/<uuid>-<name>
|
||
failure ───────► data/failed/<uuid>-<name>
|
||
```
|
||
|
||
Key properties:
|
||
|
||
- **Fast intake:** the queue only ever carries small references; file bytes live
|
||
on disk, so memory stays flat regardless of file size.
|
||
- **Backpressure:** the queue is bounded (default 1000). When full, uploads get
|
||
`503 Service Unavailable` instead of silently piling up.
|
||
- **Crash-resilient:** files survive on disk. On startup, anything left in
|
||
`data/spool/` is re-enqueued (`recovered = N` in the log).
|
||
- **Graceful shutdown:** SIGINT (Ctrl-C) and SIGTERM (systemd/Docker/k8s `stop`)
|
||
both stop accepting uploads, drain the queue, wait for in-flight files to
|
||
finish, then exit. (See "Shutdown" below for one cosmetic caveat on SIGTERM.)
|
||
- **Safe filenames:** client-supplied names are sanitized and prefixed with a
|
||
server-minted UUID before touching the filesystem (no path traversal).
|
||
|
||
## The queue seam (→ RabbitMQ later)
|
||
|
||
The HTTP handler and workers only ever call `enqueue!`, `dequeue!`, and
|
||
`close!` on a `JobQueue` (see `src/queue.jl`). Today that's an in-process
|
||
`ChannelQueue`. To move to RabbitMQ (or any broker), implement a new `JobQueue`
|
||
subtype with those three methods and swap the construction in `run` — no handler
|
||
or worker code changes.
|
||
|
||
## Running
|
||
|
||
```bash
|
||
# install deps (first time)
|
||
julia --project=. -e 'using Pkg; Pkg.instantiate()'
|
||
|
||
# start the server; -t sets the number of OS threads available to workers
|
||
julia --project=. -t auto bin/server.jl
|
||
```
|
||
|
||
## Shutdown
|
||
|
||
Both SIGINT and SIGTERM trigger the same idempotent graceful drain
|
||
(stop serving → close queue → wait for workers → exit):
|
||
|
||
- **SIGINT** is caught as an `InterruptException` (we call
|
||
`Base.exit_on_sigint(false)`), so shutdown is clean and quiet.
|
||
- **SIGTERM** can't be intercepted directly — Julia blocks it on worker threads
|
||
and handles it in its own runtime, so a user `signal()` handler never fires.
|
||
Instead we hook the drain into an `atexit` handler, which Julia's SIGTERM path
|
||
does run. Caveat: Julia prints its own `signal 15: Terminated` backtrace
|
||
*before* `atexit` runs. It's harmless noise — the drain still completes right
|
||
after it — but if you want a fully quiet stop under a process manager,
|
||
configure it to send SIGINT instead (systemd: `KillSignal=SIGINT`; Docker:
|
||
`STOPSIGNAL SIGINT`). Give the stop timeout enough headroom to drain
|
||
in-flight work (systemd: `TimeoutStopSec`).
|
||
|
||
## File classifier
|
||
|
||
Each file is scored by a fixed-structure neural network (Lux.jl) that answers a
|
||
single binary question: is this file **known** (like the types in the training
|
||
set) or **unknown**? It's novelty detection, not exact file-typing — it won't
|
||
tell you "PDF", just "this looks like something I was trained on, or not".
|
||
|
||
- **Features:** the first 16 bytes + last 16 bytes of the file, each scaled
|
||
0–255 → `[0,1]`, giving a 32-dim input. Files under 32 bytes can't form that
|
||
window and are classified `unknown` without touching the model.
|
||
- **Architecture:** `Dense(32→64,relu) → Dense(64→16,relu) → Dense(16→2)`,
|
||
raw logits; decision is `argmax` (class 1 = known, class 2 = unknown).
|
||
- **Artifact:** trained weights live in `model/classifier.jld2` (committed), so
|
||
the server just loads them at startup. Missing/unreadable ⇒ the server fails
|
||
fast rather than run without classification.
|
||
- **Effect today:** *annotate-only*. The class is logged
|
||
(`classification=known|unknown`) but every file still moves to `done/`; the
|
||
classifier can't misroute real files while it's unproven.
|
||
|
||
The architecture and byte→feature mapping are defined once in `src/model.jl` and
|
||
shared by the trainer and the server, so they can't drift apart.
|
||
|
||
### Training
|
||
|
||
Training is a separate, offline script — it never runs in the request path:
|
||
|
||
```bash
|
||
julia --project=. bin/train.jl <positives_dir> [negatives_dir]
|
||
```
|
||
|
||
- **positives_dir** — every file in it (≥32 bytes) is a "known" example.
|
||
- **negatives_dir** *(optional)* — a grab-bag of *other* real file types used as
|
||
"unknown" examples. Negatives are generated ~1:1 with positives, split 50/50
|
||
between uniform-random byte vectors and grab-bag files. With no grab-bag dir,
|
||
negatives are all random (weaker: the net may just learn "high entropy =
|
||
unknown" rather than your actual types, so a grab-bag of real off-distribution
|
||
files is recommended).
|
||
|
||
The script uses an 80/20 seeded split, reports validation accuracy, and writes
|
||
`model/classifier.jld2` (path overridable via `FS_MODEL_PATH`). A fixed seed
|
||
(`FS_TRAIN_SEED`, default 42) drives negative generation, the split, and weight
|
||
init, so the artifact is exactly regenerable from the same inputs.
|
||
|
||
## Configuration (environment variables)
|
||
|
||
| Variable | Default | Meaning |
|
||
|---------------------|----------------|------------------------------------------|
|
||
| `FS_HOST` | `127.0.0.1` | Bind address |
|
||
| `FS_PORT` | `8080` | Port |
|
||
| `FS_WORKERS` | `nthreads()` | Number of worker tasks |
|
||
| `FS_QUEUE_CAPACITY` | `1000` | Max pending jobs before `503` |
|
||
| `FS_SPOOL_DIR` | `data/spool` | Incoming files (pending) |
|
||
| `FS_DONE_DIR` | `data/done` | Files after successful processing |
|
||
| `FS_FAILED_DIR` | `data/failed` | Files whose processing threw |
|
||
| `FS_MODEL_PATH` | `model/classifier.jld2` | Classifier artifact loaded at startup |
|
||
|
||
> To get real parallelism, start Julia with enough threads (`-t N`) to match
|
||
> `FS_WORKERS`. If `FS_WORKERS` exceeds available threads you'll get a warning
|
||
> and workers will share threads.
|
||
|
||
## Usage
|
||
|
||
```bash
|
||
# health check
|
||
curl http://127.0.0.1:8080/health
|
||
# {"status":"ok"}
|
||
|
||
# upload one or more files (multipart/form-data)
|
||
curl -F "a=@report.pdf" -F "b=@data.csv" http://127.0.0.1:8080/upload
|
||
# 202 {"accepted":[{"id":"<uuid>","name":"report.pdf"}, ...]}
|
||
```
|
||
|
||
Each file in a request becomes its own job. Responses:
|
||
|
||
- `202 Accepted` — all files spooled and queued (with per-file job IDs)
|
||
- `400 Bad Request` — not multipart, or no files present
|
||
- `503 Service Unavailable` — queue full, retry later
|
||
- `500 Internal Server Error` — failed to write a file to disk
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/
|
||
FileServer.jl module + run() (startup, recovery, workers, serve, shutdown)
|
||
config.jl Config struct + env parsing
|
||
job.jl Job (the queue reference)
|
||
queue.jl JobQueue seam + in-process ChannelQueue
|
||
spool.jl filename sanitizing, spool/move, startup recovery
|
||
model.jl NN architecture + byte→feature mapping (shared with trainer)
|
||
classify.jl load artifact + classify a file at inference time
|
||
worker.jl worker loop + per-job processing (classify + move)
|
||
server.jl HTTP routes/handlers
|
||
bin/
|
||
server.jl entry point
|
||
train.jl offline training script → model/classifier.jld2
|
||
model/
|
||
classifier.jld2 committed trained weights (loaded at startup)
|
||
```
|