# FileServer A minimal Julia service that receives files over HTTP and hands them off to a pool of worker threads for processing. The HTTP endpoint does no real work: it spools each uploaded file to disk, pushes a lightweight reference onto a work queue, and responds immediately — staying free to accept the next upload. The per-file "processing" runs each file through a small neural-network classifier that labels it **known** (a file type resembling the training set) or **unknown**, and logs the result. See "File classifier" below. ## Architecture ``` POST /upload (multipart) │ ▼ ┌─────────────────┐ spool bytes to disk │ HTTP handler │────────────────────────► data/spool/- │ (Oxygen.jl) │ └────────┬─────────┘ enqueue reference (non-blocking) │ │ ▼ ▼ 202 + job IDs ┌───────────────┐ (503 if full) │ work queue │ bounded, thread-safe │ (Channel-ish)│ └───────┬───────┘ │ dequeue ┌───────────────┼───────────────┐ ▼ ▼ ▼ worker 1 worker 2 … worker N (Threads.@spawn) │ success ────┴──► data/done/- failure ───────► data/failed/- ``` Key properties: - **Fast intake:** the queue only ever carries small references; file bytes live on disk, so memory stays flat regardless of file size. - **Backpressure:** the queue is bounded (default 1000). When full, uploads get `503 Service Unavailable` instead of silently piling up. - **Crash-resilient:** files survive on disk. On startup, anything left in `data/spool/` is re-enqueued (`recovered = N` in the log). - **Graceful shutdown:** SIGINT (Ctrl-C) and SIGTERM (systemd/Docker/k8s `stop`) both stop accepting uploads, drain the queue, wait for in-flight files to finish, then exit. (See "Shutdown" below for one cosmetic caveat on SIGTERM.) - **Safe filenames:** client-supplied names are sanitized and prefixed with a server-minted UUID before touching the filesystem (no path traversal). ## The queue seam (→ RabbitMQ later) The HTTP handler and workers only ever call `enqueue!`, `dequeue!`, and `close!` on a `JobQueue` (see `src/queue.jl`). Today that's an in-process `ChannelQueue`. To move to RabbitMQ (or any broker), implement a new `JobQueue` subtype with those three methods and swap the construction in `run` — no handler or worker code changes. ## Running ```bash # install deps (first time) julia --project=. -e 'using Pkg; Pkg.instantiate()' # start the server; -t sets the number of OS threads available to workers julia --project=. -t auto bin/server.jl ``` ## Shutdown Both SIGINT and SIGTERM trigger the same idempotent graceful drain (stop serving → close queue → wait for workers → exit): - **SIGINT** is caught as an `InterruptException` (we call `Base.exit_on_sigint(false)`), so shutdown is clean and quiet. - **SIGTERM** can't be intercepted directly — Julia blocks it on worker threads and handles it in its own runtime, so a user `signal()` handler never fires. Instead we hook the drain into an `atexit` handler, which Julia's SIGTERM path does run. Caveat: Julia prints its own `signal 15: Terminated` backtrace *before* `atexit` runs. It's harmless noise — the drain still completes right after it — but if you want a fully quiet stop under a process manager, configure it to send SIGINT instead (systemd: `KillSignal=SIGINT`; Docker: `STOPSIGNAL SIGINT`). Give the stop timeout enough headroom to drain in-flight work (systemd: `TimeoutStopSec`). ## File classifier Each file is scored by a fixed-structure neural network (Lux.jl) that answers a single binary question: is this file **known** (like the types in the training set) or **unknown**? It's novelty detection, not exact file-typing — it won't tell you "PDF", just "this looks like something I was trained on, or not". - **Features:** the first 16 bytes + last 16 bytes of the file, each scaled 0–255 → `[0,1]`, giving a 32-dim input. Files under 32 bytes can't form that window and are classified `unknown` without touching the model. - **Architecture:** `Dense(32→64,relu) → Dense(64→16,relu) → Dense(16→2)`, raw logits; decision is `argmax` (class 1 = known, class 2 = unknown). - **Artifact:** trained weights live in `model/classifier.jld2` (committed), so the server just loads them at startup. Missing/unreadable ⇒ the server fails fast rather than run without classification. - **Effect today:** *annotate-only*. The class is logged (`classification=known|unknown`) but every file still moves to `done/`; the classifier can't misroute real files while it's unproven. The architecture and byte→feature mapping are defined once in `src/model.jl` and shared by the trainer and the server, so they can't drift apart. ### Training Training is a separate, offline script — it never runs in the request path: ```bash julia --project=. bin/train.jl [negatives_dir] ``` - **positives_dir** — every file in it (≥32 bytes) is a "known" example. - **negatives_dir** *(optional)* — a grab-bag of *other* real file types used as "unknown" examples. Negatives are generated ~1:1 with positives, split 50/50 between uniform-random byte vectors and grab-bag files. With no grab-bag dir, negatives are all random (weaker: the net may just learn "high entropy = unknown" rather than your actual types, so a grab-bag of real off-distribution files is recommended). The script uses an 80/20 seeded split, reports validation accuracy, and writes `model/classifier.jld2` (path overridable via `FS_MODEL_PATH`). A fixed seed (`FS_TRAIN_SEED`, default 42) drives negative generation, the split, and weight init, so the artifact is exactly regenerable from the same inputs. ## Configuration (environment variables) | Variable | Default | Meaning | |---------------------|----------------|------------------------------------------| | `FS_HOST` | `127.0.0.1` | Bind address | | `FS_PORT` | `8080` | Port | | `FS_WORKERS` | `nthreads()` | Number of worker tasks | | `FS_QUEUE_CAPACITY` | `1000` | Max pending jobs before `503` | | `FS_SPOOL_DIR` | `data/spool` | Incoming files (pending) | | `FS_DONE_DIR` | `data/done` | Files after successful processing | | `FS_FAILED_DIR` | `data/failed` | Files whose processing threw | | `FS_MODEL_PATH` | `model/classifier.jld2` | Classifier artifact loaded at startup | > To get real parallelism, start Julia with enough threads (`-t N`) to match > `FS_WORKERS`. If `FS_WORKERS` exceeds available threads you'll get a warning > and workers will share threads. ## Usage ```bash # health check curl http://127.0.0.1:8080/health # {"status":"ok"} # upload one or more files (multipart/form-data) curl -F "a=@report.pdf" -F "b=@data.csv" http://127.0.0.1:8080/upload # 202 {"accepted":[{"id":"","name":"report.pdf"}, ...]} ``` Each file in a request becomes its own job. Responses: - `202 Accepted` — all files spooled and queued (with per-file job IDs) - `400 Bad Request` — not multipart, or no files present - `503 Service Unavailable` — queue full, retry later - `500 Internal Server Error` — failed to write a file to disk ## Layout ``` src/ FileServer.jl module + run() (startup, recovery, workers, serve, shutdown) config.jl Config struct + env parsing job.jl Job (the queue reference) queue.jl JobQueue seam + in-process ChannelQueue spool.jl filename sanitizing, spool/move, startup recovery model.jl NN architecture + byte→feature mapping (shared with trainer) classify.jl load artifact + classify a file at inference time worker.jl worker loop + per-job processing (classify + move) server.jl HTTP routes/handlers bin/ server.jl entry point train.jl offline training script → model/classifier.jld2 model/ classifier.jld2 committed trained weights (loaded at startup) ```