Files
file-server/README.md
Jeffrey Ward d9f32d9aaf Add stage-5 unknown-format discovery: header clustering + calibration
Implements phase A of the DESIGN_clustering.md design: a Dirichlet-process
mixture of per-position categoricals over the first 32 header bytes (257-symbol
alphabet) that clusters the binary/ pile by file format, plus signature
extraction and promotion nomination. All base-Julia (a Lanczos loggamma keeps
the Dirichlet-multinomial marginal dependency-free).

- src/cluster.jl: header_symbols feature extraction, collapsed Gibbs sampler
  (phase A), sequential CRP-predictive assignment (phase B core), signatures/
  promotion, and ARI/V-measure calibration metrics.
- bin/cluster_calibrate.jl: grid-tunes hyperparameters against magic-collapsed
  ground truth and cross-checks a model-free NCD (gzip) baseline.
- FS_CLUSTER_*/FS_PROMOTE_* config knobs; wire cluster.jl into the module.
- Tests for the three DESIGN §10 assertions plus the model primitives.

Calibrated defaults (n=32, alpha=1.0, beta=0.1) recover known formats at
ARI 0.77 (0.885 excl. tar); docx+zip and the ELF family merge correctly and the
NCD baseline agrees. DESIGN §11 records the results and three assumptions the
data corrected (tar/ELF header-zero merge, the cold-start seeding deadlock, and
the Bernoulli signature / Occam-penalized restart scoring).
2026-07-03 16:43:52 -04:00

23 KiB
Raw Blame History

FileServer

A minimal Julia service that receives files over HTTP and hands them off to a pool of worker threads for processing. The HTTP endpoint does no real work: it spools each uploaded file to disk, pushes a lightweight reference onto a work queue, and responds immediately — staying free to accept the next upload.

The per-file "processing" runs each file through a small neural-network classifier that labels it known (a file type resembling the training set) or unknown, and logs the result. See "File classifier" below.

Architecture

The pipeline is four stages, each with its own bounded queue and its own worker pool (tuned independently, since classification is CPU-bound, known-file enrichment is process-/IO-bound, content triage is cheap IO, and language enrichment mixes CPU with a subprocess):

        POST /upload (multipart)
                 │
                 ▼
        ┌─────────────────┐     spool bytes to disk
        │  HTTP handler    │────────────────────────► data/spool/<uuid>-<name>
        │  (Oxygen.jl)     │
        └────────┬─────────┘     enqueue reference (non-blocking)
                 │                        │
                 ▼                        ▼
          202 + job IDs        ┌────────────────────┐
          (503 if full)        │  stage-1 queue     │  classification
                               └─────────┬──────────┘
                                         │ dequeue
                     ┌───────────────────┼───────────────────┐
                     ▼                   ▼                    ▼
              classify wkr 1      classify wkr 2   …   classify wkr N
                     │
        ┌────────────┴────────────┐
    :unknown                    :known
        │  move to data/unknown/,  │  move to data/known/, then
        ▼  then enqueue (blocking) ▼  enqueue (blocking backpressure)
  ┌────────────────────┐        ┌────────────────────┐
  │  unknown queue     │        │  known queue       │  enrichment
  └─────────┬──────────┘        └─────────┬──────────┘
            │ dequeue                     │ dequeue
   ┌────────┼────────┐        ┌───────────┼───────────┐
   ▼        ▼        ▼        ▼           ▼           ▼
 unk 1  unk 2 … unk K   known wkr 1  known wkr 2 … known wkr M
   │  binary-vs-text sniff       │  exiftool → normalized sidecar
   ├─► data/binary/<uuid>-<name> success ──┴──► data/done/<uuid>-<name>
   │   (terminal)                           data/done/<uuid>-<name>.meta.json
   │                                        (sidecar-first commit)
   │  :text  move to data/text/,  failure  ───────► data/failed/<uuid>-<name>
   ▼  then enqueue (blocking backpressure)
  ┌────────────────────┐
  │  text queue        │  language enrichment
  └─────────┬──────────┘
            │ dequeue
   ┌────────┼────────┐
   ▼        ▼        ▼
 txt 1  txt 2 … txt P
   │  Languages.jl (natural language) + github-linguist (programming language)
   └─► data/text_done/<uuid>-<name>  +  data/text_done/<uuid>-<name>.meta.json
                                        (sidecar-first commit)

Stages 2 (known-file enrichment) and 3 (content triage) run in parallel: stage 1 feeds both the known and unknown queues. Stage 3 in turn feeds stage 4 (language enrichment) for every file it sorts as text.

Key properties:

  • Fast intake: the queue only ever carries small references; file bytes live on disk, so memory stays flat regardless of file size.
  • Backpressure: each queue is bounded (default 1000). When the intake queue is full, uploads get 503 Service Unavailable. When the known queue is full, the stage-1 worker blocks and retries (a classified file is never dropped).
  • Crash-resilient: files survive on disk. On startup, recovery is stage-aware: leftovers in data/spool/ re-enter classification, data/known/ re-enter enrichment, data/unknown/ re-enter content triage, and data/text/ re-enter language enrichment (recovered / recovered_known / recovered_unknown / recovered_text in the log), so a file resumes at its correct stage instead of restarting from scratch.
  • Graceful shutdown: SIGINT (Ctrl-C) and SIGTERM (systemd/Docker/k8s stop) both stop accepting uploads, then drain the stages in order — close the stage-1 queue and wait out the classify workers (the only producer of the known and unknown queues), then close those queues and wait out the enrich and content-triage workers (content triage being the only producer of the text queue), then close the text queue and wait out the language-enrichment workers. (See "Shutdown" below for one cosmetic caveat on SIGTERM.)
  • Safe filenames: client-supplied names are sanitized and prefixed with a server-minted UUID before touching the filesystem (no path traversal).

Metadata enrichment (stage 2)

Files the classifier labels known are handed to a second pool that extracts metadata with exiftool (exiftool -json -G -n) — chosen because no native Julia library comes close to its multi-format coverage. The output is normalized into a small, stable, documented schema and written as a JSON sidecar next to the file in data/done/, e.g. data/done/<uuid>-<name>.meta.json. The original bytes are never modified.

Prerequisite: exiftool must be on PATH (e.g. apt install libimage-exiftool-perl). The server fails fast at startup if it's missing.

Sidecar top-level fields (all nullable — present only when available), plus the complete raw exiftool object under raw:

Field Meaning
id, original_name job id and client-supplied name
file_type, mime_type e.g. PDF / application/pdf
file_size bytes (authoritative, from intake — not exiftool)
created_date, modified_date content timestamps
author person (Author/Artist/By-line)
created_by authoring app/tool (Producer/CreatorTool/Creator/Software/…)
dimensions {width, height} for media
duration seconds, for audio/video
page_count for documents
error set on a degraded sidecar (see below)
raw full exiftool output

Each normalized field is a coalesce over a priority list of exiftool tags (src/metadata.jl); extend a field by appending tag names. If extraction fails or exiftool times out (FS_EXIFTOOL_TIMEOUT, default 30s), the file still completes to data/done/ with a degraded sidecarfile_size/file_type plus an error note — rather than being quarantined, because it's still a wanted known file. Only genuine I/O errors (can't write the sidecar or move the file) send it to data/failed/.

The sidecar is committed before the file is moved into data/done/, so a file's presence there always implies its sidecar is already present; a crash in between leaves only a harmless orphan sidecar, and recovery re-enriches idempotently.

Content triage (stage 3)

Files the classifier labels unknown are handed to a third pool that sorts them into two coarse buckets so downstream tooling can treat them differently:

  • data/binary/ — the file looks like binary data.
  • data/text/ — the file looks like text.

The test is a UTF-8 sniff: read the first 8000 bytes and call the file text when that window is valid UTF-8 and holds no control bytes outside the text-safe set (tab, newline, CR, and friends, plus ESC for ANSI-colored logs); otherwise binary. It's cheap (no full read) and Unicode-aware — unlike the older NUL-byte or printable-ASCII heuristics, it keeps non-ASCII text (accents, CJK, emoji) in text/ instead of misfiling it, while binary formats — which rarely form valid UTF-8 near their start — still land in binary/. A NUL byte is valid UTF-8 but not a text control byte, so it still reads as binary. A multi-byte character split by the 8000-byte boundary is trimmed before the check so it isn't mistaken for malformed bytes. An empty file is treated as text. binary/ is terminal on the live path (but is the input the offline stage-5 discovery sweeps — see below); text/ is handed to stage 4 (src/content.jl).

Language enrichment (stage 4)

Files that stage 3 sorts as text are handed to a fourth pool that identifies their language and writes a .meta.json sidecar, mirroring the stage-2 known-file enrichment. Two detectors run per file:

  • natural languageLanguages.jl's LanguageDetector (a Julia port of the whatlang n-gram model) reads a bounded prefix (up to LANG_SAMPLE_BYTES, 64 KiB) and reports the language's English name, ISO 639-3 code, and a confidence in [0,1]. Pure Julia, no subprocess. The detector is built once at startup and shared read-only across the pool.
  • programming language — the github-linguist CLI recognizes source and markup by extension + content heuristics (e.g. Python, Markdown). Plain prose reports as Text and unrecognized content as null; both collapse to no programming language.

The sidecar schema:

field meaning
id, original_name from intake
file_size bytes (authoritative, from intake)
content_type always "text"
language natural-language English name (e.g. English), or null
language_code ISO 639-3 code (e.g. eng), or null
language_confidence detector confidence in [0,1], or null
programming_language e.g. Python, Markdown, or null
error set if natural-language detection produced nothing

github-linguist and the git-repo quirk: run against a path inside a git repository, linguist reads the file's committed git blob, not the on-disk bytes — and an untracked file (which everything under data/ is) has no blob, so it crashes. Stage 4 sidesteps this by copying each file to a fresh temp dir under /tmp (outside any repo, preserving the name so extension heuristics still fire) and pointing linguist there.

Programming-language detection is best-effort: if github-linguist is missing (a startup warning, not a fatal error, unlike exiftool), fails, or times out (FS_LINGUIST_TIMEOUT, default 30s), programming_language is simply null and the file still completes. Natural-language detection failing produces a degraded sidecar (with an error note) rather than a quarantine, because the file is still wanted text.

Like stage 2, the sidecar is committed before the file is moved into data/text_done/, so the file's presence there always implies its sidecar is present; recovery re-enriches idempotently (src/language.jl).

Unknown-format discovery (stage 5, offline)

The binary/ sink from stage 3 is the pile of genuinely unrecognized files. Stage 5 mines it for recurring new file formats by clustering files on their header bytes — a growing catalog of discovered formats, each with a magic-byte signature that can eventually be promoted into the classifier's fast path. Unlike stages 14 it is not on the request hot path: it is a single-owner batch process (the catalog is mutable shared state, the opposite of the stateless classifier), and because promotion is human-gated nothing here is latency-sensitive. The full rationale — and the assumptions we deliberately rejected — live in model/DESIGN_clustering.md.

The model (src/cluster.jl, base-Julia, no extra deps) is a Dirichlet-process mixture of per-position categoricals over the first 32 header bytes, on a 257-symbol alphabet (byte 0255 plus a past-EOF symbol so short fixed-length formats are modeled honestly). Bytes are treated as categorical, not numeric0x89 and 0x88 are not "close" — so this deliberately does not reuse the classifier's [0,1] byte scaling. A fixed uniform background component absorbs structureless (compressed/encrypted) blobs so they don't mint spurious clusters. A cluster's spiked positions become a libmagic-style signature; clusters with enough members and enough fixed positions self-nominate for promotion (a human does the one irreversible step, redefining "known").

Status: the offline science (phase A) is implemented and calibrated; the live catalog process (phase B) is designed and its scoring core (assign_file) is in place, but its batch-runner plumbing is not yet built.

Calibration is its own offline script (like training — never in the request path), scored against magic-collapsed ground truth (so docxzip and the whole ELF family count as one format each, which is the correct answer, not an error):

julia --project=. bin/cluster_calibrate.jl [training_set_dir]   # defaults to ../training_set

It grid-tunes the hyperparameters to maximize Adjusted Rand Index against known formats and cross-checks against a model-free NCD (gzip) baseline. On the 700-file training corpus the calibrated defaults (n=32, α=1.0, β=0.1) recover the known formats at ARI 0.77 (0.885 excluding tar), with gzip, pkzip (docx+zip merged), and jpeg forming clean, promotable clusters; the NCD baseline agrees. See DESIGN_clustering.md §11 for the full results, including the one known limitation (ELF and these tarballs share a long run of header zero- padding and merge — the v2 fix is inverse-entropy position weighting).

The queue seam (→ RabbitMQ later)

The HTTP handler and workers only ever call enqueue!, dequeue!, and close! on a JobQueue (see src/queue.jl). Today that's an in-process ChannelQueue. To move to RabbitMQ (or any broker), implement a new JobQueue subtype with those three methods and swap the construction in run — no handler or worker code changes.

Running

# install deps (first time)
julia --project=. -e 'using Pkg; Pkg.instantiate()'

# external tools: exiftool (stage 2, required) and github-linguist (stage 4,
# optional — programming-language detection). e.g. on Debian/Ubuntu:
#   apt install libimage-exiftool-perl
#   gem install github-linguist

# start the server; -t sets the number of OS threads available to workers
julia --project=. -t auto bin/server.jl

Shutdown

Both SIGINT and SIGTERM trigger the same idempotent graceful drain (stop serving → close queue → wait for workers → exit):

  • SIGINT is caught as an InterruptException (we call Base.exit_on_sigint(false)), so shutdown is clean and quiet.
  • SIGTERM can't be intercepted directly — Julia blocks it on worker threads and handles it in its own runtime, so a user signal() handler never fires. Instead we hook the drain into an atexit handler, which Julia's SIGTERM path does run. Caveat: Julia prints its own signal 15: Terminated backtrace before atexit runs. It's harmless noise — the drain still completes right after it — but if you want a fully quiet stop under a process manager, configure it to send SIGINT instead (systemd: KillSignal=SIGINT; Docker: STOPSIGNAL SIGINT). Give the stop timeout enough headroom to drain in-flight work (systemd: TimeoutStopSec).

File classifier

Each file is scored by a fixed-structure neural network (Lux.jl) that answers a single binary question: is this file known (like the types in the training set) or unknown? It's novelty detection, not exact file-typing — it won't tell you "PDF", just "this looks like something I was trained on, or not".

  • Features: the first 16 bytes + last 16 bytes of the file, each scaled 0255 → [0,1], giving a 32-dim input. Files under 32 bytes can't form that window and are classified unknown without touching the model.
  • Architecture: Dense(32→64,relu) → Dense(64→16,relu) → Dense(16→2), raw logits; decision is argmax (class 1 = known, class 2 = unknown).
  • Artifact: trained weights live in model/classifier.jld2 (committed), so the server just loads them at startup. Missing/unreadable ⇒ the server fails fast rather than run without classification.
  • Effect today: annotate-only. The class is logged (classification=known|unknown) but every file still moves to done/; the classifier can't misroute real files while it's unproven.

The architecture and byte→feature mapping are defined once in src/model.jl and shared by the trainer and the server, so they can't drift apart.

Training

Training is a separate, offline script — it never runs in the request path:

julia --project=. bin/train.jl <positives_dir> [negatives_dir]
  • positives_dir — every file in it (≥32 bytes) is a "known" example.
  • negatives_dir (optional) — a grab-bag of other real file types used as "unknown" examples. Negatives are generated ~1:1 with positives, split 50/50 between uniform-random byte vectors and grab-bag files. With no grab-bag dir, negatives are all random (weaker: the net may just learn "high entropy = unknown" rather than your actual types, so a grab-bag of real off-distribution files is recommended).

The script uses an 80/20 seeded split, reports validation accuracy, and writes model/classifier.jld2 (path overridable via FS_MODEL_PATH). A fixed seed (FS_TRAIN_SEED, default 42) drives negative generation, the split, and weight init, so the artifact is exactly regenerable from the same inputs.

Configuration (environment variables)

Variable Default Meaning
FS_HOST 127.0.0.1 Bind address
FS_PORT 8080 Port
FS_WORKERS nthreads() Stage-1 (classification) worker tasks
FS_QUEUE_CAPACITY 1000 Max pending intake jobs before 503
FS_KNOWN_WORKERS nthreads() Stage-2 (enrichment) worker tasks
FS_KNOWN_QUEUE_CAPACITY 1000 Max pending enrichment jobs (then backpressure)
FS_UNKNOWN_WORKERS nthreads() Stage-3 (content triage) worker tasks
FS_UNKNOWN_QUEUE_CAPACITY 1000 Max pending triage jobs (then backpressure)
FS_TEXT_WORKERS nthreads() Stage-4 (language enrichment) worker tasks
FS_TEXT_QUEUE_CAPACITY 1000 Max pending language jobs (then backpressure)
FS_SPOOL_DIR data/spool Incoming files (pending classification)
FS_KNOWN_DIR data/known Classified-known, awaiting enrichment
FS_UNKNOWN_DIR data/unknown Classified-unknown, awaiting content triage
FS_BINARY_DIR data/binary Stage-3 sink: unknown files that look binary
FS_TEXT_DIR data/text Classified-text, awaiting language enrichment
FS_DONE_DIR data/done Enriched known files (+ .meta.json)
FS_TEXT_DONE_DIR data/text_done Enriched text files (+ .meta.json)
FS_FAILED_DIR data/failed Files whose processing threw
FS_MODEL_PATH model/classifier.jld2 Classifier artifact loaded at startup
FS_EXIFTOOL_TIMEOUT 30 Seconds before a stuck exiftool is killed
FS_LINGUIST_TIMEOUT 30 Seconds before a stuck github-linguist is killed
FS_CLUSTER_DIR data/binary Stage-5 input: the unknown/binary pile to sweep
FS_CLUSTER_N 32 Header bytes modeled per file
FS_CLUSTER_ALPHA 1.0 CRP concentration (propensity to spawn new formats)
FS_CLUSTER_PSEUDOCOUNT 0.1 Dirichlet pseudocount β (calibrated)
FS_CLUSTER_BG_MASS 5.0 Mass of the uniform background component
FS_PROMOTE_MIN_MEMBERS 20 Cluster size threshold for promotion nomination
FS_PROMOTE_MIN_MAGIC 3 Required fixed signature positions to nominate

To get real parallelism, start Julia with enough threads (-t N) to cover all pools. If FS_WORKERS + FS_KNOWN_WORKERS + FS_UNKNOWN_WORKERS + FS_TEXT_WORKERS exceeds available threads you'll get a warning (non-fatal) and workers will share threads.

Usage

# health check
curl http://127.0.0.1:8080/health
# {"status":"ok"}

# upload one or more files (multipart/form-data)
curl -F "a=@report.pdf" -F "b=@data.csv" http://127.0.0.1:8080/upload
# 202 {"accepted":[{"id":"<uuid>","name":"report.pdf"}, ...]}

Each file in a request becomes its own job. Responses:

  • 202 Accepted — all files spooled and queued (with per-file job IDs)
  • 400 Bad Request — not multipart, or no files present
  • 503 Service Unavailable — queue full, retry later
  • 500 Internal Server Error — failed to write a file to disk

Layout

src/
  FileServer.jl   module + run() (startup, recovery, workers, serve, shutdown)
  config.jl       Config struct + env parsing
  job.jl          Job (the queue reference)
  queue.jl        JobQueue seam + in-process ChannelQueue
  spool.jl        filename sanitizing, spool/move, startup recovery
  model.jl        NN architecture + byte→feature mapping (shared with trainer)
  classify.jl     load artifact + classify a file at inference time
  metadata.jl     exiftool extraction + normalized sidecar (stage 2)
  content.jl      binary-vs-text sniff for unknown files (stage 3)
  language.jl     natural + programming language enrichment for text (stage 4)
  cluster.jl      header-byte clustering for unknown-format discovery (stage 5, offline)
  worker.jl       parametrized worker loop + classify/enrich/triage/language handlers
  server.jl       HTTP routes/handlers
bin/
  server.jl           entry point
  train.jl            offline training script → model/classifier.jld2
  cluster_calibrate.jl offline stage-5 hyperparameter calibration + NCD baseline
model/
  classifier.jld2       committed trained weights (loaded at startup)
  DESIGN_clustering.md  stage-5 design rationale + calibration results