Implements phase A of the DESIGN_clustering.md design: a Dirichlet-process mixture of per-position categoricals over the first 32 header bytes (257-symbol alphabet) that clusters the binary/ pile by file format, plus signature extraction and promotion nomination. All base-Julia (a Lanczos loggamma keeps the Dirichlet-multinomial marginal dependency-free). - src/cluster.jl: header_symbols feature extraction, collapsed Gibbs sampler (phase A), sequential CRP-predictive assignment (phase B core), signatures/ promotion, and ARI/V-measure calibration metrics. - bin/cluster_calibrate.jl: grid-tunes hyperparameters against magic-collapsed ground truth and cross-checks a model-free NCD (gzip) baseline. - FS_CLUSTER_*/FS_PROMOTE_* config knobs; wire cluster.jl into the module. - Tests for the three DESIGN §10 assertions plus the model primitives. Calibrated defaults (n=32, alpha=1.0, beta=0.1) recover known formats at ARI 0.77 (0.885 excl. tar); docx+zip and the ELF family merge correctly and the NCD baseline agrees. DESIGN §11 records the results and three assumptions the data corrected (tar/ELF header-zero merge, the cold-start seeding deadlock, and the Bernoulli signature / Occam-penalized restart scoring).
412 lines
23 KiB
Markdown
412 lines
23 KiB
Markdown
# FileServer
|
||
|
||
A minimal Julia service that receives files over HTTP and hands them off to a
|
||
pool of worker threads for processing. The HTTP endpoint does no real work: it
|
||
spools each uploaded file to disk, pushes a lightweight reference onto a work
|
||
queue, and responds immediately — staying free to accept the next upload.
|
||
|
||
The per-file "processing" runs each file through a small neural-network
|
||
classifier that labels it **known** (a file type resembling the training set) or
|
||
**unknown**, and logs the result. See "File classifier" below.
|
||
|
||
## Architecture
|
||
|
||
The pipeline is four stages, each with its own bounded queue and its own worker
|
||
pool (tuned independently, since classification is CPU-bound, known-file
|
||
enrichment is process-/IO-bound, content triage is cheap IO, and language
|
||
enrichment mixes CPU with a subprocess):
|
||
|
||
```
|
||
POST /upload (multipart)
|
||
│
|
||
▼
|
||
┌─────────────────┐ spool bytes to disk
|
||
│ HTTP handler │────────────────────────► data/spool/<uuid>-<name>
|
||
│ (Oxygen.jl) │
|
||
└────────┬─────────┘ enqueue reference (non-blocking)
|
||
│ │
|
||
▼ ▼
|
||
202 + job IDs ┌────────────────────┐
|
||
(503 if full) │ stage-1 queue │ classification
|
||
└─────────┬──────────┘
|
||
│ dequeue
|
||
┌───────────────────┼───────────────────┐
|
||
▼ ▼ ▼
|
||
classify wkr 1 classify wkr 2 … classify wkr N
|
||
│
|
||
┌────────────┴────────────┐
|
||
:unknown :known
|
||
│ move to data/unknown/, │ move to data/known/, then
|
||
▼ then enqueue (blocking) ▼ enqueue (blocking backpressure)
|
||
┌────────────────────┐ ┌────────────────────┐
|
||
│ unknown queue │ │ known queue │ enrichment
|
||
└─────────┬──────────┘ └─────────┬──────────┘
|
||
│ dequeue │ dequeue
|
||
┌────────┼────────┐ ┌───────────┼───────────┐
|
||
▼ ▼ ▼ ▼ ▼ ▼
|
||
unk 1 unk 2 … unk K known wkr 1 known wkr 2 … known wkr M
|
||
│ binary-vs-text sniff │ exiftool → normalized sidecar
|
||
├─► data/binary/<uuid>-<name> success ──┴──► data/done/<uuid>-<name>
|
||
│ (terminal) data/done/<uuid>-<name>.meta.json
|
||
│ (sidecar-first commit)
|
||
│ :text move to data/text/, failure ───────► data/failed/<uuid>-<name>
|
||
▼ then enqueue (blocking backpressure)
|
||
┌────────────────────┐
|
||
│ text queue │ language enrichment
|
||
└─────────┬──────────┘
|
||
│ dequeue
|
||
┌────────┼────────┐
|
||
▼ ▼ ▼
|
||
txt 1 txt 2 … txt P
|
||
│ Languages.jl (natural language) + github-linguist (programming language)
|
||
└─► data/text_done/<uuid>-<name> + data/text_done/<uuid>-<name>.meta.json
|
||
(sidecar-first commit)
|
||
```
|
||
|
||
Stages 2 (known-file enrichment) and 3 (content triage) run in parallel: stage 1
|
||
feeds both the known and unknown queues. Stage 3 in turn feeds stage 4 (language
|
||
enrichment) for every file it sorts as text.
|
||
|
||
Key properties:
|
||
|
||
- **Fast intake:** the queue only ever carries small references; file bytes live
|
||
on disk, so memory stays flat regardless of file size.
|
||
- **Backpressure:** each queue is bounded (default 1000). When the *intake* queue
|
||
is full, uploads get `503 Service Unavailable`. When the *known* queue is full,
|
||
the stage-1 worker blocks and retries (a classified file is never dropped).
|
||
- **Crash-resilient:** files survive on disk. On startup, recovery is
|
||
stage-aware: leftovers in `data/spool/` re-enter classification, `data/known/`
|
||
re-enter enrichment, `data/unknown/` re-enter content triage, and `data/text/`
|
||
re-enter language enrichment (`recovered` / `recovered_known` /
|
||
`recovered_unknown` / `recovered_text` in the log), so a file resumes at its
|
||
correct stage instead of restarting from scratch.
|
||
- **Graceful shutdown:** SIGINT (Ctrl-C) and SIGTERM (systemd/Docker/k8s `stop`)
|
||
both stop accepting uploads, then drain the stages *in order* — close the
|
||
stage-1 queue and wait out the classify workers (the only producer of the known
|
||
*and* unknown queues), then close those queues and wait out the enrich and
|
||
content-triage workers (content triage being the only producer of the text
|
||
queue), then close the text queue and wait out the language-enrichment workers.
|
||
(See "Shutdown" below for one cosmetic caveat on SIGTERM.)
|
||
- **Safe filenames:** client-supplied names are sanitized and prefixed with a
|
||
server-minted UUID before touching the filesystem (no path traversal).
|
||
|
||
### Metadata enrichment (stage 2)
|
||
|
||
Files the classifier labels **known** are handed to a second pool that extracts
|
||
metadata with [`exiftool`](https://exiftool.org/) (`exiftool -json -G -n`) —
|
||
chosen because no native Julia library comes close to its multi-format coverage.
|
||
The output is normalized into a small, stable, documented schema and written as a
|
||
JSON **sidecar** next to the file in `data/done/`, e.g.
|
||
`data/done/<uuid>-<name>.meta.json`. The original bytes are never modified.
|
||
|
||
> **Prerequisite:** `exiftool` must be on `PATH` (e.g. `apt install
|
||
> libimage-exiftool-perl`). The server **fails fast at startup** if it's missing.
|
||
|
||
Sidecar top-level fields (all nullable — present only when available), plus the
|
||
complete raw `exiftool` object under `raw`:
|
||
|
||
| Field | Meaning |
|
||
|---|---|
|
||
| `id`, `original_name` | job id and client-supplied name |
|
||
| `file_type`, `mime_type` | e.g. `PDF` / `application/pdf` |
|
||
| `file_size` | bytes (authoritative, from intake — not exiftool) |
|
||
| `created_date`, `modified_date` | content timestamps |
|
||
| `author` | person (`Author`/`Artist`/`By-line`) |
|
||
| `created_by` | authoring app/tool (`Producer`/`CreatorTool`/`Creator`/`Software`/…) |
|
||
| `dimensions` | `{width, height}` for media |
|
||
| `duration` | seconds, for audio/video |
|
||
| `page_count` | for documents |
|
||
| `error` | set on a *degraded* sidecar (see below) |
|
||
| `raw` | full `exiftool` output |
|
||
|
||
Each normalized field is a coalesce over a priority list of exiftool tags
|
||
(`src/metadata.jl`); extend a field by appending tag names. If extraction fails
|
||
or `exiftool` times out (`FS_EXIFTOOL_TIMEOUT`, default 30s), the file still
|
||
completes to `data/done/` with a **degraded sidecar** — `file_size`/`file_type`
|
||
plus an `error` note — rather than being quarantined, because it's still a wanted
|
||
known file. Only genuine I/O errors (can't write the sidecar or move the file)
|
||
send it to `data/failed/`.
|
||
|
||
The sidecar is committed **before** the file is moved into `data/done/`, so a
|
||
file's presence there always implies its sidecar is already present; a crash in
|
||
between leaves only a harmless orphan sidecar, and recovery re-enriches
|
||
idempotently.
|
||
|
||
### Content triage (stage 3)
|
||
|
||
Files the classifier labels **unknown** are handed to a third pool that sorts
|
||
them into two coarse buckets so downstream tooling can treat them differently:
|
||
|
||
- **`data/binary/`** — the file looks like binary data.
|
||
- **`data/text/`** — the file looks like text.
|
||
|
||
The test is a **UTF-8 sniff**: read the first 8000 bytes and call the file text
|
||
when that window is valid UTF-8 and holds no control bytes outside the text-safe
|
||
set (tab, newline, CR, and friends, plus ESC for ANSI-colored logs); otherwise
|
||
binary. It's cheap (no full read) and Unicode-aware — unlike the older NUL-byte
|
||
or printable-ASCII heuristics, it keeps non-ASCII text (accents, CJK, emoji) in
|
||
`text/` instead of misfiling it, while binary formats — which rarely form valid
|
||
UTF-8 near their start — still land in `binary/`. A NUL byte is valid UTF-8 but
|
||
not a text control byte, so it still reads as binary. A multi-byte character
|
||
split by the 8000-byte boundary is trimmed before the check so it isn't mistaken
|
||
for malformed bytes. An empty file is treated as text. `binary/` is terminal on
|
||
the live path (but is the input the offline **stage-5 discovery** sweeps — see
|
||
below); `text/` is handed to stage 4 (`src/content.jl`).
|
||
|
||
### Language enrichment (stage 4)
|
||
|
||
Files that stage 3 sorts as **text** are handed to a fourth pool that identifies
|
||
their language and writes a `.meta.json` sidecar, mirroring the stage-2
|
||
known-file enrichment. Two detectors run per file:
|
||
|
||
- **natural language** — [`Languages.jl`](https://github.com/JuliaText/Languages.jl)'s
|
||
`LanguageDetector` (a Julia port of the `whatlang` n-gram model) reads a bounded
|
||
prefix (up to `LANG_SAMPLE_BYTES`, 64 KiB) and reports the language's English
|
||
name, ISO 639-3 code, and a confidence in `[0,1]`. Pure Julia, no subprocess.
|
||
The detector is built once at startup and shared read-only across the pool.
|
||
- **programming language** — the [`github-linguist`](https://github.com/github-linguist/linguist)
|
||
CLI recognizes source and markup by extension + content heuristics (e.g.
|
||
`Python`, `Markdown`). Plain prose reports as `Text` and unrecognized content as
|
||
`null`; both collapse to *no programming language*.
|
||
|
||
The sidecar schema:
|
||
|
||
| field | meaning |
|
||
|---|---|
|
||
| `id`, `original_name` | from intake |
|
||
| `file_size` | bytes (authoritative, from intake) |
|
||
| `content_type` | always `"text"` |
|
||
| `language` | natural-language English name (e.g. `English`), or `null` |
|
||
| `language_code` | ISO 639-3 code (e.g. `eng`), or `null` |
|
||
| `language_confidence` | detector confidence in `[0,1]`, or `null` |
|
||
| `programming_language` | e.g. `Python`, `Markdown`, or `null` |
|
||
| `error` | set if natural-language detection produced nothing |
|
||
|
||
> **`github-linguist` and the git-repo quirk:** run against a path *inside* a git
|
||
> repository, linguist reads the file's committed git blob, not the on-disk bytes
|
||
> — and an untracked file (which everything under `data/` is) has no blob, so it
|
||
> crashes. Stage 4 sidesteps this by copying each file to a fresh temp dir under
|
||
> `/tmp` (outside any repo, preserving the name so extension heuristics still
|
||
> fire) and pointing linguist there.
|
||
>
|
||
> Programming-language detection is **best-effort**: if `github-linguist` is
|
||
> missing (a startup warning, not a fatal error, unlike `exiftool`), fails, or
|
||
> times out (`FS_LINGUIST_TIMEOUT`, default 30s), `programming_language` is simply
|
||
> `null` and the file still completes. Natural-language detection failing produces
|
||
> a **degraded sidecar** (with an `error` note) rather than a quarantine, because
|
||
> the file is still wanted text.
|
||
|
||
Like stage 2, the sidecar is committed **before** the file is moved into
|
||
`data/text_done/`, so the file's presence there always implies its sidecar is
|
||
present; recovery re-enriches idempotently (`src/language.jl`).
|
||
|
||
### Unknown-format discovery (stage 5, offline)
|
||
|
||
The `binary/` sink from stage 3 is the pile of genuinely *unrecognized* files.
|
||
Stage 5 mines it for **recurring new file formats** by clustering files on their
|
||
header bytes — a growing catalog of discovered formats, each with a magic-byte
|
||
signature that can eventually be promoted into the classifier's fast path. Unlike
|
||
stages 1–4 it is **not on the request hot path**: it is a single-owner *batch*
|
||
process (the catalog is mutable shared state, the opposite of the stateless
|
||
classifier), and because promotion is human-gated nothing here is
|
||
latency-sensitive. The full rationale — and the assumptions we deliberately
|
||
rejected — live in [`model/DESIGN_clustering.md`](model/DESIGN_clustering.md).
|
||
|
||
The model (`src/cluster.jl`, base-Julia, no extra deps) is a Dirichlet-process
|
||
mixture of **per-position categoricals** over the first 32 header bytes, on a
|
||
257-symbol alphabet (byte `0–255` plus a `past-EOF` symbol so short fixed-length
|
||
formats are modeled honestly). Bytes are treated as **categorical, not numeric**
|
||
— `0x89` and `0x88` are not "close" — so this deliberately does *not* reuse the
|
||
classifier's `[0,1]` byte scaling. A fixed uniform **background** component
|
||
absorbs structureless (compressed/encrypted) blobs so they don't mint spurious
|
||
clusters. A cluster's spiked positions become a libmagic-style signature;
|
||
clusters with enough members and enough fixed positions self-**nominate** for
|
||
promotion (a human does the one irreversible step, redefining "known").
|
||
|
||
**Status:** the offline science (phase A) is implemented and calibrated; the live
|
||
catalog process (phase B) is designed and its scoring core (`assign_file`) is in
|
||
place, but its batch-runner plumbing is not yet built.
|
||
|
||
Calibration is its own offline script (like training — never in the request
|
||
path), scored against magic-collapsed ground truth (so `docx`≡`zip` and the whole
|
||
ELF family count as one format each, which is the *correct* answer, not an error):
|
||
|
||
```bash
|
||
julia --project=. bin/cluster_calibrate.jl [training_set_dir] # defaults to ../training_set
|
||
```
|
||
|
||
It grid-tunes the hyperparameters to maximize Adjusted Rand Index against known
|
||
formats and cross-checks against a model-free NCD (gzip) baseline. On the 700-file
|
||
training corpus the calibrated defaults (`n=32`, `α=1.0`, `β=0.1`) recover the
|
||
known formats at **ARI 0.77** (0.885 excluding tar), with `gzip`, `pkzip`
|
||
(`docx`+`zip` merged), and `jpeg` forming clean, promotable clusters; the NCD
|
||
baseline agrees. See `DESIGN_clustering.md` §11 for the full results, including the
|
||
one known limitation (ELF and these tarballs share a long run of header zero-
|
||
padding and merge — the v2 fix is inverse-entropy position weighting).
|
||
|
||
## The queue seam (→ RabbitMQ later)
|
||
|
||
The HTTP handler and workers only ever call `enqueue!`, `dequeue!`, and
|
||
`close!` on a `JobQueue` (see `src/queue.jl`). Today that's an in-process
|
||
`ChannelQueue`. To move to RabbitMQ (or any broker), implement a new `JobQueue`
|
||
subtype with those three methods and swap the construction in `run` — no handler
|
||
or worker code changes.
|
||
|
||
## Running
|
||
|
||
```bash
|
||
# install deps (first time)
|
||
julia --project=. -e 'using Pkg; Pkg.instantiate()'
|
||
|
||
# external tools: exiftool (stage 2, required) and github-linguist (stage 4,
|
||
# optional — programming-language detection). e.g. on Debian/Ubuntu:
|
||
# apt install libimage-exiftool-perl
|
||
# gem install github-linguist
|
||
|
||
# start the server; -t sets the number of OS threads available to workers
|
||
julia --project=. -t auto bin/server.jl
|
||
```
|
||
|
||
## Shutdown
|
||
|
||
Both SIGINT and SIGTERM trigger the same idempotent graceful drain
|
||
(stop serving → close queue → wait for workers → exit):
|
||
|
||
- **SIGINT** is caught as an `InterruptException` (we call
|
||
`Base.exit_on_sigint(false)`), so shutdown is clean and quiet.
|
||
- **SIGTERM** can't be intercepted directly — Julia blocks it on worker threads
|
||
and handles it in its own runtime, so a user `signal()` handler never fires.
|
||
Instead we hook the drain into an `atexit` handler, which Julia's SIGTERM path
|
||
does run. Caveat: Julia prints its own `signal 15: Terminated` backtrace
|
||
*before* `atexit` runs. It's harmless noise — the drain still completes right
|
||
after it — but if you want a fully quiet stop under a process manager,
|
||
configure it to send SIGINT instead (systemd: `KillSignal=SIGINT`; Docker:
|
||
`STOPSIGNAL SIGINT`). Give the stop timeout enough headroom to drain
|
||
in-flight work (systemd: `TimeoutStopSec`).
|
||
|
||
## File classifier
|
||
|
||
Each file is scored by a fixed-structure neural network (Lux.jl) that answers a
|
||
single binary question: is this file **known** (like the types in the training
|
||
set) or **unknown**? It's novelty detection, not exact file-typing — it won't
|
||
tell you "PDF", just "this looks like something I was trained on, or not".
|
||
|
||
- **Features:** the first 16 bytes + last 16 bytes of the file, each scaled
|
||
0–255 → `[0,1]`, giving a 32-dim input. Files under 32 bytes can't form that
|
||
window and are classified `unknown` without touching the model.
|
||
- **Architecture:** `Dense(32→64,relu) → Dense(64→16,relu) → Dense(16→2)`,
|
||
raw logits; decision is `argmax` (class 1 = known, class 2 = unknown).
|
||
- **Artifact:** trained weights live in `model/classifier.jld2` (committed), so
|
||
the server just loads them at startup. Missing/unreadable ⇒ the server fails
|
||
fast rather than run without classification.
|
||
- **Effect today:** *annotate-only*. The class is logged
|
||
(`classification=known|unknown`) but every file still moves to `done/`; the
|
||
classifier can't misroute real files while it's unproven.
|
||
|
||
The architecture and byte→feature mapping are defined once in `src/model.jl` and
|
||
shared by the trainer and the server, so they can't drift apart.
|
||
|
||
### Training
|
||
|
||
Training is a separate, offline script — it never runs in the request path:
|
||
|
||
```bash
|
||
julia --project=. bin/train.jl <positives_dir> [negatives_dir]
|
||
```
|
||
|
||
- **positives_dir** — every file in it (≥32 bytes) is a "known" example.
|
||
- **negatives_dir** *(optional)* — a grab-bag of *other* real file types used as
|
||
"unknown" examples. Negatives are generated ~1:1 with positives, split 50/50
|
||
between uniform-random byte vectors and grab-bag files. With no grab-bag dir,
|
||
negatives are all random (weaker: the net may just learn "high entropy =
|
||
unknown" rather than your actual types, so a grab-bag of real off-distribution
|
||
files is recommended).
|
||
|
||
The script uses an 80/20 seeded split, reports validation accuracy, and writes
|
||
`model/classifier.jld2` (path overridable via `FS_MODEL_PATH`). A fixed seed
|
||
(`FS_TRAIN_SEED`, default 42) drives negative generation, the split, and weight
|
||
init, so the artifact is exactly regenerable from the same inputs.
|
||
|
||
## Configuration (environment variables)
|
||
|
||
| Variable | Default | Meaning |
|
||
|---------------------|----------------|------------------------------------------|
|
||
| `FS_HOST` | `127.0.0.1` | Bind address |
|
||
| `FS_PORT` | `8080` | Port |
|
||
| `FS_WORKERS` | `nthreads()` | Stage-1 (classification) worker tasks |
|
||
| `FS_QUEUE_CAPACITY` | `1000` | Max pending intake jobs before `503` |
|
||
| `FS_KNOWN_WORKERS` | `nthreads()` | Stage-2 (enrichment) worker tasks |
|
||
| `FS_KNOWN_QUEUE_CAPACITY` | `1000` | Max pending enrichment jobs (then backpressure) |
|
||
| `FS_UNKNOWN_WORKERS` | `nthreads()` | Stage-3 (content triage) worker tasks |
|
||
| `FS_UNKNOWN_QUEUE_CAPACITY` | `1000` | Max pending triage jobs (then backpressure) |
|
||
| `FS_TEXT_WORKERS` | `nthreads()` | Stage-4 (language enrichment) worker tasks |
|
||
| `FS_TEXT_QUEUE_CAPACITY` | `1000` | Max pending language jobs (then backpressure) |
|
||
| `FS_SPOOL_DIR` | `data/spool` | Incoming files (pending classification) |
|
||
| `FS_KNOWN_DIR` | `data/known` | Classified-known, awaiting enrichment |
|
||
| `FS_UNKNOWN_DIR` | `data/unknown` | Classified-unknown, awaiting content triage |
|
||
| `FS_BINARY_DIR` | `data/binary` | Stage-3 sink: unknown files that look binary |
|
||
| `FS_TEXT_DIR` | `data/text` | Classified-text, awaiting language enrichment |
|
||
| `FS_DONE_DIR` | `data/done` | Enriched known files (+ `.meta.json`) |
|
||
| `FS_TEXT_DONE_DIR` | `data/text_done` | Enriched text files (+ `.meta.json`) |
|
||
| `FS_FAILED_DIR` | `data/failed` | Files whose processing threw |
|
||
| `FS_MODEL_PATH` | `model/classifier.jld2` | Classifier artifact loaded at startup |
|
||
| `FS_EXIFTOOL_TIMEOUT` | `30` | Seconds before a stuck exiftool is killed |
|
||
| `FS_LINGUIST_TIMEOUT` | `30` | Seconds before a stuck github-linguist is killed |
|
||
| `FS_CLUSTER_DIR` | `data/binary` | Stage-5 input: the unknown/binary pile to sweep |
|
||
| `FS_CLUSTER_N` | `32` | Header bytes modeled per file |
|
||
| `FS_CLUSTER_ALPHA` | `1.0` | CRP concentration (propensity to spawn new formats) |
|
||
| `FS_CLUSTER_PSEUDOCOUNT` | `0.1` | Dirichlet pseudocount β (calibrated) |
|
||
| `FS_CLUSTER_BG_MASS` | `5.0` | Mass of the uniform background component |
|
||
| `FS_PROMOTE_MIN_MEMBERS` | `20` | Cluster size threshold for promotion nomination |
|
||
| `FS_PROMOTE_MIN_MAGIC` | `3` | Required fixed signature positions to nominate |
|
||
|
||
> To get real parallelism, start Julia with enough threads (`-t N`) to cover all
|
||
> pools. If `FS_WORKERS + FS_KNOWN_WORKERS + FS_UNKNOWN_WORKERS + FS_TEXT_WORKERS`
|
||
> exceeds available threads you'll get a warning (non-fatal) and workers will
|
||
> share threads.
|
||
|
||
## Usage
|
||
|
||
```bash
|
||
# health check
|
||
curl http://127.0.0.1:8080/health
|
||
# {"status":"ok"}
|
||
|
||
# upload one or more files (multipart/form-data)
|
||
curl -F "a=@report.pdf" -F "b=@data.csv" http://127.0.0.1:8080/upload
|
||
# 202 {"accepted":[{"id":"<uuid>","name":"report.pdf"}, ...]}
|
||
```
|
||
|
||
Each file in a request becomes its own job. Responses:
|
||
|
||
- `202 Accepted` — all files spooled and queued (with per-file job IDs)
|
||
- `400 Bad Request` — not multipart, or no files present
|
||
- `503 Service Unavailable` — queue full, retry later
|
||
- `500 Internal Server Error` — failed to write a file to disk
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/
|
||
FileServer.jl module + run() (startup, recovery, workers, serve, shutdown)
|
||
config.jl Config struct + env parsing
|
||
job.jl Job (the queue reference)
|
||
queue.jl JobQueue seam + in-process ChannelQueue
|
||
spool.jl filename sanitizing, spool/move, startup recovery
|
||
model.jl NN architecture + byte→feature mapping (shared with trainer)
|
||
classify.jl load artifact + classify a file at inference time
|
||
metadata.jl exiftool extraction + normalized sidecar (stage 2)
|
||
content.jl binary-vs-text sniff for unknown files (stage 3)
|
||
language.jl natural + programming language enrichment for text (stage 4)
|
||
cluster.jl header-byte clustering for unknown-format discovery (stage 5, offline)
|
||
worker.jl parametrized worker loop + classify/enrich/triage/language handlers
|
||
server.jl HTTP routes/handlers
|
||
bin/
|
||
server.jl entry point
|
||
train.jl offline training script → model/classifier.jld2
|
||
cluster_calibrate.jl offline stage-5 hyperparameter calibration + NCD baseline
|
||
model/
|
||
classifier.jld2 committed trained weights (loaded at startup)
|
||
DESIGN_clustering.md stage-5 design rationale + calibration results
|
||
```
|