Switch stage-3 triage from NUL sniff to UTF-8 validity
The NUL-byte heuristic misfiled any non-ASCII UTF-8 text (accents, CJK, emoji) as binary and let non-NUL control bytes through as text. is_binary now calls a file text when its 8000-byte sniff window is valid UTF-8 with no control bytes outside the text-safe set (tab/newline/CR/ESC/etc). - trim_truncated_utf8 drops a multi-byte char split by the window edge so it isn't mistaken for malformed bytes. - NUL still classifies as binary (valid UTF-8 scalar, non-text control). - Expanded tests: Unicode, ANSI logs, stray control byte, malformed UTF-8, boundary-split char; updated README stage-3 description.
This commit is contained in:
17
README.md
17
README.md
@@ -125,12 +125,17 @@ them into two coarse buckets so downstream tooling can treat them differently:
|
|||||||
- **`data/binary/`** — the file looks like binary data.
|
- **`data/binary/`** — the file looks like binary data.
|
||||||
- **`data/text/`** — the file looks like text.
|
- **`data/text/`** — the file looks like text.
|
||||||
|
|
||||||
The test is the classic **NUL-byte sniff** (the same heuristic `git` and
|
The test is a **UTF-8 sniff**: read the first 8000 bytes and call the file text
|
||||||
`file(1)` use): read the first 8000 bytes and, if any is NUL, call it binary,
|
when that window is valid UTF-8 and holds no control bytes outside the text-safe
|
||||||
else text. It's cheap (no full read) and reliable in practice — text encodings
|
set (tab, newline, CR, and friends, plus ESC for ANSI-colored logs); otherwise
|
||||||
don't embed NUL bytes, while binary formats almost always do near the start. An
|
binary. It's cheap (no full read) and Unicode-aware — unlike the older NUL-byte
|
||||||
empty file has no NUL, so it's treated as text. This is deliberately simple for
|
or printable-ASCII heuristics, it keeps non-ASCII text (accents, CJK, emoji) in
|
||||||
now; richer handling can hang off either bucket later (`src/content.jl`).
|
`text/` instead of misfiling it, while binary formats — which rarely form valid
|
||||||
|
UTF-8 near their start — still land in `binary/`. A NUL byte is valid UTF-8 but
|
||||||
|
not a text control byte, so it still reads as binary. A multi-byte character
|
||||||
|
split by the 8000-byte boundary is trimmed before the check so it isn't mistaken
|
||||||
|
for malformed bytes. An empty file is treated as text. Richer handling can hang
|
||||||
|
off either bucket later (`src/content.jl`).
|
||||||
|
|
||||||
## The queue seam (→ RabbitMQ later)
|
## The queue seam (→ RabbitMQ later)
|
||||||
|
|
||||||
|
|||||||
@@ -2,23 +2,64 @@
|
|||||||
#
|
#
|
||||||
# A file that stage-1 couldn't recognize is still sorted into one of two coarse
|
# A file that stage-1 couldn't recognize is still sorted into one of two coarse
|
||||||
# buckets so downstream tooling can treat them differently: `text/` for
|
# buckets so downstream tooling can treat them differently: `text/` for
|
||||||
# human-readable content, `binary/` for everything else. The test is the classic
|
# human-readable content, `binary/` for everything else. We sniff only the first
|
||||||
# "NUL byte in the first sniff window" heuristic that git and file(1) use — cheap
|
# `CONTENT_SNIFF_BYTES` (no full read) and ask two questions: does the window
|
||||||
# (no full read), and reliable in practice: text encodings don't embed NUL bytes,
|
# decode as valid UTF-8, and are any of its control bytes ones that don't belong
|
||||||
# while binary formats almost always do near the start.
|
# in text? This is the Unicode-aware successor to the classic "NUL byte" test —
|
||||||
|
# it accepts non-ASCII text (accents, CJK, emoji) instead of misfiling it as
|
||||||
|
# binary, while still rejecting binary formats, which almost never form valid
|
||||||
|
# UTF-8 near their start (and a NUL is never a valid UTF-8 scalar, so it still
|
||||||
|
# reads as binary for free).
|
||||||
|
|
||||||
const CONTENT_SNIFF_BYTES = 8000
|
const CONTENT_SNIFF_BYTES = 8000
|
||||||
|
|
||||||
|
# Control bytes (< 0x20) that appear legitimately in text: BS, TAB, LF, VT, FF,
|
||||||
|
# CR, and ESC (ANSI-colored logs). Any *other* control byte is a binary signal.
|
||||||
|
const TEXT_CONTROL_BYTES = (0x08, 0x09, 0x0a, 0x0b, 0x0c, 0x0d, 0x1b)
|
||||||
|
|
||||||
|
# Drop a trailing UTF-8 sequence that the sniff window cut in half, so a
|
||||||
|
# multi-byte character straddling the boundary isn't mistaken for invalid bytes.
|
||||||
|
# Continuation bytes are 0x80–0xBF; a lead byte encodes its own sequence length
|
||||||
|
# in its high bits. We walk back over the trailing continuation bytes, and if
|
||||||
|
# the lead byte we land on expects more bytes than the window actually holds,
|
||||||
|
# trim the whole incomplete sequence.
|
||||||
|
function trim_truncated_utf8(chunk::AbstractVector{UInt8})
|
||||||
|
n = length(chunk)
|
||||||
|
n == 0 && return chunk
|
||||||
|
# Find the start of the final byte sequence: skip back over continuations.
|
||||||
|
i = n
|
||||||
|
while i > 0 && (chunk[i] & 0xc0) == 0x80
|
||||||
|
i -= 1
|
||||||
|
end
|
||||||
|
i == 0 && return chunk # all continuations; leave as-is
|
||||||
|
lead = chunk[i]
|
||||||
|
# How many bytes does this lead byte announce?
|
||||||
|
expected = lead < 0x80 ? 1 : # ASCII
|
||||||
|
lead < 0xe0 ? 2 : # 110xxxxx
|
||||||
|
lead < 0xf0 ? 3 : # 1110xxxx
|
||||||
|
4 # 11110xxx
|
||||||
|
have = n - i + 1
|
||||||
|
return have < expected ? view(chunk, 1:i-1) : chunk
|
||||||
|
end
|
||||||
|
|
||||||
"""
|
"""
|
||||||
is_binary(path) -> Bool
|
is_binary(path) -> Bool
|
||||||
|
|
||||||
Classify a file as binary (`true`) or text (`false`) by sniffing its first
|
Classify a file as binary (`true`) or text (`false`) by sniffing its first
|
||||||
`CONTENT_SNIFF_BYTES` bytes for a NUL byte. An empty file has no NUL, so it is
|
`CONTENT_SNIFF_BYTES` bytes. A file is text when that window (minus any
|
||||||
treated as text.
|
multi-byte character truncated by the window edge) is valid UTF-8 and contains
|
||||||
|
no control bytes outside the text-safe set (`TEXT_CONTROL_BYTES`). An empty file
|
||||||
|
is treated as text.
|
||||||
"""
|
"""
|
||||||
function is_binary(path::AbstractString)::Bool
|
function is_binary(path::AbstractString)::Bool
|
||||||
open(path, "r") do io
|
open(path, "r") do io
|
||||||
chunk = read(io, CONTENT_SNIFF_BYTES)
|
chunk = read(io, CONTENT_SNIFF_BYTES)
|
||||||
return any(==(0x00), chunk)
|
isempty(chunk) && return false
|
||||||
|
window = trim_truncated_utf8(chunk)
|
||||||
|
# Malformed UTF-8 → binary.
|
||||||
|
isvalid(String(copy(window))) || return true
|
||||||
|
# Valid UTF-8, but a stray non-text control byte still means binary.
|
||||||
|
# (NUL is a valid UTF-8 scalar, so it's rejected here, not above.)
|
||||||
|
return any(b -> b < 0x20 && !(b in TEXT_CONTROL_BYTES), window)
|
||||||
end
|
end
|
||||||
end
|
end
|
||||||
|
|||||||
@@ -141,24 +141,54 @@ end
|
|||||||
end
|
end
|
||||||
end
|
end
|
||||||
|
|
||||||
@testset "is_binary: NUL-byte sniff" begin
|
@testset "is_binary: UTF-8 sniff" begin
|
||||||
mktempdir() do root
|
mktempdir() do root
|
||||||
# Plain text → text.
|
# Plain ASCII text → text.
|
||||||
txt = joinpath(root, "notes.txt")
|
txt = joinpath(root, "notes.txt")
|
||||||
write(txt, "hello, world\nsecond line\n")
|
write(txt, "hello, world\nsecond line\n")
|
||||||
@test is_binary(txt) == false
|
@test is_binary(txt) == false
|
||||||
|
|
||||||
# A NUL byte anywhere in the sniff window → binary.
|
# Non-ASCII UTF-8 (accents, CJK, emoji) is valid text — the whole
|
||||||
|
# point of moving off the printable-ASCII/NUL heuristic.
|
||||||
|
uni = joinpath(root, "unicode.txt")
|
||||||
|
write(uni, "café — 日本語 — 🚀\n")
|
||||||
|
@test is_binary(uni) == false
|
||||||
|
|
||||||
|
# ANSI-colored log: ESC + other text control bytes are text-safe.
|
||||||
|
ansi = joinpath(root, "colored.log")
|
||||||
|
write(ansi, "\e[31merror\e[0m: tab\there\r\nnext\n")
|
||||||
|
@test is_binary(ansi) == false
|
||||||
|
|
||||||
|
# A NUL byte anywhere in the sniff window → binary (it's a control
|
||||||
|
# byte outside the text-safe set, even though it's valid UTF-8).
|
||||||
bin = joinpath(root, "blob.dat")
|
bin = joinpath(root, "blob.dat")
|
||||||
write(bin, UInt8[0x01, 0x02, 0x00, 0x03])
|
write(bin, UInt8[0x01, 0x02, 0x00, 0x03])
|
||||||
@test is_binary(bin) == true
|
@test is_binary(bin) == true
|
||||||
|
|
||||||
# Empty file has no NUL → treated as text.
|
# A non-NUL, non-text control byte (e.g. 0x07 BEL) → binary.
|
||||||
|
ctrl = joinpath(root, "ctrl.dat")
|
||||||
|
write(ctrl, UInt8[UInt8('h'), UInt8('i'), 0x07])
|
||||||
|
@test is_binary(ctrl) == true
|
||||||
|
|
||||||
|
# Malformed UTF-8 (lone continuation / bad lead byte) → binary.
|
||||||
|
bad = joinpath(root, "bad.dat")
|
||||||
|
write(bad, UInt8[UInt8('a'), 0xff, 0xfe, 0xc3, 0x28])
|
||||||
|
@test is_binary(bad) == true
|
||||||
|
|
||||||
|
# A multi-byte char split by the sniff boundary must NOT read as
|
||||||
|
# binary: pad to one byte short of the window, then a 2-byte 'é'
|
||||||
|
# (0xc3 0xa9) so only its lead byte lands inside the window.
|
||||||
|
split = joinpath(root, "split.txt")
|
||||||
|
write(split, vcat(fill(UInt8('a'), FileServer.CONTENT_SNIFF_BYTES - 1),
|
||||||
|
UInt8[0xc3, 0xa9]))
|
||||||
|
@test is_binary(split) == false
|
||||||
|
|
||||||
|
# Empty file → treated as text.
|
||||||
empty = joinpath(root, "empty")
|
empty = joinpath(root, "empty")
|
||||||
write(empty, UInt8[])
|
write(empty, UInt8[])
|
||||||
@test is_binary(empty) == false
|
@test is_binary(empty) == false
|
||||||
|
|
||||||
# A NUL past the sniff window is not seen → still text.
|
# Binary garbage past the sniff window is not seen → still text.
|
||||||
far = joinpath(root, "far.txt")
|
far = joinpath(root, "far.txt")
|
||||||
write(far, vcat(fill(UInt8('a'), FileServer.CONTENT_SNIFF_BYTES), UInt8[0x00]))
|
write(far, vcat(fill(UInt8('a'), FileServer.CONTENT_SNIFF_BYTES), UInt8[0x00]))
|
||||||
@test is_binary(far) == false
|
@test is_binary(far) == false
|
||||||
|
|||||||
Reference in New Issue
Block a user