fix(upload): drop what no name can hold BEFORE the dot rule; a cut never manufactures a kind

Heid bug hunt, hulda, second round on 92c774e:

- A lone surrogate was dropped at the final decode, after the leading-dot
  rule had already run, so "\ud800.forever" came out as .forever, the
  keep marker, and "\ud800.." as "..". The NUL and every unencodable
  character now go first, in one pass, so nothing dropped later can shield
  a dot. Starlette decodes a multipart filename strictly (utf-8, else
  latin-1), so this was not reachable over HTTP; the helper is now right by
  construction regardless.
- A suffix too long to keep was cut like text, and the cut could land on a
  shorter suffix that means something: "….png" out of "….pngxxxx…"
  became an image. A cut that changes classify/doc_kind now has its dots
  neutralised.
- The 16-byte extension threshold was unguarded (every test suffix was 4
  bytes); a .jpeg case pins it.

Falsifiers: tests/mutations/upload_names.toml, 7/7 proved. Not taken here,
as they sit in the upload route rather than this helper: the pickup-id
mkdir outside the try (a FileExistsError race), rmtree(ignore_errors)
hiding a failed cleanup, and a CancelledError skipping cleanup.
This commit is contained in:
vh
2026-09-24 17:04:35 -07:00
parent 92c774e105
commit 225ba32209
3 changed files with 87 additions and 25 deletions
+23 -14
View File
@@ -987,28 +987,37 @@ UPLOAD_NAME_MAX_BYTES = 200 # NAME_MAX is 255 bytes; the rest is _dedupe_name's
def safe_upload_name(name: str, fallback: str) -> str:
"""Reduce a client-supplied filename to a safe basename (no path, no hidden)
that the filesystem can actually hold.
that the filesystem can actually hold, and whose kind is the one it came in
with.
Two names used to reach `open()` and raise, a 500 with the booth torn down
(r3 heid bug hunt, hulda): a NUL, the one byte no POSIX name can hold
(ValueError), and a name over NAME_MAX, which is 255 BYTES — a 200-character
cap let 200 two-byte characters through (ENAMETOOLONG). The NUL goes FIRST,
so it cannot shield a leading dot from the hide rule. The cap is 200 UTF-8
bytes, leaving room for `_dedupe_name`'s suffix, and it comes out of the
STEM: the extension is what `classify` reads, so a cut `.png` would no
longer be an image. A cut through a multibyte character drops the partial
character, and a lone surrogate, which no filename can encode either, goes
the same way.
cap let 200 two-byte characters through (ENAMETOOLONG).
ORDER IS THE POINT. Everything no filename can hold goes FIRST — the NUL,
and a lone surrogate, which no UTF-8 name can encode — so nothing dropped
later can shield a leading dot from the hide rule: dropped last, the
surrogate in "\\ud800.forever" left `.forever`, which is the keep marker
(hulda, second round). Then the basename and the dot rule, then the cap:
200 UTF-8 bytes, leaving room for `_dedupe_name`'s suffix, taken out of the
STEM so the extension `classify` reads survives, and never through the middle
of a character. A suffix too long to be an extension is cut like any other
text, and a cut that lands on a SHORTER suffix meaning something else
(`….png` out of `….pngxxxx…`) has its dots neutralised, so the kind is never
manufactured.
"""
base = (name or "").replace("\x00", "").replace("\\", "/").split("/")[-1].strip()
base = (name or "").replace("\x00", "").encode("utf-8", "surrogatepass").decode("utf-8", "ignore")
base = base.replace("\\", "/").split("/")[-1].strip()
base = base.lstrip(".") # a leading dot would hide the file from every listing
stem, dot, ext = base.rpartition(".")
tail = dot + ext if stem and len((dot + ext).encode("utf-8", "surrogatepass")) <= 16 else ""
tail = dot + ext if stem and len((dot + ext).encode("utf-8")) <= 16 else ""
head = stem if tail else base
room = UPLOAD_NAME_MAX_BYTES - len(tail.encode("utf-8", "surrogatepass"))
base = (head.encode("utf-8", "surrogatepass")[:room]
+ tail.encode("utf-8", "surrogatepass")).decode("utf-8", "ignore")
return base or fallback
room = UPLOAD_NAME_MAX_BYTES - len(tail.encode("utf-8"))
cut = head.encode("utf-8")[:room].decode("utf-8", "ignore") + tail
if (classify(cut), doc_kind(cut)) != (classify(base), doc_kind(base)):
cut = cut.replace(".", "_")
return cut or fallback
def _dedupe_name(name: str, used: set) -> str: