375244ad05
Entity resolution, deterministic rename augmentation, packing and the pilot trainer. Qwen3-0.6B-Base is training now: 507 steps, 11.2 s/it, ~1h35m. D2 -- gender resolution is TITLE-FIRST, and that is a change from F02's method rather than a port of it. F02 used pronoun proximity and recorded that it is structurally blind to the first-person narrator, whose name appears mainly in dialogue surrounded by other people's pronouns. Measured here, proximity called JANE MALE -- the narrator of Jane Eyre and the single worst entity to get wrong. Titles have no such blind spot: Miss Eyre, Mrs. Fairfax, Mr. Rochester, Madame Beck, M. Paul, and a 19th-century novel is saturated with them. Measured: 16 entities resolved, zero wrong, every ambiguous case landing on HELD -- shared family surnames like Helstone and Pelet genuinely belong to both a man and a woman and hold as they should. Held means ungendered, not unrenamed. A HELD entity is still renamed, from the gender-neutral surname pool, because the operator's Yarros directive was "rename all proper nouns" and holding a place leaks it -- Thornfield appears 100 times in Jane Eyre and is as author-specific as Riders Quadrant was. Substituting a neutral token makes no gender claim, so no gender claim can be wrong. D3 -- pool is French + English per the operator, weighted per work by setting: Brussels novels 60% French, Yorkshire novels 25%. Locales restricted to fr_FR/fr_BE/en_GB/en_IE; en_US and en_AU carry modern surnames that are wrong register for the 1840s. The pool is filtered against Brontë's own 75-letter alphabet, so French accents stay and Czech/Latvian marks do not. Two collision defects found by running the leak gate rather than trusting it: `Burns` and `Marie` were drawn as replacements while being Brontë characters -- F02's collision filter was built against Yarros and does not carry -- and then `Pierre-Yves` passed a whole-string filter while `Pierre` (Mademoiselle St. Pierre) is a Villette character. The filter now compares by COMPONENT. Final gate: 0 of 203 source entities survive in any of 24 copy-files. Trainer records what the run RESOLVED to rather than what it requested -- attention implementation, dtype, device, corpus sha and harness cleanliness are read back off the live objects. transformers 5.x has dropped warmup_ratio, caught by reading the signature after the first launch failed on it; the 3% warmup is computed into warmup_steps instead.
41 lines
923 B
JSON
41 lines
923 B
JSON
{
|
|
"run": "r49-h02-pilot",
|
|
"base": "/home/infra-ops/carriers/Qwen3-0.6B-Base",
|
|
"corpus": "/home/infra-ops/r49-corpus-renamed",
|
|
"corpus_sha256_16": "3959036cf851bf62",
|
|
"seq_len": 4096,
|
|
"lora_rank": 32,
|
|
"lora_alpha": 64,
|
|
"targets": [
|
|
"q_proj",
|
|
"k_proj",
|
|
"v_proj",
|
|
"o_proj",
|
|
"gate_proj",
|
|
"up_proj",
|
|
"down_proj"
|
|
],
|
|
"lr": 0.0001,
|
|
"epochs": 3.0,
|
|
"batch": 1,
|
|
"grad_accum": 8,
|
|
"seed": 4919,
|
|
"train_blocks": 1349,
|
|
"train_tokens": 5525504,
|
|
"val_blocks": 24,
|
|
"trainable_params": 20185088,
|
|
"total_params": 616235008,
|
|
"trainable_pct": 3.276,
|
|
"steps_per_epoch": 169,
|
|
"planned_steps": 507,
|
|
"resolved": {
|
|
"attn_implementation": "sdpa",
|
|
"dtype": "torch.bfloat16",
|
|
"device": "NVIDIA GB10",
|
|
"torch": "2.14.0+cu130",
|
|
"adapted_modules": 196
|
|
},
|
|
"harness_commit": "",
|
|
"harness_dirty_at_launch": false,
|
|
"launched_at": "2026-09-10T07:03:00-0700"
|
|
} |