feat(r49): D2/D3 complete and the H02 pilot is training on gx10
Entity resolution, deterministic rename augmentation, packing and the pilot trainer. Qwen3-0.6B-Base is training now: 507 steps, 11.2 s/it, ~1h35m. D2 -- gender resolution is TITLE-FIRST, and that is a change from F02's method rather than a port of it. F02 used pronoun proximity and recorded that it is structurally blind to the first-person narrator, whose name appears mainly in dialogue surrounded by other people's pronouns. Measured here, proximity called JANE MALE -- the narrator of Jane Eyre and the single worst entity to get wrong. Titles have no such blind spot: Miss Eyre, Mrs. Fairfax, Mr. Rochester, Madame Beck, M. Paul, and a 19th-century novel is saturated with them. Measured: 16 entities resolved, zero wrong, every ambiguous case landing on HELD -- shared family surnames like Helstone and Pelet genuinely belong to both a man and a woman and hold as they should. Held means ungendered, not unrenamed. A HELD entity is still renamed, from the gender-neutral surname pool, because the operator's Yarros directive was "rename all proper nouns" and holding a place leaks it -- Thornfield appears 100 times in Jane Eyre and is as author-specific as Riders Quadrant was. Substituting a neutral token makes no gender claim, so no gender claim can be wrong. D3 -- pool is French + English per the operator, weighted per work by setting: Brussels novels 60% French, Yorkshire novels 25%. Locales restricted to fr_FR/fr_BE/en_GB/en_IE; en_US and en_AU carry modern surnames that are wrong register for the 1840s. The pool is filtered against Brontë's own 75-letter alphabet, so French accents stay and Czech/Latvian marks do not. Two collision defects found by running the leak gate rather than trusting it: `Burns` and `Marie` were drawn as replacements while being Brontë characters -- F02's collision filter was built against Yarros and does not carry -- and then `Pierre-Yves` passed a whole-string filter while `Pierre` (Mademoiselle St. Pierre) is a Villette character. The filter now compares by COMPONENT. Final gate: 0 of 203 source entities survive in any of 24 copy-files. Trainer records what the run RESOLVED to rather than what it requested -- attention implementation, dtype, device, corpus sha and harness cleanliness are read back off the live objects. transformers 5.x has dropped warmup_ratio, caught by reading the signature after the first launch failed on it; the 3% warmup is computed into warmup_steps instead.
This commit is contained in:
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"run": "r49-h02-pilot",
|
||||
"base": "/home/infra-ops/carriers/Qwen3-0.6B-Base",
|
||||
"corpus": "/home/infra-ops/r49-corpus-renamed",
|
||||
"corpus_sha256_16": "3959036cf851bf62",
|
||||
"seq_len": 4096,
|
||||
"lora_rank": 32,
|
||||
"lora_alpha": 64,
|
||||
"targets": [
|
||||
"q_proj",
|
||||
"k_proj",
|
||||
"v_proj",
|
||||
"o_proj",
|
||||
"gate_proj",
|
||||
"up_proj",
|
||||
"down_proj"
|
||||
],
|
||||
"lr": 0.0001,
|
||||
"epochs": 3.0,
|
||||
"batch": 1,
|
||||
"grad_accum": 8,
|
||||
"seed": 4919,
|
||||
"train_blocks": 1349,
|
||||
"train_tokens": 5525504,
|
||||
"val_blocks": 24,
|
||||
"trainable_params": 20185088,
|
||||
"total_params": 616235008,
|
||||
"trainable_pct": 3.276,
|
||||
"steps_per_epoch": 169,
|
||||
"planned_steps": 507,
|
||||
"resolved": {
|
||||
"attn_implementation": "sdpa",
|
||||
"dtype": "torch.bfloat16",
|
||||
"device": "NVIDIA GB10",
|
||||
"torch": "2.14.0+cu130",
|
||||
"adapted_modules": 196
|
||||
},
|
||||
"harness_commit": "",
|
||||
"harness_dirty_at_launch": false,
|
||||
"launched_at": "2026-09-10T07:03:00-0700"
|
||||
}
|
||||
Reference in New Issue
Block a user