* fix: Check response prefix.
In case if the model adds <think> at the end of user prompt and only generates </think> end, then the detection goes wrong.
* fix: Handle whitespace in CoT, remove redundant checks
* fix: Type checker
* fix: try looking at tag positions, not text end
regex
* ruff, fix import sorting
* fix: rich markup
* fix: I missed other ones
* feat: Update SHA256SUMS file hashes in the tests.
A major change that affects reproducibility.
* fix: It is now sensible to also update the extra SHA256SUMS.ci2 file.
* fix: Only consider whitespace, no other text or instructions.
because mistral-3 as additional reasoning instructions in its chat template. And I suppose many other models can have it too.
* fix: Update windows hashes.
* fix: Update CI hashes too
* as always update case two of mistral-3 (ci2 hash)
* docs: update comment
* feat: Handle the edge case for models having additional instructions.
* fix: Update windows hash for mistral-3
* docs: remove a line from the comments
because I'm not sure about GPT-OSS models' thinking tags and it cannot be confirmed using an untrained tiny GPT-OSS model. And inference fallback would of course generate gibberish as the model cannot understand additional instructions about 'how to generate response and how to think' from the chat_template.
* fix: a few things.
* fix: Update hash for qwen3.5 after the whitespace fix for its response prefix.
* docs: Update comment
* fix: Update qwen3.5 hash for CI
* fix: Remove Case 2 which only serves tests
unnecessary
* fix: Hash
* fix: concern is valid enough, so we use a small text.
add a comment too
* docs: minor
* fix: minor change
use `W_org` matrix where needed...
* Update model.py
* Update model.py
* fix: Windows hash, remove BOM marker
* docs: Add info about test cases
* feat: Tests for row_normalization PRE & NONE
* feat: CI hash files for row_normalization PRE & NONE models
* feat: Documentation instructions about test suite
* add recommendation
* fix: remove notebook input shims
Closes#280
* feat: support headless operation (no interactive input)
* fix: prevent infinite loops
* feat: add end-to-end tests
* ci: run tests in CI
* ci: fix test output ordering
* fix: replace home-cooked `set_seed` function with Transformers builtin
* feat: print PyTorch config when running tests
* feat: print additional information
* experiment: try to standardize test environment
* fix: revert environment changes
* feat: support multiple valid hashes for each output file
* feat: add test output hashes for CI
* feat: add test output hashes for CI (alternative environment)
* feat: add hashes for Windows (#394)
* fix: Hash on windows
* trigger ci
* fix: prefer .yaml (used widely than .toml for model configs)
* use removeprefix
* docs: restore commet
* use removeprefix again
* tests: Add windows hash files for all test models
* trigger ci
* fix: minor cleanup
* clean merge mismatch
* remove unnecessary CRLF replace, now that we support more SUMS files
* fix: use binary mode for hashes everywhere
---------
Co-authored-by: Vinay Umrethe <umrethevinay@gmail.com>