A handful of failure examples says little about a model. It doesn't tell you how often the model fails, or whether the fix is more data or a different setting. So instead of collecting examples, this project runs a small, structured suite against Qwen3.5-4B-Base, a 4B-parameter pre-trained model with Gated DeltaNet attention and a sparse MoE.
Each of the 12 reasoning categories, including arithmetic, logic, commonsense, coreference, Bayesian reasoning and Winograd schemas, gets 7 prompts: 5 that probe for a failure and 2 controls the model should pass. For a targeted fine-tune, that is enough to find the repeat blind spots faster than a large benchmark would.
Two labels, two kinds of fix
Every output gets two separate labels: is it readable (coherence), and is it right (correctness)? The split matters because the fixes are different. Coherence failures point to inference settings or instruction tuning. Correctness failures point to chain-of-thought or domain-specific data. Mixing them up wastes fine-tuning effort.
The dtype finding
In float16, the model failed across every category, even on the easy controls. The cause was numerical overflow, not reasoning. For this model, float16 isn't a small drop in quality: it makes every result invalid. The float16 outputs stay in the dataset as a reference, and the lesson is to check the dtype before trusting any evaluation.
Built with Hugging Face Transformers and greedy decoding for reproducible runs, on Google Colab (L4 or A10G) or Modal. Prompts, full records and labelled results are stored as JSONL.