Qwen3.5-4B Base Blind Spots

A careful evaluation framework that tests Qwen3.5-4B-Base with 84 structured prompts in 12 reasoning categories. It labels failures on two axes and found a critical float16 overflow issue.

What changed

  • Showed that float16 overflows in all 12 categories and makes the output unreadable, so a float16 evaluation of a Qwen-class model is invalid. bfloat16 is required.
  • Turned vague quality concerns into failure rates for each category, with a fine-tuning dataset and scale range mapped to each root cause.

What I worked on

  • I designed the prompt suite and the two-axis labels that separate coherence failures from correctness failures.
  • I published the outputs for both dtypes, with auto-labels, as a dataset on Hugging Face.

A handful of failure examples says little about a model. It doesn't tell you how often the model fails, or whether the fix is more data or a different setting. So instead of collecting examples, this project runs a small, structured suite against Qwen3.5-4B-Base, a 4B-parameter pre-trained model with Gated DeltaNet attention and a sparse MoE.

Each of the 12 reasoning categories, including arithmetic, logic, commonsense, coreference, Bayesian reasoning and Winograd schemas, gets 7 prompts: 5 that probe for a failure and 2 controls the model should pass. For a targeted fine-tune, that is enough to find the repeat blind spots faster than a large benchmark would.

Two labels, two kinds of fix

Every output gets two separate labels: is it readable (coherence), and is it right (correctness)? The split matters because the fixes are different. Coherence failures point to inference settings or instruction tuning. Correctness failures point to chain-of-thought or domain-specific data. Mixing them up wastes fine-tuning effort.

The dtype finding

In float16, the model failed across every category, even on the easy controls. The cause was numerical overflow, not reasoning. For this model, float16 isn't a small drop in quality: it makes every result invalid. The float16 outputs stay in the dataset as a reference, and the lesson is to check the dtype before trusting any evaluation.

Built with Hugging Face Transformers and greedy decoding for reproducible runs, on Google Colab (L4 or A10G) or Modal. Prompts, full records and labelled results are stored as JSONL.