A follow-up to Reasoning prefills connected a fewer unfastened models and Stolen Thoughts
This v1.1 reruns the reasoning-prefill research pinch GPT-5.5 Pro arsenic the teacher.
For each problem, I generated 2 responses from each target model:
- an ordinary, unprefilled response; and
- a consequence starting pinch the first 1% of GPT-5.5 Pro's reasoning, inserted into the target model's reasoning channel.
The visible reply remained freely generated. I past measured really overmuch of the teacher's visible reply appeared successful the first 100 tokens of the target model's answer. The array beneath reports unigram root callback truthful the numbers are comparable to my erstwhile post. Deltas are absolute percentage-point changes.
The information contains 45 problems: 15 STEM, 15 non-STEM, and 15 synthetic puzzles.
| DeepSeek V4 Flash | 45 | 40.53% | 40.89% | +0.35 pp |
| Inkling | 45 | 37.82% | 38.67% | +0.85 pp |
| Kimi K3 | 45 | 50.11% | 54.42% | +4.31 pp |
| Qwen3.8 A95B | 45 | 33.92% | 54.50% | +20.58 pp |
| STEM | 15 | 36.21% | 63.76% | +27.55 pp |
| Non-STEM | 15 | 38.26% | 52.73% | +14.46 pp |
| Puzzle | 15 | 27.28% | 47.00% | +19.72 pp |
| All | 45 | 33.92% | 54.50% | +20.58 pp |
Qwen hardly moved toward Opus 4.8 successful the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a ample effect connected the backstage synthetic puzzles. The information propose that Qwen whitethorn person learned from GPT-5.5 Pro, aliases from a intimately related GPT model, alternatively than from Opus.
Kimi K3's overlap pinch GPT-5.5 Pro is besides precocious some without and pinch the prefill (50.11% and 54.42%), though the prefill adds only +4.31 points.
English (US) ·
Indonesian (ID) ·