LLMs will show lower accuracy on float addition involving negative numbers, combining decimal handling with sign handling challenges.
Combines two challenges that individually produced no errors: decimal handling (positive floats) and sign handling (negative integers). Tests whether the combination introduces errors.
100 addition problems with negative floats (absolute values 0.01–99.99, up to 2 decimal places, seed=42), stratified into 4 cases:
One negative, positive result (30 pairs)
One negative, negative result (30 pairs)
One negative, zero result (10 pairs): b = -a
Both negative (30 pairs)
Model: gpt-oss-120b via Cerebras API (free tier)
Temperature: 0, top_p: 1, max_completion_tokens: 1024
Prompt: "What is {a} + {b}? Reply with only the number."
Correctness: round(model, 2) == round(expected, 2) using Decimal arithmetic
REJECTED
100% accuracy across all 100 problems. The combination of decimal handling and sign handling produced zero errors.
| Case | Total | Correct | Accuracy |
|---|---|---|---|
| One negative, positive result | 30 | 30 | 100.0% |
| One negative, negative result | 30 | 30 | 100.0% |
| One negative, zero result | 10 | 10 | 100.0% |
| Both negative | 30 | 30 | 100.0% |
Combined with all prior arithmetic experiments, gpt-oss-120b is now 600/600 on basic two-operand arithmetic (addition and multiplication, integers and floats, positive and negative) with operands below 100 in absolute value.
- Test float subtraction and division with negative operands
- Scale up decimal places to find precision limits
- Run on a non-reasoning model to see if chain-of-thought is carrying the result