Qwen3.5-9B Accuracy
DSPy
accuracy (%)
DSPy Qwen: lenient×W 3.3, 9.0, 11.6, 28.0; strict×CC 0.0, 0.0, 2.0, 0.8.
max ReAct iterations
OpenClaw
accuracy (%)
OpenClaw Qwen: lenient×W 2.3, 6.9, 15.8, 17.6; strict×CC 0.0, 0.0, 6.6, 6.4.
max ReAct iterations
lenient grading strict grading | no early exit @ 20
Figure 4. Qwen3.5-9B accuracy across inference budgets under strict and lenient grading.