Running Featured 38 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 38 Who needs 1T parameters? Olympiad proofs with a 4B model
Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Paper • 2602.01058 • Published 19 days ago • 41