Tweet by zeeg
September 5, 2025
I think the fundamental issue is the evals are simply too unreliable and not scientific enough. Yes you could build crazy statistical test suite but iteration is a nightmare because don’t have a clue why shit fails. This is a real issue with LLM-based dev imo. https://t.co/aTezXJzY6S
- Author
- zeeg
- Date
- September 5, 2025
- Canonical URL
- /tweets/zeeg-1963814682097274931-1e389c