
Most AI research tools point a language model at the web and trust the output. Saarth Shah built Sixtyfour around the opposite instinct: grade everything, ship only what improves the score. Saarth Shah keeps a scoreboard. Every build of Sixtyfour’s research agents is graded against questions a team…
No discussion yet. Be the first to share your thoughts!