Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If I am being honest, the value came from doing evals and testing against different models.

Essentially all I needed was a way to upload a data set, run tests against that data set and spit out a percentage of pass fail.

Braintrust makes this pretty easy, but If I was to do it again I would vibecode the same functionality.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: