All the prompts / features currently implemented have been created with trial and error across different models. I think it’s time we create evals that allow us to iterate and see gradual improvements over time. I expect this would help us to:
- Get higher-quality results via prompt engineering
- Move towards smaller open source models while maintaining output quality
All the prompts / features currently implemented have been created with trial and error across different models. I think it’s time we create evals that allow us to iterate and see gradual improvements over time. I expect this would help us to: