We benchmark every new model
When a new open-weight model comes out, it goes through our benchmark of 33 fixed tasks. The best model of the moment becomes the base for Loes. Later we will choose per question, depending on load, which model gives you the best answer.