Our summary
Choosing an AI model starts with knowing what work it must do. Glevd describes BenchLM's selector, where a visitor writes a short description of the job. Jev reads that description and fills a questionnaire about the task, budget and requirements. The visitor can see and change every inferred answer.
The same code builds the shortlist whether a person clicks the answers or Jev supplies them. When Jev is unsure about the job, it shows alternatives and asks for confirmation; other uncertain requirements stay unanswered. Jev chooses from supplied options, including unknown. Changing an answer makes the code calculate the shortlist again.
BenchLM has not measured how often these readings are wrong or whether its cutoffs match real accuracy. One successful request is only a basic check, not a benchmark. The manual questionnaire remains available at request limits, and sending a description to a hosted service still needs a separate privacy decision.
Key takeaways
- Show people the requirements read from their description and let them change those answers.
- Keep model selection and arithmetic in code, with an unknown answer when the description does not establish a requirement.
- Test cutoffs on descriptions with known correct answers. A confidence number shows how firmly the model chose, not how often that choice is right.