How we pick the models behind Front Door
We tested 32 models on 468 made-up cases. A model now has to pass our safety checks before it can answer you.

Front Door now tests the models that write its answers and read its letters. A new model has to pass before it can answer you, and the team checks the results each week on a private report.
- 468 made-up cases: questions in all 50 states and DC, in Spanish, Arabic, Vietnamese, Chinese and six other languages, people in crisis, immigrant families with no Social Security number, 60 agency letters, and short follow-ups. No real person’s words are used.
- Every answer is checked. Crisis lines must come first. Phone numbers, links and dollar amounts must come from what we looked up, never from a model’s memory. A personal number is never repeated, and no answer promises that someone qualifies. A model that misses any of these on any case does not answer residents, however well it does otherwise.
- A second, stronger model reads the answers and scores how helpful, accurate, plain and warm they are, and says why.
- Your words, letters and documents only go to providers that neither keep them nor train on them. Every call that carries them now asks for that, in code.
- New models, including free ones, are tested every week. When one is as good and cheaper or faster, the team gets a proposal to review. Nothing changes on its own.
- The test found a bug in how we hide names before a question is read: “my partner hits me” came out as “my partner [name]”. It is fixed, so the danger in those words now reaches the step that reads them.