OpenAI says GPT-6 Sol makes about half as many mistakes, on a test built from chats where the old model was caught wrong
According to OpenAI's announcement, the factuality test was built from conversations "where users flagged mistakes by our models", and the new models "are not yet available in Chat."
OpenAI released two cheaper GPT-6 models, Sol and Luna, on September 22. The claim worth slowing down on is about accuracy, and it comes with two catches.
According to OpenAI's announcement, "GPT‑6 Sol makes about half as many mistakes as its predecessor, approaching Astra-level reliability at much lower cost." Note the word approaching, not matching.
Catch one is the test. OpenAI describes it as its "internal factuality evaluation, which is based on de-identified real-world conversations where users flagged mistakes by our models". In other words, it was scored on conversations where people had already caught the older model getting a fact wrong. OpenAI adds its own caution: "These error-inducing conversations are not representative of typical usage, where factual errors are more rare."
Catch two is where the models run. The announcement says "GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users." Free users get only the lower-priced model: "Free and Go users can access GPT‑6 Luna in the desktop app." And: "These models are not yet available in Chat."
For a small business, the practical step has not changed. If you want to know what ChatGPT tells people about your business, ask it the way a customer would and check every detail it gives, from phone number to hours. A new model somewhere else in ChatGPT does not fix an answer you have not read.
OpenAI also cut developer prices, "reducing API prices for Sol and Luna by 50% compared with their GPT‑5.6 promotional pricing."
Open it and check the claim yourself. That is the point.