GPT-5.6 Sol Deep Research: 20 Real Queries, 17 Passed — What It Still Gets Wrong

Analysis·2026-09-04·Alex Chen
GPT-5.6 Sol deep research report with cited source cards on screen

The 20-Query Test Design

Deep research has been the most-hyped ChatGPT feature since it landed, so I decided to test it the way I'd test an intern: with real work queries, not demo questions. I ran 20 queries across five buckets: market research (4), technical research (4), pricing and procurement (4), academic-style literature reviews (4), and competitive intelligence (4). Every query was something I actually needed an answer to in the last two months, so I had independent knowledge of most answers already — which is exactly the setup that exposes fabrication.

Each query got one run in deep research mode with no follow-ups, the way a busy user would actually use it. I scored on four axes: answer quality (is the conclusion right), source quality (are the citations real and relevant), freshness (does it use current data), and actionability (could I act on the report without more digging).

GPT-5.6 Sol Deep Research: 20 Real Queries, 17 Passed — What It Still Gets Wrong

The 17 Wins: What Deep Research Does Well

The wins are not subtle. Query one: 'what changed in the small-business payroll software market in the last 12 months' produced a 2,800-word report with 31 citations, correct market-share framing, and a summary of three acquisitions I had independently tracked. Query six: a literature review on graph-based RAG techniques returned 19 papers, 17 of which I'd have included myself, with correct one-paragraph summaries of each. That's genuinely better than the first pass most human researchers would produce.

The pricing bucket was the standout — procurement-style queries ('what does enterprise observability tooling actually cost in 2026') returned structured comparison tables with real vendor pricing pages cited. I verified seven of eight numbers directly. The competitive-intelligence bucket was the weakest of the winning group: conclusions were sound but it leaned heavily on vendor blogs and press releases, which is exactly what a junior analyst would do before learning to find primary sources.

The 3 Fails: One Root Cause

All three failures share a single root cause: confident answers on topics where the recent public record is thin. Query: 'which mid-sized European cities are building the most EV charging infrastructure per capita in 2026' — the report came back with confident rankings and a table of numbers. I checked three of the cited sources: one was a 2023 report the model had quietly extrapolated from, one was a regional tourism page that mentioned charging stations in passing, and one URL did not exist at all. The rankings were plausible and wrong.

Failure two was a query about a niche open-source license dispute that broke in the last 30 days. The report described the dispute accurately up to a point — then confidently narrated an 'outcome' that hadn't happened, citing a developer's GitHub issue as if it were a resolution. Failure three was an obscure regulatory question where the model synthesized two unrelated rulings into a single fictional precedent. The pattern is consistent: when recent, specific information is scarce, deep research fills the gap with the most plausible story, complete with citations. Plausibility is not accuracy — that's the sentence to tattoo somewhere visible.

GPT-5.6 Sol Deep Research: 20 Real Queries, 17 Passed — What It Still Gets Wrong

The Citation Audit: 86% Real, 9% Stale, 5% Fabricated

I audited all 374 citations across the 20 reports. 86% pointed to real pages that said what the report claimed they said — that's the headline, and it's genuinely impressive. 9% were real pages that were outdated relative to the report's claim, usually a year or more old being used to support a present-tense statement. 5% were fabricated: URLs that returned 404 or never existed, typically inserted in the exact spot where a real citation was missing. The fabrication rate varied wildly by bucket — near zero in pricing queries with rich vendor pages, up to 14% in the thin-record regulatory queries.

The practical implication: citation quality is a leading indicator of report quality. In the 17 winning reports, I found one fabricated citation total. In the 3 failures, 18. If your report's sources start looking thin or oddly generic, the conclusions deserve double scrutiny. I now run every deep research report through a quick spot check: I open five random citations and confirm they say what the report claims. Five minutes, and it catches the failure mode every time.

Cost and Speed vs Doing It Yourself

The economics are the real story. My 20 queries averaged 4.2 minutes and roughly 900 tokens of output per minute of thinking — a typical report ran 2-8 minutes and maybe 15,000-25,000 output tokens. On Plus-tier pricing, that's a fraction of a cent per report. The same work manually: I estimated 2-3 hours per query for the market and pricing buckets, because most of that time is tab-opening, skim-reading, and deciding which sources matter. At any reasonable hourly rate, deep research pays for itself on the first report of the week.

One honest caveat on speed: the reports that ran longest (7-8 minutes) were not meaningfully better than the 3-4 minute ones. Depth setting matters more than patience. And the reports degrade gracefully under interruption — you can ask follow-ups after delivery, but each follow-up restarts the search loop rather than continuing it, which burns another minute or two on already-answered questions.

When to Trust It, Not To

After 20 queries I have a clean mental model. Trust it without much review for: market landscapes, vendor and pricing research, literature discovery, competitor feature tracking, and any question where lots of recent public sources exist. The richer the source ecosystem, the better the report — this model is a mirror of the web in the same way everything else is. Spot-check it always for: regulatory questions, anything where a wrong answer has legal or financial cost, fast-moving niche topics, and any question where you suspect the public record is thin. Those are exactly the conditions where the 5% fabrication rate clusters.

My standard workflow now: run deep research for the landscape pass, then verify the three or four load-bearing claims manually, then write. Total time for a topic that used to take me an afternoon: about 45 minutes, with materially better source coverage. If you're already using Sol for data work, pair deep research with the data analysis test to see how it handles the numbers side, and the context window stress test to understand how much raw material it can hold before quality slips.

Verdict

GPT-5.6 Sol's deep research is the first AI research feature I'd pay for out of my own pocket. Seventeen of twenty queries returned work that would have taken me hours, with citations that mostly check out and conclusions that mostly hold. It has replaced my first-pass research entirely — I no longer open twenty tabs to understand a new topic; I open the report, then verify.

The boundary is the thin-record failure mode. When the public web has the answer, deep research is extraordinary. When it doesn't, the model invents a plausible answer with confident citations — and nothing in the interface warns you which situation you're in. Keep the five-minute citation spot check in your workflow, treat regulatory and fast-moving topics with suspicion, and you have a research assistant that would cost $60,000 a year as a human. For everything else, it's the feature that finally makes the phrase 'do the research' mean 'ask Sol'.

Frequently Asked Questions

Does GPT-5.6 Sol have a deep research mode?

Yes — deep research is built into GPT-5.6 Sol on Plus and Pro plans. It runs iterative web searches, reads and cross-references sources, and compiles a cited report instead of answering from memory. Reports typically take 2-8 minutes depending on depth.

How accurate is GPT-5.6 Sol deep research?

In my 20-query test, 17 of 20 reports were genuinely useful and 86% of the citations pointed to real, relevant pages. The failure mode to watch is stale data: it prefers authoritative-looking older sources when recent information is scarce, and it occasionally fabricates a URL to fill a gap.

Is deep research worth the cost?

For the time it saves, yes. My 20 queries averaged 4.2 minutes each and would have taken me roughly 2-3 hours per query manually. At Plus-tier pricing the cost per report is trivial compared to a paid research assistant.

Can deep research replace a human researcher?

Not yet. It misses nuance in conflicting sources, and its citation failures mean every critical claim needs a spot check. Treat it as a world-class research intern who works fast but needs review before anything ships.

A
Alex Chen