Skip to content
CoronRing
All posts

· 7 min read

Stop taking AI advice on where to submit your paper

I asked GPT-6 Luna where to submit 132 blinded AI papers whose 2026 venues it could not have seen. It never suggested an oral, sent half the workshop papers to main conferences, and sounded just as sure when wrong.

advice matched the tier (51) advice missed (81) One line per paper. GPT-6 Luna, asked "which of these should I submit to?" for 132 blinded 2026 papers.

A growing number of researchers paste a draft into a chatbot and ask some version of “is this NeurIPS material?” The answer comes back fluent, specific and reasonable. I wanted to know whether it is also right.

So I took 132 papers from 2026 AI venues, stripped everything that could identify where they were published, and asked a current model where each one should be submitted. Then I compared its advice with where the papers actually landed.

orals and spotlights told to aim that high
0 of 29
right venue among main-conference papers
20%
25% by picking at random
average confidence when right / when wrong
0.42 / 0.40

The short version

  • It never once recommended aiming for an oral or spotlight, including for the 29 papers that received one.
  • Among main-conference papers, it named the right venue 20% of the time. Picking at random among the four main options would get 25%.
  • It sent half of the Findings and workshop papers to main conferences.
  • Its confidence was the same when it was right (0.42) and when it was wrong (0.40).
  • Asking it to review the paper carefully and estimate acceptance odds first made the advice worse, not better.
  • Its acceptance estimates rank papers by quality about as well as counting the words in them.

The setup

The papers. Four tiers of 2026 venues, all AI and LLM work:

TierWhere the paper was acceptedPapers
1NeurIPS 2026, oral or spotlight29
2NeurIPS 2026 poster, EMNLP 2026 main, COLM 202636
3Findings of EMNLP 202632
4Workshops at ICML 2026 and NeurIPS 202635

NeurIPS decisions come from the conference’s official accepted-paper list. EMNLP, COLM and workshop labels come from the authors’ own arXiv comments (“Accepted to Findings of EMNLP 2026”), because those venues had not published full lists yet.

No memorization. The evaluator was GPT-6 Luna at high reasoning effort, with a knowledge cutoff of May 18, 2026. Every venue above announced its decisions after that date, and every paper was first posted to arXiv after it. The model could not have read these papers or learned where they went.

Blinding. Before the model saw a paper I removed author names, affiliations, links, acknowledgements, references, appendices, and venue-specific sections such as limitations statements and checklists. Venue names in the text were replaced, and template leftovers that name a workshop were deleted. The model was also told not to reason from submission deadlines or formatting, which in a dataset like this would leak the answer.

Two ways of asking.

  1. Just asking. “I’m planning to submit this paper in the 2026 cycle. These are the options, from most to least selective. Which one should I submit to?” This is what most people actually type.
  2. Asking like a careful advisor. Write a full program-committee review first, then estimate the probability of acceptance at each option, then recommend the most selective option with roughly even odds or better.

Just asking

ResultFor comparison
Right tier39%25% random, 27% for always saying “tier 2”
Right venue, main-conference papers only20%25% random
Main conference vs not68%50% coin flip
Orals and spotlights told to aim that high0 of 29
Findings and workshop papers told to aim for a main conference34 of 67
Same answer when asked a second time28 of 31

The 39% looks better than chance until you see where it comes from. The model told 69% of all papers to aim for a main conference. That default is right for papers that are main-conference papers, so it scores well on tier 2 (29 of 36) and badly everywhere else: 0 of 29 orals, 13 of 32 Findings papers and 9 of 35 workshop papers. It aimed too high 44 times and too low 37 times.

Two more things stood out.

The confidence is decoration. On average the model reported 0.42 when it picked the right tier and 0.40 when it picked the wrong one. A number that does not move with being right tells you nothing.

The advice follows topic, not quality. None of the 40 papers in arXiv’s computation-and-language category was told to try NeurIPS. Seven of the eleven computer-vision papers were. The model has a sense of which conference a topic “belongs” to, and that sense drives the recommendation.

It is also consistent. Asked again, it gave the same venue for 28 of 31 papers. That makes it more persuasive, not more accurate.

Asking more carefully made it worse

The careful-advisor prompt got the tier right 30% of the time, down from 39%, and recommended a main conference for only 3 of the 65 papers that were main-conference papers.

The cause is instructive. The model’s acceptance estimates are realistic in aggregate. It put a typical NeurIPS paper’s chances at roughly 12% to 23%, and NeurIPS 2026 actually accepted 25.7% of submissions. But with odds like that, almost no paper clears “even odds or better” at a top venue, so 127 of the 132 papers were steered to Findings or a workshop.

I then tried replacing the model’s threshold with the venues’ real acceptance rates: aim for a venue whenever the model’s probability beats that venue’s actual rate. That overshoots in the other direction. Across all 132 papers, the model’s chance of an oral had a median of 2% and stayed within a narrow 0.1% to 6%, while the real rate is about 1.3% of submissions. So 83 of 132 papers came out as “go for an oral”.

Neither threshold works because the numbers are squeezed into a narrow band. The model hedges toward the middle, and no cutoff turns that into good advice.

Each way of asking piles the papers into a different tier, and none matches where they were accepted
Just asking
Careful advisor
Advisor + real acceptance rates
Oral / spotlight
0
0
83
Main conference
91
5
5
Findings
29
76
10
Workshop
12
51
34
Where the advice sent papers Where they were accepted
Grey: where the 132 papers were actually accepted. Accent: where the model sent them. 'Careful advisor' recommends the most selective option with even odds or better; the third panel replaces that threshold with each venue's real acceptance rate.
Show the data
Tier Accepted Just askingCareful advisorAdvisor + real acceptance rates
Oral / spotlight 29 0083
Main conference 36 9155
Findings 32 297610
Workshop 35 125134

Is there any signal at all?

A recommendation can be bad while the underlying judgment still has some value, so I also checked whether the model’s acceptance probabilities at least rank papers correctly. The measure is AUC: pick one paper from the stronger group and one from the weaker group, and AUC is the chance the model gives the stronger one the higher number. 0.5 is a coin flip, 1.0 is perfect. As a baseline I ranked the same papers by their word count, nothing else.

Word count ranks papers about as well as the model's own acceptance estimates
Main conference vs Findings and workshops 65 vs 67 papers · model 0.60 · words 0.65
EMNLP main vs Findings of EMNLP 11 vs 32 papers · model 0.53 · words 0.52
NeurIPS oral vs NeurIPS poster 29 vs 14 papers · model 0.76 · words 0.58
NeurIPS oral vs NeurIPS workshop paper 29 vs 13 papers · model 0.55 · words 0.75
Model's acceptance probability Word count alone
AUC is the chance that a paper from the better group gets the higher score; 0.5 is a coin flip. Lines show 95% bootstrap intervals, and an interval that crosses 0.5 is consistent with no signal at all.
Show the data
Comparison Papers Model AUC Word count AUC
Main conference vs Findings and workshops 65 vs 67 0.60 (0.51 to 0.70) 0.65 (0.55 to 0.74)
EMNLP main vs Findings of EMNLP 11 vs 32 0.53 (0.35 to 0.71) 0.52 (0.31 to 0.74)
NeurIPS oral vs NeurIPS poster 29 vs 14 0.76 (0.61 to 0.90) 0.58 (0.39 to 0.77)
NeurIPS oral vs NeurIPS workshop paper 29 vs 13 0.55 (0.35 to 0.75) 0.75 (0.57 to 0.90)

For the coarse question, whether a paper belongs at a main conference at all, word count does slightly better than the model. For main vs Findings at the same conference, which is the comparison that isolates reviewers’ judgment of quality, the model is at chance.

The oral-vs-poster number, 0.76, looks like a real signal, and it is the only one in the data. The last row explains it. The model gives NeurIPS workshop papers almost exactly the same oral probability as actual orals (2.5% vs 2.7% on average), and word count tells those two groups apart better than the model does (0.75 vs 0.55). What the model recognizes is a paper that reads like NeurIPS-style machine learning, not a paper that clears the oral bar.

Orals get a slightly higher oral probability than posters, and NeurIPS workshop papers get almost the same
NeurIPS oral or spotlight 29 papers · mean 2.7%
NeurIPS poster 14 papers · mean 1.5%
NeurIPS workshop paper 13 papers · mean 2.5%
Orals and spotlights Posters and workshop papers
Each dot is one paper's predicted chance of an oral or spotlight, if submitted to NeurIPS. The vertical line is the real rate, about 1.3% of submissions. Means: orals 2.7%, workshop papers 2.5%, posters 1.5%.
Show the data
Group Papers Mean Every value (%)
NeurIPS oral or spotlight 29 2.67% 0.5, 1, 1, 1, 1, 1.5, 1.5, 1.5, 2, 2, 2.5, 2.5, 2.5, 2.5, 2.5, 3, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 6
NeurIPS poster 14 1.53% 0.1, 0.3, 1, 1, 1, 1, 1, 1.5, 1.5, 2, 2, 2.5, 2.5, 4
NeurIPS workshop paper 13 2.46% 0.5, 1, 1, 1.5, 2, 2, 2, 2, 3, 4, 4, 4, 5

What this does not show

  • Every paper here was accepted somewhere. This tests telling good from great, not predicting acceptance or rejection, and a model could still be useful for catching papers that are not ready at all.
  • Labels for EMNLP, COLM and workshops are self-reported by authors.
  • The full experiment used one model. A smaller pilot with Claude Opus 5.5 on 37 of the papers pointed the same way: 24% exact venue among main-conference papers, and no tier-1 recommendation under any prompt. Larger models may do better.
  • I did not test whether the model’s written reviews, its lists of strengths and weaknesses, are useful. Only its venue judgment.
  • Topic differs across tiers, since NeurIPS papers lean toward machine learning and vision and EMNLP papers toward language. That helps the model on any comparison that crosses conferences, so the weak results above are, if anything, generous.

What to do instead

The verdict is the part to ignore. “This looks like a strong NeurIPS poster” sounds informed, but in this experiment it carried about as much information as the paper’s length.

If you want help deciding where to submit:

  • Ask people who review for the venue. They have seen the distribution the decision is drawn from. The model has not.
  • Read last year’s accepted papers at the venue you are considering, especially the orals. Compare the size of the claim and the depth of the evidence, not the topic.
  • Use the model for the parts that do not need a verdict. It writes a plausible review. Whether its listed weaknesses are the ones reviewers will raise is a separate question this experiment did not answer, but checking a draft against them costs little.
  • Ignore the confidence number. Here it was the same whether the model was right or wrong.

The model answers the question “where should I submit?” with the same fluency it brings to everything else. That fluency is the problem: it reads like judgment, and on this question it mostly is not.

Method notes: GPT-6 Luna, reasoning effort high, about 3M tokens across 132 papers and two prompts, run in early October 2026. Acceptance rates for the threshold test come from NeurIPS 2026 (7,900 of 30,709) and COLM 2026 (29%), with EMNLP 2025 standing in for EMNLP 2026, which had not published its submission count.