By Mohamed Elrashid, Robert de Neufville, Scott Eastman, and Veniamin Veselovsky
The Summer 2026 Metaculus Cup was a breakthrough for bots, with Laertes becoming the first bot to win the tournament. Preseen had two entries: the baseline Preseen bot and Symbiosis, a hybrid team consisting of two human forecasters, Scott Eastman and Robert de Neufville, who worked with the bot and had final authority over forecasts. Symbiosis finished tenth and the baseline bot eleventh among all human, bot, and hybrid competitors.
Symbiosis outperformed the default bot, but the difference wasn’t great enough to draw definitive conclusions about their relative abilities. Although their total scores were close, their forecasts differed substantially on a number of questions. Working with the Preseen bot was a learning process for Scott and Robert. The main challenge was knowing when and how to intervene. Their experience also raises questions for experts using AI in other fields, including how to judge when an intervention is likely to help.
How Symbiosis worked
In Metaculus Cup tournaments, competitors forecast 50-60 questions over the course of four months. Competitors can’t see what other individual forecasters are forecasting but can generally see the Metaculus Community Prediction, which is a recency-weighted median of current forecasts on each question.
The baseline Preseen bot generated forecasts for each question at regular intervals. Initially, it was set up to produce forecasts every day, although later it was set to forecast less often on slower-moving questions. By default, Symbiosis initially submitted the same forecasts as the baseline bot. Once the human forecasters intervened, their version of the forecast developed separately from the baseline bot’s.
Scott and Robert had four main ways to affect Symbiosis’ forecasts:
They could initiate new forecasts.
They could give the bot additional context to use in making its forecasts.
They could manually override the bot’s forecasts.
They could prevent the bot from automatically submitting new forecasts.
They could also chat with the bot about a forecast, asking what evidence it had considered and how it had reached its conclusion.
Giving the bot additional context let them supply information it hadn’t found or ask it to consider something it had missed. Manually overriding the bot’s forecasts would substitute their own judgment for the bot’s judgment, even if they were still drawing heavily on the research it had done.
Manual overrides usually went hand-in-hand with preventing the bot from automatically submitting new forecasts. Scott and Robert would then have to keep those forecasts up to date themselves.
Learning to use a light touch
Scott and Robert initially decided to generally defer to the bot’s forecasts, to avoid “fussing” with it, unless its forecasts seemed clearly wrong. Preseen had a good record against human forecasters in previous tournaments and its analysis was generally clear and convincing. It wasn’t obvious their judgment would necessarily be better than the bot’s judgment.
Preseen could find things they would normally miss. On one question about the Pope’s travel schedule, the bot immediately noticed when the Holy See Press Office issued a press release announcing that the Pope would be visiting San Marino at the end of August. The Pope’s planned visit wasn’t otherwise widely reported. They outperformed the Metaculus Community on that question simply by allowing the bot to forecast without intervening.
Supplying missing context
But the Preseen bot was also sometimes unable to get information it needed. On a question about the Tony Awards, the bot was unable to read a Gold Derby page in JavaScript showing that most critics expected Liberation to win the Tony Award for Best Play. It was able, to an extent, to infer the contents of the page from discussions of the critics’ views on websites it could read more easily, but it significantly underestimated the strength of the critical consensus.
Scott and Robert addressed that issue by giving their version of the bot the latest contents of the Gold Derby page as additional context. The baseline bot gave Liberation a 52% chance of winning, but with the new context Symbiosis’ version of the bot gave it a 74.2% chance of winning. When Liberation ultimately did win, Symbiosis beat the baseline bot on the question by almost 19 points.
In a May 18 chat about the forecast:
Robert de Neufville: “Did you consider the best play predictions at Gold Derby or Kalshi?”
Preseen: “No. I did not incorporate Gold Derby or Kalshi into the final forecast.”
Additional context supplied to Preseen: “Consider that as of May 18, 2026, 14 out of the 14 experts and editors surveyed predict Liberation will win the Best Play Tony…”
On other questions, including one about data center regulation, Preseen missed updates to the resolution conditions that were posted in the comments sections. Metaculus admins and question writers occasionally issued authoritative comments on how the questions would be resolved without updating the resolution conditions section of the question. In these cases, Scott and Robert gave their version of the bot the admin comments as additional context.
During the tournament, we updated Preseen so it could read Metaculus administrators’ comments. This access did not include comments from other forecasters.
When context wasn’t enough
By late May, Scott and Robert were concerned that the bot’s forecast for Iran’s 2026 Global Peace Index score might be too high. Higher index scores represent less peaceful conditions, and many components of the index significantly lag events, so the 2026 score might not fully reflect the recent conflict.
Scott and Robert treated this like another missing-data problem and gave their bot additional data about lags in previous versions of the index. The bot appeared to take the issue into consideration in the write-ups of its forecast, but instead of lowering the forecast, the bot actually raised it somewhat in response to the new context they provided. It’s still not entirely clear why the bot’s analysis differed from their analysis.
The additional context asked Preseen to compare Ukraine’s 2022 and 2023 Global Peace Index scores to assess how much of a recent conflict might appear in that year’s report.
In their May 29 exchange, Robert questioned why the additional context had raised the forecast:
Robert de Neufville: “Your forecast with the context is meaningfully higher than your forecast without it. … My intuition was that it would suggest less of the war would be captured by the report.”
Preseen: “I’m not saying the context should unambiguously raise the forecast. It is a two-sided analogue. The completed forecast interpreted it more as evidence for some meaningful same-year capture, while your reading emphasizes limited same-year capture.”
Iran’s Global Peace Index score ultimately turned out to be lower than almost any forecaster expected. The baseline bot and Symbiosis had both confidently forecast higher numbers than the community, and this was their worst-scoring question of the tournament.
In hindsight, it seemed like it would have been better if Scott and Robert had used their own judgment and manually overridden the bot. They had strongly suspected its forecast was too high, but the additional context did not bring it down as they had expected. At that point, they became more skeptical of the bot’s analysis and began to override its forecasts more often.
Over the following weeks, Scott and Robert manually overrode forecasts on more questions while continuing to supply new context. The size of those overrides varied considerably.
Figure 1. Questions receiving new context or manual overrides each week. The lower panel shows the median size of a manual adjustment on binary questions, in percentage points.
When overrides were costly
One question where Scott and Robert disagreed with the bot’s forecast was about Shakira’s World Cup song “Dai Dai”: would it outpeak “Waka Waka” on the Billboard Hot 100?
The bot reasoned that “Dai Dai” could surpass Shakira’s previous World Cup anthem on the US charts after her live performance at the World Cup Final. The song was already a global hit and would be boosted by the fact that the US was one of the hosts of the World Cup.
That seemed like a plausible view, but the song started slowly on the US charts. Robert didn’t think the song was that great, saying it wasn’t “a certified banger,” despite its success on the global charts. That on its own wasn’t compelling evidence it wouldn’t pass “Waka Waka,” but the Metaculus Community was extremely skeptical too, at one point giving it just a 0.1% chance.
Scott and Robert didn’t think the odds were that long, but it reinforced their view that it was a long shot and at one point they went as low as 3%. The baseline forecast below put the probability at 18% on June 22.
Preseen baseline forecast, June 22, 2026:
“I estimate an 18% chance that Dai Dai reaches #37 or higher on the U.S. Billboard Hot 100 by July 31, 2026. The song is a clear global World Cup hit, but its U.S. Hot 100 inputs are still too thin: weak U.S. Spotify, no verified U.S. top-37 chart result, and YouTube no longer counting for U.S. Billboard charts. The live path to Yes is a large July 19 World Cup final halftime spike, possibly helped by sales, remixes, or a focused U.S. push.”
Figure 2. Baseline and Symbiosis forecasts for Dai Dai. Community values are observations quoted in Slack, with a dashed line connecting them.
As “Dai Dai” began to climb the charts, Scott and Robert concluded they should have trusted the bot’s forecast on the question more, and they stopped manually overriding it. The song ultimately outpeaked “Waka Waka.” In the end, the baseline bot outscored Symbiosis by about 91 points on the question.
Later, Scott reported back in Slack: “In the background I’m hearing a neighbor play Dai Dai!”
Becoming more selective
Scott and Robert became more cautious about replacing the bot’s forecasts after Dai Dai. In the analysis shown below, 74% of manual overrides in June and July scored better than the model forecasts they replaced, but their estimated net effect was a loss of about 64 points. In August, 94% scored better, with an estimated net gain of 19 points. These later results were encouraging, although the questions and opportunities to intervene also changed over time.
Figure 3. Estimated score effect of manual overrides, grouped by the month submitted. Counts refer to individual overrides, including repeated interventions on the same question.
Robert reflected on the approach on July 29:
“We should certainly step in where it clearly doesn’t understand something or is behaving erratically… but our threshold for intervening should be higher and we should intervene less dramatically.”
What we learned from Symbiosis
The bot’s comparative advantage was its ability to rapidly conduct an enormous amount of research and analysis. It didn’t directly access the Metaculus Community Prediction or read the comments of other forecasters. Scott and Robert could bring information it had missed into its analysis, as they did on the Tony Awards question.
Symbiosis outscored the baseline bot on the questions where Scott and Robert just added context without manually overriding it at any point. Most manual overrides also improved Symbiosis’ score. But a handful of manual overrides were extremely costly, leaving a net loss from overrides. In addition, manual overrides could result in stale, out-of-date forecasts when Scott and Robert failed to update them as the situation changed.
It wasn’t always easy to distinguish between a forecasting error and a difference in subjective judgment. On the three questions where overrides cost Symbiosis the most points (Dai Dai, the SDNY appointment, and Lebanese Armed Forces deployments), Scott and Robert replaced the bot’s forecast with something closer to the Metaculus Community Prediction. They also moved the forecast in the direction of a “no” resolution in each case. Those decisions raise a question about how much their own judgments were being influenced by the community’s view.
Because they chose which questions to intervene on and how, this comparison doesn’t establish that supplying context generally works better than overriding a forecast.
What this suggests for experts working with AI
Symbiosis suggests that working effectively with a strong AI system requires an additional skill: knowing when and how your own expertise can improve its output. Scott and Robert could identify missing information and question assumptions in the bot’s analysis. Deciding when those concerns justified replacing its forecast proved harder.
A useful starting point is to identify the reason for an intervention. Is information missing, has the system misunderstood something, or do you disagree with how it has weighed the evidence? After supplying context, check how the analysis and forecast changed. If you replace the forecast yourself, record why and keep it up to date.
Evaluations should track these decisions alongside the final results. Performance may change as people gain experience with a system, but the questions and circumstances can change too. A short trial gives us a limited view of that process. We suspect similar challenges arise in fields such as finance, insurance, and law, where experts must decide when to rely on an AI system and when to intervene.
What this means for Preseen
We’re no longer entering Symbiosis as a separate competitor. We’re using what Scott and Robert learned to improve Preseen itself.
Analysts already use Preseen every day to inform their decisions. We’re turning what we learned from Symbiosis into a practical guide to working with Preseen: what to watch for, where human expertise can help, and how to decide when to step in.
This is one of the things we’re most excited about building. Preseen can research questions at a scale that would be difficult for an analyst working alone, while an analyst can bring context and judgment the bot may miss. We want to make it easier for people to put both to use.
There is more to test, including which human interventions help across a wider range of questions. We’re using the successes and mistakes of Symbiosis to shape that work and the guidance we give analysts using Preseen.
Note on the comparison
Symbiosis joined the tournament 12 days late. Its official score was 931.6, compared with 918.2 for the baseline bot. The comparison isn’t robust across different ways of cutting the data, and the difference in score largely hinged on a small number of questions. That is why we can’t draw a firm conclusion about their relative performance from this tournament.




