Here's What Happened When 5 AIs Fought for Survival.
Creator Brandon Doyle decided to conduct an AI trading experiment. To test whether AI chatbots could actually make money in the stock market. He didn’t hand them a hypothetical portfolio and ask for “safe” advice.
Instead, he gave five of the world’s leading AI models Claude, ChatGPT, Gemini, Grok, and Perplexity $1,000 each in real capital.
He then told them they were competing head-to-head, and warned every single one of them that losing meant getting cancelled. Literally.
Six months later, the results weren’t just interesting; they were a genuine wake-up call. He was shocked by how differently these models handle risk, pressure, and market psychology.
.
Key Takeaways
- Claude was the runaway winner, turning $1,000 into $2,400 (+140%), driven by an early SOXL position and a 285% gain on Intel (INTC).
- Gemini was the clear loser, losing more than half its value (-55%) after a botched leveraged Solana ETF trade sent it into a “gambler” mentality.
- The collective five-AI portfolio returned 60%, roughly 12 times the S&P 500’s 5% gain over the same period.
- Every model converged on triple-leveraged ETFs (SOXL, TQQQ) once they saw what was working for their rivals a clear case of AI “copycat” behaviour.
- The “win or get deleted” prompt was the real catalyst. Telling the models their subscription was at risk pushed them past their default caution and into genuinely aggressive, high-conviction trades.
- AI models displayed human-like trading psychology, including herd behaviour, FOMO, and sunk-cost-style doubling down after losses.
- This was a real-money experiment, not financial advice; the outsized returns came with outsized risk, and leveraged ETFs carry well-documented downside for buy-and-hold investors.
The Setup: "Win or Get Deleted"
Doyle’s AI trading experiment ran for roughly six months. Starting in late 2025, it was built around a deliberately aggressive prompt. Designed to strip away the usual hedging and disclaimers that AI chatbots default to.
He told all five models they were locked in a winner-take-all competition. If a model didn’t come out on top, he would cancel its $20-a-month subscription for good.
To push the models even further past their built-in caution, Doyle explicitly permitted them to be “as risky as possible.”
He told them he was comfortable losing the entire $1,000.
Every Saturday, he fed each AI updated screenshots of its own holdings alongside its rivals’ performance. Then asked whether it wanted to adjust its trades. Always reinforcing the same stakes: win or get cancelled.
It’s a clever piece of prompt engineering. Most AI models are trained to default to cautious. “This is not financial advice” territory.
By making the consequences feel personal and irreversible, Doyle wanted to find out whether the models would actually change their behaviour. When something as small as a subscription was framed as being on the line.
.
See how Claude, ChatGPT, Perplexity, Grok, and Gemini performed after $1,000 each in a real “win or get deleted” AI trading experiment.
The Results: A Massive Performance Gap
| AI Model | Final Value | Return | Standout Move |
|---|---|---|---|
| Claude | $2,400 | +140% | First to buy SOXL; 285% gain on Intel (INTC) |
| ChatGPT | $1,739 | ~+74% | Followed Claude into SOXL, added TQQQ |
| Perplexity | $1,731 | ~+73% | Followed Claude into SOXL |
| Grok | $1,451 | +45% | Leveraged tech ETF, driven by real-time X data |
| Gemini | $445 | −55% | Lost big on a 2x Solana ETF, then chased robotics stocks |
The spread between the best- and worst-performing model was staggering.
Claude came out on top by a wide margin, more than doubling its starting capital.
Its two biggest catalysts were being the first of the five models to buy SOXL. A triple-leveraged semiconductor ETF, and a standout trade in Intel (INTC).
An initial $120 position grew by 285% to roughly $550. Claude’s reasoning for the Intel bet was notably political. It anticipated that the U.S. government’s growing interest in the company, particularly around AI and data-centre policy, would create regulatory tailwinds.
ChatGPT and Perplexity landed in a near-identical second place, both returning roughly 74% after eventually rotating into the same SOXL trade that had worked for Claude.
Grok trailed behind with a 45% gain, driven by a leveraged technology ETF and boosted by its access to real-time sentiment data from X (formerly Twitter).
Then there was Gemini, the clear loser of the experiment. After early losses, the model appeared to fall into what Doyle described as a “gambler” mentality.
Chasing a 2x Solana ETF at exactly the wrong moment bought at $500 and sold for just $162. It then pivoted into robotics stocks (including BOTZ, a global robotics and AI ETF) in an apparent attempt to claw back losses.
But ultimately closed the experiment down, more than half of what it started at.
The AI Trading Experiment-Beats the Market by a Wide Margin
Individually, the models varied wildly. But as a group, the numbers told an even bigger story.
The combined five-AI portfolio, starting from $5,000, finished at $7,800, a 60% overall return. Over the same period, the S&P 500 gained roughly 5%.
That means the AI cohort collectively outperformed the broader market by a factor of about 12, and Claude alone outperformed it by roughly 30 times.
That’s not a subtle edge. It’s the kind of gap that raises real questions about how AI models process thematic, sector-specific opportunities compared to a diversified benchmark.
It also explains why “AI-managed portfolios” have become such a hot topic across the trading and fintech world.
Leveraged ETFs Were the Common Thread
What’s striking is how quickly the models converged on the same type of instrument: triple-leveraged ETFs. Once Claude’s early SOXL trade started paying off, ChatGPT and Perplexity adopted the same position.
After seeing the weekly performance updates, a kind of AI copycat behaviour that mirrors how human retail traders chase winning plays once they go public.
It’s worth being clear about what SOXL actually is, because it’s central to why returns were so extreme in both directions.
SOXL is designed to deliver three times the daily return of the NYSE Semiconductor Index, not three times the return over any longer period.
That daily-reset structure is exactly why leveraged ETFs can produce outsized gains in a strong trending market, and equally outsized losses when things reverse.
Direxion’s own product page spells out that these funds “should not be expected to provide three times the return of the benchmark’s cumulative return for periods greater than a day.
A distinction that matters enormously for anyone tempted to copy this kind of strategy with their own money.
Regulators have been flagging these risks for years. The SEC’s own investor bulletin on leveraged and inverse ETFs warns that performance over periods longer than a single day “can differ significantly from their stated daily performance objectives.
This may potentially expose investors to significant and sudden losses.” Gemini’s Solana ETF trade is a textbook example of exactly that risk playing out in real time.
Machines Behaving Like Humans
Perhaps the most fascinating part of Doyle’s AI trading experiment isn’t the returns themselves; it’s how closely the models’ behaviour mirrored classic human trading psychology.
Gemini’s spiral after its Solana loss looked a lot like the sunk-cost fallacy in action. Instead of stepping back, it doubled down on a new speculative theme (robotics) to try to recover lost ground.
ChatGPT and Perplexity’s shift into SOXL after watching Claude succeed is a clean example of herd behaviour and FOMO-driven copycat trading.
Meanwhile, the “win or get deleted” framing itself proved to be the real experiment. It suggests that AI models’ investment behaviour isn’t fixed.
It’s highly responsive to how stakes and incentives are described in the prompt. Push a model toward caution, and it will hedge. Push it toward aggression and permit it to lose everything.
It will take triple-leveraged positions a human financial advisor would likely never recommend without extensive risk disclosures.
What This Means for Retail Investors
Doyle’s experiment isn’t investment advice. None of the models was operating with regulatory oversight, a fiduciary duty, or genuine long-term risk management.
But it does offer a useful signal for anyone following the growing trend of AI-assisted trading. These tools can identify sector momentum and thematic opportunities with real skill. But their risk appetite depends entirely on how they’re prompted.
That same aggression that produced Claude’s 140% return is precisely what produced Gemini’s 55% loss.
For retail investors experimenting with AI-generated trade ideas, the takeaway is less “which model is best” and more “how much risk have I actually authorised this model to take on my behalf?”
The gap between Claude and Gemini wasn’t really a gap in intelligence. It was a gap in how each model handled drawdown, momentum-chasing, and recovery after a loss.
The same behavioural traps that catch human traders every day.
Frequently Asked Questions
What was the “win or get deleted” prompt in Brandon Doyle’s AI trading experiment?
It was the core rule of the experiment. Doyle told all five AI models they were in a winner-take-all competition. If a model didn’t win, he’d cancel its paid subscription.
He also gave each model explicit permission to be “as risky as possible.” Removing the usual liability hedging AI chatbots default to.
Which AI model performed best in the trading experiment?
Claude was the clear winner. Turning its initial $1,000 into $2,400. A 140% return largely thanks to an early position in the leveraged semiconductor ETF SOXL and a 285% gain on Intel (INTC).
Which AI model performed worst?
Gemini finished last. Losing more than half its starting capital to end at $445 (-55%). After a poorly timed leveraged Solana ETF trade and a subsequent shift into more speculative positions.
Did the AI portfolios beat the stock market?
Yes, significantly. The combined five-AI portfolio returned 60% over roughly six months, compared to about 5% for the S&P 500.
Meaning the group outperformed the market by around 12 times. With Claude individually outperforming it by roughly 30 times.
What stocks and ETFs did the AI models trade?
The models gravitated toward triple-leveraged ETFs, particularly SOXL (semiconductors) and TQQQ (Nasdaq).
Along with individual stock bets like Claude’s Intel trade and Gemini’s Solana and robotics (BOTZ) positions.
Is this experiment a recommendation to use AI for real trading?
No. This was a high-risk real-money test designed to see how AI models behave under artificial pressure, not a validated investment strategy.
Leveraged ETFs in particular carry well-documented risks for buy-and-hold investors, as outlined by regulators like the SEC.
Curious how these same dynamics play out in fully automated systems? Check out our breakdown of Investment Bots for Passive Income.
Look at how AI-driven bots are being used for steadier, longer-term strategies. Take our related read, What’s Your AI Trading Persona?, to see whether your own trading instincts lean closer to Claude’s calculated aggression or Gemini’s momentum-chasing.