AI trading agents: what actually works (and what's hype)
“AI agent that trades the market and prints money” is one of the most-searched and least-honest pitches in AI right now. The technology underneath is real and genuinely interesting — but the gap between a slick backtest and live P&L is where most of these projects quietly die. Here’s the honest version, from people who build agents for a living.
What’s genuinely new
Traditional algorithmic trading is hardcoded rules. The shift in 2026 is using LLMs as reasoning engines: an agent reads earnings transcripts, filings, news, and price data, then reasons about what it means — instead of following fixed if-then logic. Open frameworks like TradingAgents and ai-hedge-fund (agents modeled after named investors) show the pattern: multiple specialist agents (a researcher, a risk manager, a trader) debating a decision.
This is real and useful — for research and signal extraction. LLMs are good at turning messy unstructured text (a 60-page 10-K, an earnings call) into structured signal. A small team can now do data work that used to need a desk of analysts.
Why the backtest lies
In backtesting, AI trading agents can look spectacular. In live trading, the gap is brutal, and it comes from things a backtest hides:
- Transaction costs & slippage — the price you backtested isn’t the price you get.
- Data leakage — the single most common (and most flattering) bug: the model “saw” information it wouldn’t have had in real time. Almost every too-good-to-be-true backtest has it.
- Regime change — markets shift; a model tuned on the last two years can break in a month.
- Sentiment fragility — LLM sentiment reads are noisy and easy to overfit.
- Limited predictability — short-term price movement is close to random; no amount of LLM reasoning changes that.
If someone shows you a backtest and not a live track record with costs included, you’re looking at marketing, not a strategy.
What actually works
The durable applications aren’t “agent decides the trade.” They’re the boring, high-value layers:
- Research automation — summarizing filings, earnings calls, and news into structured, cited briefs a human acts on.
- Signal extraction — turning unstructured text into features for an existing quant pipeline.
- Monitoring & alerts — flagging risk, anomalies, or compliance issues across positions.
- Analyst copilots — surfacing context before a human makes the call.
The common thread: AI as decision support with a human in the loop, not an autonomous money printer.
The engineering that separates real from hype
If you’re building in this space, the work that matters is exactly the work the hype skips:
- Leakage-proof evaluation — point-in-time data, walk-forward testing, costs modeled in. This is the same discipline as a real eval harness, applied to markets.
- Grounding, not guessing — answers tied to real documents via RAG, with citations, so a human can verify.
- Guardrails — an agent that takes actions on real money needs hard limits, approvals, and audit logs.
How we approach it
At Malgary Labs we’re happy to build AI for finance — research agents, signal pipelines, analyst copilots — but we’ll tell you straight: we build decision-support systems with rigorous, leakage-proof evaluation, not autonomous trading bots that promise returns. The honest version is less exciting and far more useful.
Building something real in fintech AI? Book a free call — we’ll help you separate the part that works from the part that just demos well. See also AI for fintech: 6 use cases that ship.
Sources: LLMs for stock forecasting — a hedge-fund review (arXiv) · TradingAgents framework · Best AI trading agents 2026