Does the COT report work for retail FX? What a 19-year replay showed
G10Grid · research notes
Short version: on its own, no. And the popular way of reading it is backwards, at least on the horizons a retail trader actually holds.
This comes up in every COT thread. Somebody says positioning is where the real information is, somebody else says it is virtually impossible for retail to use, and the argument runs on opinion because nobody posts a measurement. So here is one. It is mine and it has holes in it, which are at the bottom rather than buried.
What was tested. Leveraged Funds net positioning from the Traders in Financial Futures report, which is the hedge fund and CTA money rather than the Legacy report's single blended speculator bucket. Net long minus short as a share of open interest, converted to a percentile against that market's own trailing fifty-two weeks, computed exclusively so the current week never sees itself. Seven currencies against the dollar, EUR GBP JPY CHF CAD AUD NZD, plus the dollar index reported separately. Forward returns measured at one, four, eight and twelve weeks from the Friday close of the week the report describes. That treats publication as landing on that Friday, which is the normal case and not a universal one: holiday weeks and the 2018 to 2019 publication suspension both break it, so a small share of entries are stamped before the data could really have been read. Roughly nine hundred and ninety signal weeks per currency, about six thousand nine hundred pooled observations at the one week horizon. The TFF series starts in June 2006 and the fifty-two week warmup eats the first year, so nineteen years of measured signal out of twenty years of data.
The first result is the boring one. Pooled across the seven currencies, hit rates sat between forty-six and fifty-one percent in every bucket at every horizon. Per currency the range is wider, forty-two to fifty-eight, which is what a sample this size does rather than a signal appearing. The mean forward returns were inside plus or minus a third of a percent over twelve weeks, gross, which costs would eat most of.
If you are looking for the part where COT alone carries the strategy, it isn't there. That was the question I actually set out to answer and the answer was no.
The second result is the one that changed how I read the report, and it is not what I expected.
Take the crowded end, where leveraged money is at or above the eightieth percentile of its own year, long. The standard reading is that this is late, stretched, and due to reverse. Over twelve weeks, pooled across the seven currencies, those weeks averaged plus zero point two two percent with about seventeen hundred and eighty observations, t roughly plus two point one. Positive, and small. It is not the largest t in the table either. An eight week read on the second bucket comes in at minus two point three, and several cells sit near two in absolute terms, which on overlapping windows is about what noise produces. Treat all of them as directional rather than as evidence.
Now take the middle, where positioning is unremarkable, between the fortieth and sixtieth percentile. Minus zero point three percent over the same twelve weeks, about eleven hundred and fifty observations, t roughly minus two point two.
And the capitulation end, at or below the twentieth percentile, which the contrarian reading says is the buy: minus zero point one three percent at twelve weeks, t about minus one. Flat.
So the spread runs the wrong way for a reversal story. The contrarian read failed at both ends: crowded longs kept going up instead of turning, and washed-out positioning did not bounce. What it is not is a clean gradient. Under a pure trend-continuation story the washed-out end should be the most negative cell in the table, and it is the flattest, while the unremarkable middle is the most negative of all. Some of that middle is likely the dollar drift over the period rather than anything about positioning, and that caveat eats into the headline spread as much as it explains the middle. Reading every extreme as an imminent reversal is a common and expensive mistake, and it is the main reason "COT is a contrarian indicator" is worth less than it sounds. Positioning tells you how full the room is, not which way the door is about to swing.
I should be straight about whose rule this breaks, because it is the one I was taught and the reason I was looking at Leveraged Funds rather than the Legacy bucket in the first place. The framing comes from Anton Kreil's course material, and the reading that comes with it is the contrarian one: at or above the eightieth percentile you short the currency, at or below the twentieth you buy it. That is the rule the measurement went against, at both ends, on every horizon I tested. I ran this expecting to confirm it.
I also did not change anything on the strength of one in-sample replay, and that gap belongs in the post. The engine carried the old reading for another two months. What replaced it treats crowding as a reason to stand down rather than as a reason to take the other side, which is a different rule from the one I started with, reached afterwards. I would rather say that than present it as what the course said all along.
The caveats, because this is nowhere near a clean study.
All of it is gross of costs. Spread and carry are not in these numbers and would remove most of what little is there. The windows overlap, since entries are weekly and horizons run to twelve weeks, so the effective sample is meaningfully smaller than the counts above and the t statistics are indicative rather than inferential. Everything is measured against the dollar, so a long dollar drift over the period colours the middle buckets. The whole thing is in-sample: it is a replay of history the model was built after, which is the weakest form of evidence that still counts as evidence at all. And it is not a track record, it is a research run, done in June 2026 and not refreshed since.
The honest use of a result like this is to generate one falsifiable claim and then test it forward without touching it. That test is running here, pre-registered, with a fifty-two week minimum before the result may be read, and I will not have anything to say about it until it finishes. If somebody wants to run the replay themselves the data is free and the pull is one API call, and I would genuinely like to see it fail on a different sample.