If there is no alpha, how do I measure my performance?
What is skill, respected sir? Is it, just make money more than others if the amount of risk and capital are assumed to be the same?
Also, I know its a bit overeaching, but could you tell a bit about how to find a statistical edge in our trading, I've been working on it for a while, I code, read papers and do research, but I have yet to find a stable edge that I could execute on.
Yes. Just make more money than others while not taking insane risks. It’s very difficult to say anything that is both true and more useful than that. For finding edges, it takes a very long time. You have two costs: opportunity cost of taking the time to find the edge and then the lack of experience cost where even if you find an edge, it may not exist, or may somehow lose you money no matter what you do for reasons you don’t understand. That is why I always suggest to anyone starting off: put the base rates in your favor. To put base rates in your favor, build trend following models. Once you know what you are doing, you can specialize from there.
Update: I took your advice to heart and built several trend-following models. I found a lot of success in both the backtests and out-of-sample testing, but now I’m facing a different issue.
Sorry if I’m being annoying, but I haven’t been able to find a clear answer because the industry seems divided on this. Some people say that if the in-sample p-value is low and the t-stat is strong, you should simply run the strategy. Others say you must account for the number of searches and parameter combinations you tested.
The other day, I read your article, “The Emperor Has No Alpha.” It was a great read, and I want to thank you on behalf of your readers for sharing your knowledge and experience with us.
Coming back to my problem, I usually start with a hypothesis and turn it into a set of rules. The initial rules are often only okay, so I run a parameter search and adjust them using the in-sample data. Eventually, I find a version with a good p-value and t-stat.
I then perform several additional checks, but to keep it brief, I test the strategy out of sample. The profitability often persists, and the strategy remains profitable, but the p-value and t-stat usually become weaker.
Because I searched through many parameter combinations, the process is essentially a form of data mining. The industry says I should therefore penalize the backtest using methods such as the Deflated Sharpe Ratio, Holm-adjusted p-values, and similar corrections. However, whenever I apply these penalties, I am almost guaranteed not to find a statistically significant edge.
I’ve watched many of your interviews, and I’ve heard you explain that robustness should also be evaluated practically. For example, nearby parameter values should continue to work, and the signal should weaken gradually as the parameters move further away from the optimum. We should also add noise to the system and observe whether its performance degrades gracefully. Ideally, we want to see a broad, stable region of good performance rather than one isolated spike.
My question, then, is how to reconcile this with the problem of multiple testing. In your article, you argue that data mining works, that statistical strength matters more than theoretical elegance, and that we should test repeatedly until we have genuine confidence in what we are doing. But the more I test different parameters, filters, markets, timeframes, and variations the larger my effective search space becomes.
The conventional statistical view is that I must account for every search I perform. Once I apply corrections such as the Deflated Sharpe Ratio, Holm-adjusted p-values, or similar multiple-testing penalties, however, almost nothing remains statistically significant. So when you refer to the strength of the in-sample evidence, should I look at the raw t-statistic of the final strategy, a statistic adjusted for the entire research process, or place more weight on untouched out-of-sample performance and practical robustness?
More broadly, if data mining is valid, where should its boundaries be drawn? Is there a practical limit to how many ideas or parameter combinations one can test on the same development data before that data loses its evidentiary value? Can repeated mining remain valid as long as a genuinely untouched test set is preserved, or does the size of the search still matter even when the final strategy continues to perform out of sample?
I am also unsure how broad the search space should be. Should we search any data that may contain useful information, or should the variables and transformations still be constrained by some plausible connection to the market? In other words, how do you distinguish productive data mining from searching so widely that finding an apparently successful strategy becomes almost inevitable?
Sorry for bombarding you with so many questions. I also understand that, given your role, some aspects of your research process may be proprietary or simply not something you are able to discuss publicly. I completely respect that, so please feel free to answer only what you are comfortable sharing. Even a general framework for how you think about these issues would be incredibly valuable. Thank you again for sharing so much of your knowledge and experience with all of us.
The short answer is that p values are not useful. The look elsewhere effect is also not useful. What is useful is a very flat parameter space and robustness to noise. That is usually enough to produce a useful signal. Additionally, the only thing that matters is tested profitability after all estimated costs. RMSE and other such estimators are useless.
If there is no alpha, how do I measure my performance?
What is skill, respected sir? Is it, just make money more than others if the amount of risk and capital are assumed to be the same?
Also, I know its a bit overeaching, but could you tell a bit about how to find a statistical edge in our trading, I've been working on it for a while, I code, read papers and do research, but I have yet to find a stable edge that I could execute on.
Thank you
Yes. Just make more money than others while not taking insane risks. It’s very difficult to say anything that is both true and more useful than that. For finding edges, it takes a very long time. You have two costs: opportunity cost of taking the time to find the edge and then the lack of experience cost where even if you find an edge, it may not exist, or may somehow lose you money no matter what you do for reasons you don’t understand. That is why I always suggest to anyone starting off: put the base rates in your favor. To put base rates in your favor, build trend following models. Once you know what you are doing, you can specialize from there.
Update: I took your advice to heart and built several trend-following models. I found a lot of success in both the backtests and out-of-sample testing, but now I’m facing a different issue.
Sorry if I’m being annoying, but I haven’t been able to find a clear answer because the industry seems divided on this. Some people say that if the in-sample p-value is low and the t-stat is strong, you should simply run the strategy. Others say you must account for the number of searches and parameter combinations you tested.
The other day, I read your article, “The Emperor Has No Alpha.” It was a great read, and I want to thank you on behalf of your readers for sharing your knowledge and experience with us.
Coming back to my problem, I usually start with a hypothesis and turn it into a set of rules. The initial rules are often only okay, so I run a parameter search and adjust them using the in-sample data. Eventually, I find a version with a good p-value and t-stat.
I then perform several additional checks, but to keep it brief, I test the strategy out of sample. The profitability often persists, and the strategy remains profitable, but the p-value and t-stat usually become weaker.
Because I searched through many parameter combinations, the process is essentially a form of data mining. The industry says I should therefore penalize the backtest using methods such as the Deflated Sharpe Ratio, Holm-adjusted p-values, and similar corrections. However, whenever I apply these penalties, I am almost guaranteed not to find a statistically significant edge.
I’ve watched many of your interviews, and I’ve heard you explain that robustness should also be evaluated practically. For example, nearby parameter values should continue to work, and the signal should weaken gradually as the parameters move further away from the optimum. We should also add noise to the system and observe whether its performance degrades gracefully. Ideally, we want to see a broad, stable region of good performance rather than one isolated spike.
My question, then, is how to reconcile this with the problem of multiple testing. In your article, you argue that data mining works, that statistical strength matters more than theoretical elegance, and that we should test repeatedly until we have genuine confidence in what we are doing. But the more I test different parameters, filters, markets, timeframes, and variations the larger my effective search space becomes.
The conventional statistical view is that I must account for every search I perform. Once I apply corrections such as the Deflated Sharpe Ratio, Holm-adjusted p-values, or similar multiple-testing penalties, however, almost nothing remains statistically significant. So when you refer to the strength of the in-sample evidence, should I look at the raw t-statistic of the final strategy, a statistic adjusted for the entire research process, or place more weight on untouched out-of-sample performance and practical robustness?
More broadly, if data mining is valid, where should its boundaries be drawn? Is there a practical limit to how many ideas or parameter combinations one can test on the same development data before that data loses its evidentiary value? Can repeated mining remain valid as long as a genuinely untouched test set is preserved, or does the size of the search still matter even when the final strategy continues to perform out of sample?
I am also unsure how broad the search space should be. Should we search any data that may contain useful information, or should the variables and transformations still be constrained by some plausible connection to the market? In other words, how do you distinguish productive data mining from searching so widely that finding an apparently successful strategy becomes almost inevitable?
Sorry for bombarding you with so many questions. I also understand that, given your role, some aspects of your research process may be proprietary or simply not something you are able to discuss publicly. I completely respect that, so please feel free to answer only what you are comfortable sharing. Even a general framework for how you think about these issues would be incredibly valuable. Thank you again for sharing so much of your knowledge and experience with all of us.
The short answer is that p values are not useful. The look elsewhere effect is also not useful. What is useful is a very flat parameter space and robustness to noise. That is usually enough to produce a useful signal. Additionally, the only thing that matters is tested profitability after all estimated costs. RMSE and other such estimators are useless.
Thanks a lot, I’ll move in that direction
Got it, I understand what you mean, thanks.