As a follow-up to my original post, I’ve decided to dig into the data and evaluate the often debated F-Score. In my experience, the general consensus is that the F-Score does not work with Japanese Net-Nets. This conclusion may very well be misguided, but remarks from other bloggers and practitioners, as well as my earlier crude calculations, the general consensus seemed to hold true.
Part of the reason for this conclusion might be that price momentum doesn’t appear to work well in Japan, either. So it stands to reason that if price momentum doesn’t work, why should fundamental momentum work? Nevertheless, in a 2011 article, Cliff Asness argues that Japan is the exception to the success of momentum strategies in other markets and asset classes, and that the failure to make money in Japan for nearly 30 years is “within the range of statistical noise”. That all sounds dandy, but how does a simple practitioner work with “statistical noise”?
Another reason for the potentially misguided conclusion is that the F-Score was not designed with Japanese Net-Nets in mind. In 2000, Joseph Piotroski published the now famous paper on using fundamental momentum to separate the winners and losers from a large basket of cheap stocks curried by their price-to-book ratio. In the paper, he documents that, although a low P/B strategy works in aggregate, less than 44% of all low P/B stocks earn positive market-adjusted returns in the two years after portfolio formation. He concludes that the majority of low P/B stocks don’t work because they are “financially distressed”. The F-Score, in a sense, is designed to invert the process and eliminate the distressed firms. This logic does not hold well in the land of Japanese Net-Nets where I certainly would not describe this collection of stocks as distressed. Perhaps, then, the act of separating Japanese Net-Nets vis-a-vis distressed and non-distressed is moot.
Logically speaking, fundamental momentum should work wonders with Net-Nets. They are in the sweet-spot of Piotroki’s definition of neglected and under-followed stocks. And given their relative price to liquidation value, improving business operations should have an outsized impact on the stock price.
There is not much research out there on Japanese Net-Nets and the F-Score. Noma (2010) demonstrates that the F-Score works in Japan (+17.6% CAGR). Whenever possible, I like to see what Dan Rasmussen and Verdad Advisers has to say about things. In a blog post, they mention that they use three of Piotroski’s tests of fundamental health, improvements in asset turnover, reductions in long-term debt, and positive cash from operations, and incorporate those data points into their larger model (with a focus on highly leveraged firms). They conclude this exercise, nicknamed The Piotroski Synthesis, proves rewarding.
In all, there is plenty of evidence that the efficacy of implementing the F-Score is sound. But does it work with Japanese Net-Nets?
Enough of my blabbering. Let’s see the data.
Below is a count sorted by F-Score. Contrary to Piotroski, I did not drop holdings that produced a NULL in one particular (or multiple) data field(s). (The missing data was more pronounced prior to 2014). The sample size was large enough in the original study (14,043 stocks) to handle dropping stocks that did not produce all of the sufficient financial statement data. I, though, do not have that luxury. You might say these results are left-skewed a bit, (i.e., there could be more 8s and 9s and less 0s and 1s, etc.) if I had all the data. But the results are close enough, in my opinion, and that’s good enough in this case. As is the case in the original paper, most of the observations are clustered around F-Scores between 3 and 7. Piotroski describes this as “conflicting performance signals”.

Piotroski’s paper aggregates low F-Scores (0 or 1) and high F-Scores (8 or 9) and sets up a long/short strategy. The F-Scores of 2 to 7 are more or less ignored. Given the distribution of the low F-Scores, Piotroski states one could use low F-Scores of less than or equal to 2 and that results are “qualitatively similar to those presented throughout the paper”. It also appears to be standard practice on most financial websites to use a low F-Score of 0 to 2. Due to the lack of 0s and 1s, I chose to use a low F-Score of 0 to 2.
Below is a count of grouped F-Scores by year. You should note there are years in which we have inadequate representation in the low and high F-Score buckets. For example, in 2018, only F-Scores of 3 to 7 were represented.

The mean, or equal-weight, return by year promptly reveals some problems with the low F-Score, as it catches some high-fliers. The high F-Score appears do well year-by-year and outperforms 6 years out of the 11 years it is represented.

The results of the horse race are below. Given that the F-Scores of 0 to 2 produced some big outliers and given that the low F-Score portfolio will hold concentrated positions, the low F-Score outperforms (23.91% CAGR). However, the high F-Score also shows well (22.88% CAGR), while the “conflicted” bucket of 3 to 7 underperforms. The high F-Score is not represented in 3 of the 14 years and the low F-Score misses out on 1 year. Would the high F-Score outperform if it had been included in every year? Feel free to draw your own conclusions.

The results are a bit puzzling to decipher and I think help explain why it’s been difficult for me to figure out if the F-Score works with Japanese Net-Nets. Because there are some big winners when the F-Score is horrible, it can really muddy the water. The problem here is that the results do not represent the “monotonic positive relationship between F_SCORE and subsequent returns” that the original paper clearly demonstrated. Perhaps it simply can be explained by the small sample size. But whatever the reason, it’s difficult to draw strong conclusions.
Perhaps one could buy low F-Scores to capture the big winners while enduring the increased volatility and buy high F-Scores due to sound logic and better performance. The problem is is that approach does not follow Piotroski’s method, as he wrote, “…any strategy that can eliminate the left tail of the return distribution (i.e., the negative return observations) will greatly improve the portfolio’s mean return performance.” In this study, the low F-Score did a poor job of eliminating the left tail, while the high F-Score outperformed the baseline portfolio with fewer drawdowns. Perhaps this study produces more questions than answers. Does Piotroski’s F-Score work with Japanese Net-Nets?
Yes, and no.
For those wanting to dig into the data, below is mean performance by F-Score.
Thanks for reading!
| MEAN RETURNS | F-SCORE | |||||||||
| Start Date | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
| 2010-04-20 | -0.52593 | -0.48750 | 0.07379 | 0.11868 | -0.01720 | 0.00390 | -0.02861 | 0.06538 | 0.16124 | |
| 2011-04-20 | -0.06184 | -0.00725 | 0.14340 | 0.26344 | 0.02129 | 0.38465 | 0.62581 | 0.51952 | ||
| 2012-04-20 | 0.04575 | 0.48738 | 0.65031 | 0.55046 | 0.49267 | 0.24087 | 0.66334 | 0.53570 | ||
| 2013-04-20 | 0.35244 | 0.18196 | 0.45328 | 0.34265 | 0.22234 | 0.07078 | 0.18313 | 0.42750 | ||
| 2014-04-20 | 4.16949 | 0.84978 | 0.49925 | 0.43231 | 0.71563 | 0.32781 | 0.33619 | 0.37751 | 0.61155 | |
| 2015-04-20 | -0.00521 | 0.19820 | -0.06710 | -0.05197 | -0.11091 | -0.17896 | 0.05782 | |||
| 2016-04-20 | 1.64883 | 0.22834 | 0.21767 | 0.48591 | 0.28574 | 0.20402 | ||||
| 2017-04-20 | 0.54642 | 0.48733 | 0.22084 | 0.48967 | 0.35226 | 0.26939 | 0.51603 | |||
| 2018-04-20 | -0.22876 | -0.32390 | -0.24699 | -0.04730 | -0.03125 | |||||
| 2019-04-20 | -0.17358 | 0.11965 | -0.07510 | -0.20072 | -0.07456 | -0.10715 | -0.18357 | 0.13331 | ||
| 2020-04-20 | 0.31120 | 0.48103 | 0.51850 | 0.44375 | 0.54349 | 0.42297 | 0.24886 | 0.34354 | ||
| 2021-04-20 | -0.03721 | 0.06539 | 0.05889 | 0.33051 | -0.03274 | -0.01672 | ||||
| 2022-04-20 | 0.09894 | 0.07250 | 0.12130 | 0.11407 | 0.19426 | 0.27352 | 0.38831 | 0.25150 | ||
| 2023-04-20 | -0.02183 | 0.14846 | 0.35205 | 0.48885 | 0.39511 | 0.40234 | 0.61073 | 0.30536 |