Summary
An insider buying stock in their own company with their own money is the cleanest signal in public disclosure. There is no ambiguity about what it means to buy. We measured what actually followed 47,458 of those purchases, at one month, three months, six months, one year and two years, and asked whether the size of the cheque changed the answer.
The headline number depends entirely on what you compare against, and most published versions of this study never say which comparison they used. Measured against the S&P 500, stocks that insiders bought lagged by 12.95% over the following year and only 36.78% of them beat the index. Measured against the median listed stock over the identical window, the same purchases led by 2.56%, and 53.26% beat it. Both numbers are correct. The gap between them is not insider behaviour, it is the fact that insiders buy small companies and the S&P 500 is not made of small companies.
The 2.56% lead does not survive a falsification test. If we take exactly the same purchases, shift each measurement window forward by 180 trading days, and measure again, we get 2.78%. Comparing the two measurements event by event, the difference is 1.40 percentage points with a t-statistic of 1.58. At one month, three months and six months the difference is indistinguishable from zero. In other words, these stocks outperformed the median listed stock during windows that had nothing to do with the purchase. The signal is in which companies insiders buy, not in when they buy them.
A long silence does not make the purchase more predictive. Among 858 events where nobody had bought for more than two years, the median stock returned only 0.74% over the following year. It beat a weak broad-stock benchmark, but lagged both the typical S&P 500 constituent and, by much more, the mega-cap-led index. Because the same relative weakness appears in shifted windows, the drought describes the companies in the sample; it does not support buying or shorting the purchase date.
The dollar amount does survive that test, and it points the wrong way. Sorting purchases into ten groups by dollar value, the largest decile, a median cheque of 3.53 million dollars, underperformed the smallest decile by 3.28 percentage points over three months, 7.77 over six months and 17.10 over one year. Run the identical comparison on the shifted windows and the spread is -0.60, 0.35 and -1.97 points, none of them statistically distinguishable from zero. The size effect is real and the level effect is not.
Two years out, the size effect is at its strongest and its cleanest. On the panel of purchases old enough to measure that far, the largest decile trails the smallest by 15.26 points over two years with a t-statistic of -5.40, the most significant result anywhere in this study, against a placebo of -1.66 points and a t-statistic of -0.27. The effect does not mean-revert by two years. It plateaus.
It is a top-decile effect rather than a smooth relationship. Nine of the ten deciles show a positive one-year abnormal return, ranging from 0.72% to 4.89%. Only the largest is negative, at -1.80%. Part of that is composition, because 41.00% of the largest decile is purchases by 10% shareholders buying with no officer or director involved, against 6.74% in the smallest decile. But it is not only composition. Remove every 10%-owner-only purchase and re-sort the remainder, and the largest decile still trails the smallest by 6.96 points at six months and 13.36 points at one year, with the placebo still flat.
Finally, we tested the one cut the academic literature says should work, and it did not. Cohen, Malloy and Pomorski (2012) find insider predictive power sits entirely in opportunistic trades rather than routine ones. Splitting our sample the same way, the level effect is larger in routine trades, not smaller, and neither class shows a timing premium that survives the placebo.
Abstract
We measure buy-and-hold returns following 47,458 insider open-market purchase events, aggregated from 112,647 Form 4 transaction rows filed by 16,111 reporting owners on 3,859 US-listed common stocks between January 2020 and August 2025, and test whether the dollar value of the purchase carries information beyond the fact of the purchase itself.
Returns are measured at 21, 63, 126, 252 and 504 trading days, which we report as one month, three months, six months, one year and two years. Each event is compared against three benchmarks over the identical calendar window: the S&P 500, the median listed stock in the covered universe, and the median stock in the same entry-price quintile. The median event return, rather than the mean, is the primary statistic throughout, because the one-year cross-section has a skew of 15.0 and 90 of 47,458 observations exceed +1000%.
Against the median listed stock, purchases show a median abnormal return of 0.12%, 0.80%, 1.07% and 2.56% across the first four horizons, with date-clustered Newey-West t-statistics of 1.47, 2.47, 2.74 and 5.84. Price-quintile matching gives 0.01%, 0.64%, 0.96% and 2.68%. A placebo that shifts each event's measurement window forward by 180 trading days and compares the two measurements on the same events produces paired differences of 0.06, 0.64, 0.19 and 1.40 percentage points, with t-statistics of 0.13, 1.29, 0.14 and 1.58. We therefore find no evidence of a timing premium.
The 858 purchases that end a company-wide buying drought of more than two years sharpen the benchmark problem. Their median one-year return is only 0.74%. That clears a weak broad-stock median by 2.15 points, but trails the typical S&P 500 constituent by 6.55 points and the capitalization-weighted index by 16.47 points. The constituent comparison shows broad large-cap underperformance; the still-larger index gap reflects mega-cap concentration. Because both deficits recur in shifted windows, the drought supports neither a long nor a short trade.
The dollar-size relationship behaves differently. The spread between the largest and smallest dollar decile, computed on per-date medians, is -0.91, -3.28, -7.77 and -17.10 percentage points, with t-statistics of -1.33, -2.82, -3.93 and -3.20. Under the placebo the same spread is -0.71, -0.60, 0.35 and -1.97 points, with no t-statistic exceeding 1.56 in absolute value. The result holds under a 252-day shift, where the one-year windows are fully disjoint, and it survives the removal of all 10%-owner-only purchases. On a separate panel restricted to purchases old enough to measure two years, the spread reaches -15.26 points with a t-statistic of -5.40 against a placebo t-statistic of -0.27.
Splitting the sample into routine and opportunistic insiders after Cohen, Malloy and Pomorski (2012) does not reproduce their result on this data. The level effect is larger among routine traders, at 9.55% against 4.35% at one year under our stricter classification, and the paired placebo difference is insignificant for both classes, at t-statistics of 0.61 and 0.25. We read the whole set of results as evidence that insider purchases identify a type of company rather than a moment, and that large insider purchases are concentrated in situations that subsequently underperform.
Background and research question
The academic case for insider trades carrying information is old and reasonably settled in direction. Lakonishok and Lee (2001) found insider purchases predict returns, with the effect concentrated in small firms. Jeng, Metrick and Zeckhauser (2003) put a performance-evaluation frame around it and estimated an abnormal return near six percent a year on purchases. Cohen, Malloy and Pomorski (2012) sharpened it considerably by separating routine from opportunistic trades and showing that the predictive content sits entirely in the latter.
What is much less settled, and much less often stated plainly, is how big the effect is once you subtract the right thing. Insider purchases are overwhelmingly a small-company phenomenon. Any study that benchmarks them against a large-cap index will report an effect that is mostly a size premium with the sign flipped, and any study that benchmarks them against a broad equal-weighted universe will report an effect that is mostly the same size premium with the sign the right way round. The choice of benchmark can move the headline by more than fifteen percentage points a year, which is larger than the effect anyone is trying to measure.
The second question is more specific and, as far as we can tell, less studied. Insiders buy in wildly different sizes. The purchases in our sample run from a few hundred dollars to hundreds of millions. If the purchase is a signal, a natural reading is that a bigger cheque is a stronger signal, because the insider is putting more at risk. We wanted to test that directly rather than assume it.
Data and sample construction
Form 4 filings come from the SEC and are parsed into one row per reported transaction, carrying the issuer, the reporting owner, the transaction date, the date the filing was accepted, the transaction code, whether shares were acquired or disposed, the share count, the price per share, and the owner's stated relationship to the issuer.
What counts as an insider purchase
We keep only non-derivative transactions with a transaction code of Purchase and an acquired-or-disposed flag of Acquired, which is the open-market buy. Option exercises, grants, awards, gifts, and every derivative-security line are excluded. This is a deliberate narrowing: the transactions that remain are the ones where the insider paid cash at a price they did not choose.
Amendments are handled by the filing chain rather than by inference. A filing that has been superseded is dropped, and an amendment is anchored at the original filing's date, so a correction filed two years later does not create a fresh event in the wrong year.
From transactions to events
The unit of analysis is not the transaction. It is the pair of one stock and one disclosure date. If a chief executive buys on three consecutive days and reports all three lines on one Form 4, that is one event, not three. If four directors each file separately on the same day about the same company, that is also one event. The dollar value of the event is the sum of all purchase lines it contains.
This matters for the dollar-size question specifically. Treating each line separately would let a single large purchase split across a day's trading appear as several medium ones, which would flatten exactly the relationship we are trying to measure.
Entry is the first market close strictly after the disclosure date. A Form 4 is generally accepted after the close, so entering at the disclosure date's close would credit the strategy with a price nobody could have paid.
Horizons on the market calendar
Each horizon is counted on the market's own trading calendar, taken from the index series, rather than on the individual stock's sequence of rows. A stock that stops printing prices returns a null at the horizon instead of a stale close from whenever it last traded. That converts a silent bias into a measurable one.
The same calendar governs the benchmark. That sounds obvious and is easy to get wrong: the set of dates on which at least one stock printed a price is not the set of dates the market was open. In our data one stock printed a bar on 3 July 2026, a session the market was closed, and indexing benchmark windows on the union of print dates rather than on the market calendar silently shifted 131 one-year benchmark windows by a session.
Attrition
| Stage | Events |
|---|---|
| Purchase transaction rows, after amendment handling | 131,691 |
| Distinct stock and disclosure-date events | 56,379 |
| Less: no price coverage at entry | 50 |
| Less: one-year horizon not yet matured at the data cut | 8,545 |
| Less: window contains a split whose recorded ratio contradicts the price | 94 |
| Less: window contains an unrecorded corporate action | 231 |
| Less: no benchmark for that window | 1 |
| Usable at every horizon, the nested panel | 47,458 |
Attrition among matured events is 0.73% at one year, and it does not vary much by dollar band, ranging from 0.50% in the 250,000 to 1 million dollar band to 1.00% below 10,000 dollars.
All results in the first part of this note use the nested panel, meaning the same 47,458 events at every horizon. Reporting one-month results on a larger sample than one-year results would make the horizon comparison partly a comparison of different companies. The two-year section builds its own nested panel for the same reason, and it is a different and smaller one.
Empirical design
Three benchmarks
Every event is compared against three things measured over its own identical calendar window:
The S&P 500, through the index tracking series, buy and hold.
The median listed stock, meaning the median buy-and-hold return across every stock in the covered universe over the same window. We use the median rather than a mean or an index, and the reason is worth showing. The mean 21-day buy-and-hold return across this universe is +136.58%, because 1.05% of those returns exceed +100% and the largest is +137,999,904%, which are unrecorded reverse splits rather than returns. Chaining a daily-rebalanced equal-weighted index instead is worse, not better: it gives a mean 21-day return of +337.39% and an annualised drift of 9,331%, because daily rebalancing across thousands of illiquid names also harvests the bid-ask bounce documented by Blume and Stambaugh (1983). A buy-and-hold median is immune to both.
The median stock in the same entry-price quintile, which is a crude size and quality control. We do not use market capitalisation because the shares-outstanding field in our own store is known to be unreliable, and using a field we know to be wrong is worse than using a coarse proxy we know to be coarse. Entry price is coarse but it is correct.
Statistics
The median is the primary statistic. The one-year cross-section of raw returns has a skew of 15.0, and the mean is set by the tail: 90 events out of 47,458 returned more than +1000%. Means are reported alongside, and where they diverge from the median we say so.
Inference is date-clustered. Purchases cluster heavily on filing dates, so treating 47,458 events as independent would overstate precision by a wide margin. We collapse to one observation per disclosure date, take the median across events on that date, and compute a Newey-West t-statistic on the resulting series with a lag length of the fourth root of the number of dates.
The falsification test
The core problem with any event study on a non-random sample is that the sample is not random. Insiders buy small, cheap, volatile companies. If those companies happen to do well over the period studied, the study reports an insider effect that is really a small-cap effect.
The test that separates them is to measure the same companies over windows that have nothing to do with the purchase. We take each event, shift its entry forward by 180 trading days, and repeat the entire measurement, benchmark included. If the effect is about timing, the shifted measurement should be flat. If it is about which companies insiders like, the shifted measurement should look the same as the real one.
We run this as a paired difference on the events where both measurements exist, rather than as two separate averages, because a shifted window runs past the end of the data for late events and comparing two differently composed samples is not a test of anything. We also repeat it at a 252-day shift, where the one-year windows do not overlap at all. The exclusion rule that protects the real windows from corporate-action artefacts is applied to the shifted windows too. Without that symmetry the placebo can inherit an unrecorded reverse split the real window was protected from, which we observed directly during development.
We deliberately do not use the window before the purchase as the placebo, even though it is available for far more events. Insiders are documented contrarian buyers, so the pre-purchase window is mechanically depressed for this sample, and a paired difference against it would manufacture exactly the timing premium the test is supposed to be capable of rejecting.
What happens after an insider buys
Raw returns
| Horizon | Purchases, median | Purchases, mean | Median listed stock | S&P 500 |
|---|---|---|---|---|
| 1 month | -0.15% | 0.89% | 0.00% | 2.05% |
| 3 months | 0.95% | 5.08% | 0.02% | 5.29% |
| 6 months | 0.95% | 8.47% | -0.41% | 9.89% |
| 1 year | 2.36% | 19.36% | 0.00% | 17.84% |
The mean and the median disagree by a factor of eight at one year. That gap is the whole reason to be careful here. A mean of +19.36% describes a portfolio in which a handful of positions multiplied. A median of +2.36% describes what happened to a typical purchase. Both are true and they support completely different sentences.
Note the third and fourth columns. Over these same windows, the median listed stock returned 0.00% at one year while the S&P 500 returned 17.84%. The 2020 to 2025 period was one in which the median stock went nowhere and the index went up a great deal. Any benchmark decision made without noticing that will produce a confident and wrong answer.
Abnormal returns
| Horizon | vs S&P 500 | t | Beat | vs median stock | t | Beat | vs price quintile | t |
|---|---|---|---|---|---|---|---|---|
| 1 month | -1.54% | -5.76 | 43.17% | 0.12% | 1.47 | 50.51% | 0.01% | 1.37 |
| 3 months | -3.56% | -7.54 | 41.44% | 0.80% | 2.47 | 52.01% | 0.64% | 2.40 |
| 6 months | -7.10% | -9.48 | 38.36% | 1.07% | 2.74 | 52.02% | 0.96% | 2.33 |
| 1 year | -12.95% | -11.68 | 36.78% | 2.56% | 5.84 | 53.26% | 2.68% | 6.27 |
The two right-hand benchmarks agree closely, which is reassuring, because they are constructed differently. The S&P 500 column is a different question with a different answer: it says that a portfolio of insider-bought stocks would have badly trailed an index fund over this period, which is true, and which is a statement about small companies in 2020 to 2025 rather than about insiders.
The level effect does not survive a placebo
Here is the same measurement on the same events, with entry shifted forward 180 trading days, compared event by event.
| Horizon | Real | Placebo | Paired difference | t | Events |
|---|---|---|---|---|---|
| 1 month | 0.12% | 0.17% | 0.06 pts | 0.13 | 47,458 |
| 3 months | 0.80% | 0.49% | 0.64 pts | 1.29 | 47,458 |
| 6 months | 1.09% | 1.89% | 0.19 pts | 0.14 | 45,811 |
| 1 year | 2.61% | 2.78% | 1.40 pts | 1.58 | 41,140 |
The placebo is not null. Shifting these events six trading months into the future produces abnormal returns of the same size and the same sign as the real ones. At every horizon the paired difference is statistically indistinguishable from zero.
Three of the thirteen paired placebo specifications we ran on the level effect clear a t-statistic of 2. At a 252-day shift on this panel the one-year difference is 2.00 points, t 2.46. On the two-year panel described below, where the shift is 504 days, the one-year difference is 0.93 points, t 2.53, and the two-year difference is 3.20 points, t 2.64. All three sit at a horizon of a year or more, and two of the three are on the smaller two-year panel, whose events are weighted toward the 2020 and 2021 cohort. We report them as a residual rather than a result. A level effect that appears in three specifications out of thirteen, never inside six months, and mostly in one covid-weighted subsample, is not something a reader should act on. The one-month and three-month placebos are fully disjoint from their real windows at every shift, and all six are flat.
The straightforward reading is that these companies outperformed the median listed stock over this period whether or not an insider had just bought. Insiders buy a particular kind of company, that kind of company did modestly better than the typical listed name between 2020 and 2025, and the purchase date adds nothing to that.
When nobody has bought for two years
A purchase looks more interesting when it ends a long silence. If nobody inside a company has bought shares for years, the first person to step in may know that something has changed. We tested that version directly.
We count a drought only when there is an earlier open-market purchase in the record and the next one arrives more than two calendar years later. That avoids calling the start of the database a quiet period. The earliest qualifying event is 25 February 2022. The nested panel contains 858 of them across 830 companies and 481 disclosure dates; the shortest gap is 731 days and the median is 972.
Against the broad listed universe, the one-year result initially looks encouraging: the drought stocks finished 2.15 points ahead of the median listed stock. But their own median return was only 0.74%. They cleared a weak benchmark; they did not produce a strong return.
| Horizon | Drought stocks | Median S&P 500 constituent | Difference | Beat the constituent median |
|---|---|---|---|---|
| 1 month | 0.72% | 0.88% | -0.28 pts | 48.48% |
| 3 months | 1.78% | 2.81% | -1.23 pts | 46.97% |
| 6 months | 0.87% | 4.35% | -3.00 pts | 45.34% |
| 1 year | 0.74% | 7.50% | -6.55 pts | 43.59% |
The S&P constituent comparison changes the reading. After a year, the typical drought stock earned 0.74% while the typical constituent earned 7.50%, and only 43.59% of the purchases beat that middle constituent. The shortfall was therefore not confined to missing a handful of spectacular mega-cap winners; most drought stocks also lagged an ordinary S&P company.
This is a point-in-time comparison, not a test against today's winners. For each purchase we take the latest complete quarterly SPY holdings snapshot on or before the entry date, then measure the median constituent over the same market sessions as the purchased stock. Between 464 and 491 constituents have both a resolved company identity and usable prices. The list can be up to one quarter stale; substituting the following quarter's list leaves the one-year deficit unchanged at 6.55 points.
The capitalization-weighted index sets a much harder hurdle. The median event-level deficit widens from 6.55 points against the typical constituent to 16.47 points against the index. That difference matters because it identifies concentration, not a second insider effect: the largest S&P 500 companies carried far more of the period's return than a typical member.
The placebo test decides whether any of these gaps began with the purchase. In the 571-event paired sample, the one-year deficit to the median constituent was 7.89 points around the actual purchase and 5.89 points in a window shifted 180 trading days forward. That difference is not statistically reliable (t = -0.76). A fully disjoint 252-session shift reaches the same conclusion (t = -1.18), as do the pooled tests after clustering events by date.
That leaves a useful, but narrower, conclusion. A first insider purchase after two quiet years identifies companies that behaved differently from large-cap stocks throughout the surrounding period. It does not identify when that relative weakness began. The evidence therefore supports neither buying the drought purchase nor shorting it: both trades would mistake a persistent company profile for information contained in the filing date.
The size of the cheque does
Sorting the same 47,458 events into ten equal groups by dollar value gives a different picture.
| Decile | Events | Median cheque | Median entry price | 1 month | 3 months | 6 months | 1 year |
|---|---|---|---|---|---|---|---|
| 1 smallest | 4,746 | $1,200 | $7.79 | 0.69% | 1.73% | 1.65% | 1.11% |
| 2 | 4,746 | $4,989 | $11.61 | 0.02% | 0.35% | 1.20% | 3.06% |
| 3 | 4,750 | $10,974 | $13.16 | 0.12% | -0.29% | 0.34% | 0.72% |
| 4 | 4,741 | $20,686 | $15.00 | 0.27% | 1.64% | 2.58% | 4.67% |
| 5 | 4,746 | $35,231 | $15.69 | 0.30% | 0.22% | 1.62% | 4.59% |
| 6 | 4,746 | $58,927 | $14.79 | -0.05% | 0.41% | 1.14% | 4.89% |
| 7 | 4,745 | $103,367 | $16.15 | 0.00% | 0.96% | 1.42% | 3.46% |
| 8 | 4,746 | $208,973 | $17.04 | -0.03% | 1.11% | 1.17% | 2.76% |
| 9 | 4,746 | $514,828 | $18.30 | 0.08% | 1.48% | 1.14% | 2.70% |
| 10 largest | 4,746 | $3,526,520 | $20.86 | -0.19% | -0.52% | -1.53% | -1.80% |
Values are median abnormal returns against the median listed stock.
Nine deciles are positive at one year. One is negative. The relationship is not a gradient in any meaningful sense, it is a top-decile effect, and it is worth being precise about that because "bigger purchases do worse" implies a dose response the data does not show. The Spearman rank correlation between log dollars and one-year abnormal return is -0.014, and the per-date average rank correlation is -0.029 with a t-statistic of -2.60. At three months the rank correlation is -0.008 with a p-value of 0.07. These are the numbers of a weak effect concentrated at one end, which is exactly what the decile table shows.
And it survives the placebo
| Horizon | Real spread | t | Placebo spread | t | 252-day real | t | 252-day placebo | t |
|---|---|---|---|---|---|---|---|---|
| 1 month | -0.91 pts | -1.33 | -0.71 pts | -1.35 | -0.95 pts | -1.38 | 0.40 pts | 0.75 |
| 3 months | -3.28 pts | -2.82 | -0.60 pts | -0.53 | -3.67 pts | -3.16 | 1.64 pts | 1.56 |
| 6 months | -7.77 pts | -3.93 | 0.35 pts | 0.24 | -8.60 pts | -4.26 | 0.18 pts | 0.13 |
| 1 year | -17.10 pts | -3.20 | -1.97 pts | -0.99 | -16.61 pts | -2.92 | -0.71 pts | -0.31 |
Spread is decile 10 minus decile 1, computed on per-date medians, on the events where both the real and shifted measurements exist. On the full nested panel the real spreads are -0.91, -3.28, -7.00 and -11.77 points, with t-statistics of -1.33, -2.82, -3.54 and -2.39.
This is the opposite pattern to the level result. The real effect is large and significant at three, six and twelve months. The placebo effect is absent at every horizon under both shifts, with no t-statistic above 1.56 in magnitude. Whatever is driving the underperformance of the largest purchases is tied to the window in which the purchase happened, not to a persistent property of those companies.
Two years out
A two-year return exists only for purchases old enough to have one, so this section uses its own nested panel: 38,836 events on 3,543 stocks across 1,148 disclosure dates, 74.6 billion dollars, entries from 6 January 2020 to 13 August 2024. Its numbers at the shorter horizons do not equal the ones above, and that is composition rather than disagreement. Deciles are re-cut inside this panel so that the comparison is internal to it.
| Horizon | Purchases, median | Median listed stock | S&P 500 | vs median stock | t | Beat |
|---|---|---|---|---|---|---|
| 1 month | -0.15% | 0.00% | 2.09% | 0.15% | 1.36 | 50.65% |
| 3 months | 1.09% | 0.00% | 5.11% | 0.90% | 2.13 | 52.24% |
| 6 months | 1.31% | -0.86% | 10.03% | 1.58% | 3.81 | 52.96% |
| 1 year | 2.27% | -1.98% | 17.93% | 2.84% | 6.25 | 53.82% |
| 2 years | 2.53% | -1.43% | 36.75% | 6.49% | 10.58 | 56.26% |
The abnormal return against the median listed stock keeps growing, reaching 6.49% at two years with a t-statistic of 10.58, and the beat rate reaches 56.26%. Against the S&P 500 it goes the other way, to -27.15%. Both facts are driven by the same thing: over these two-year windows the median listed stock lost 1.43% while the index gained 36.75%.
The gradient at two years
| Horizon | Real spread | t | Placebo spread | t | Events |
|---|---|---|---|---|---|
| 1 month | -1.13 pts | -1.83 | 0.19 pts | 0.34 | 38,272 |
| 3 months | -2.58 pts | -1.98 | 0.41 pts | 0.46 | 36,619 |
| 6 months | -6.54 pts | -2.77 | -0.95 pts | -0.75 | 34,601 |
| 1 year | -16.58 pts | -2.32 | -1.40 pts | -0.64 | 29,904 |
| 2 years | -16.17 pts | -5.25 | -1.66 pts | -0.27 | 20,661 |
The placebo here is a 504-day forward shift, the only one that leaves a two-year window disjoint from its own. On the full two-year panel, without the placebo's sample restriction, the real spreads are -1.22, -3.30, -8.27, -16.70 and -15.26 points, with t-statistics of -2.01, -2.62, -3.79, -2.94 and -5.40.
Two things stand out. The two-year spread is the most statistically significant result in the study, at a t-statistic of -5.40 on the full panel and -5.25 on the placebo subset, and its placebo is flat. And the spread does not widen further between one year and two, it holds. Whatever the largest purchases are marking, it is fully expressed within a year and it does not reverse afterwards.
The decile table at two years has the same shape as at one year, more sharply. Deciles 1 through 9 run from 3.83% to 12.65%, and decile 10 is at -3.14%.
The one caveat this section carries is its cohort. Requiring 1,008 sessions of data after entry means the two-year placebo can only use events entered on or before 10 August 2022, so its sample is weighted toward 2020 and 2021. That is a covid-regime sample, and it is the reason we report the two-year level result as corroboration rather than as a separate finding.
Who fills the top decile
The composition of the largest decile is unusual, and it is the first thing to check.
| Decile | Purchases by 10% owners only | Purchases involving an officer |
|---|---|---|
| 1 smallest | 6.74% | 61.25% |
| 5 | 6.36% | 44.54% |
| 9 | 20.65% | 36.68% |
| 10 largest | 41.00% | 20.71% |
Two thirds of the largest decile is not what most people picture when they hear "insider buying". A chief executive putting 40,000 dollars into their own company is a different act from a private-equity holder adding to a 15% stake, and the second is where the large dollar amounts are. 10%-owner-only purchases are 12.21% of events but 50.76% of the dollars.
Roles are composition, not timing
The obvious next step is to cut by role, and the obvious conclusion from that cut is wrong. Here are the two extremes, with their placebos.
| Group | Events | Real, 1 year | Placebo, 1 year | Paired difference | t |
|---|---|---|---|---|---|
| 10% owner only | 4,876 | -1.15% | -5.22% | 3.32 pts | 1.15 |
| Director, no officer | 15,898 | 5.27% | 6.68% | 0.34 pts | 0.15 |
On the full nested panel the same cuts give -1.94% and 5.26% at one year, a seven-point gap that looks like a strong result about who to follow.
It is not. Both groups' placebos look like their real measurements. Companies with a large outside shareholder underperform whether or not that shareholder has just bought, and companies where directors buy outperform whether or not they have just bought. The role tells you something about the company. It tells you nothing about the moment.
The dollar effect is not only composition
Since 10% owners fill the top decile and 10% owners underperform, the dollar result could be nothing more than that. We tested it by removing every 10%-owner-only event and re-sorting the remaining 41,664 into fresh deciles, so the comparison is between a large purchase and a small purchase by the same kinds of insider.
| Horizon | Real spread | t | Placebo spread | t | Events |
|---|---|---|---|---|---|
| 1 month | -1.49 pts | -2.70 | -0.74 pts | -1.31 | 41,664 |
| 3 months | -2.14 pts | -1.88 | -0.25 pts | -0.23 | 41,664 |
| 6 months | -6.96 pts | -3.32 | 0.24 pts | 0.15 | 40,259 |
| 1 year | -13.36 pts | -2.33 | -0.32 pts | -0.13 | 36,264 |
The effect weakens at three months and holds at six months and one year. So composition explains part of the top-decile result and not all of it. A large purchase by an officer or director is still followed by worse relative returns than a small one.
Routine versus opportunistic insiders
Cohen, Malloy and Pomorski (2012) is the most cited refinement in this literature. Their claim is that pooling all insider trades hides the signal, because most insiders trade on a schedule, and that once you set the schedule-followers aside the predictive power sits entirely in the remaining opportunistic trades. If any cut of this data should show a timing premium, it is that one.
An insider is routine, in their construction, if they traded in the same calendar month in each of three consecutive prior years. We apply it per reporting person, then let an event inherit the classification: routine only if every filer behind it is routine, opportunistic only if every filer is opportunistic, mixed otherwise. Classification uses both purchases and sales, because routineness is a property of the trading pattern rather than of buying.
Two things constrain this test and we state them rather than work around them. First, our transaction corpus effectively begins in 2020, so three prior years of history exists only for events filed from 2023 onward: 22,334 nested events, of which 10,187 have a filer we cannot classify at all. Second, on this corpus the rule as literally stated is nearly automatic, because someone who trades in most months of every year always shares a month with themselves. Routine filers under the plain rule traded in a median of 5.00 distinct months per prior year, against 2.33 for opportunistic filers, which means the classification is capturing trade frequency as much as predictability. We therefore also report a stricter variant that additionally requires the filer to have traded in at most four distinct months in each prior year, which moves the split from 8,652 routine and 3,248 opportunistic to 2,424 and 9,433.
The placebo below is the 180-day shift. Note that the real column is a pooled median while its t-statistic is computed on the series of per-date medians, so the two can disagree in sign when the event count per date is uneven, as it is in the plain-rule opportunistic row.
| Rule | Class | Events | Real, 1 year | t | Placebo, 1 year | Paired difference | t |
|---|---|---|---|---|---|---|---|
| Plain | Opportunistic | 2,410 | -0.65% | 2.67 | 2.80% | -0.01 pts | 1.17 |
| Plain | Routine | 6,157 | 7.63% | 8.32 | 7.47% | 3.22 pts | 0.86 |
| Strict | Opportunistic | 6,809 | 4.35% | 6.12 | 5.28% | 2.07 pts | 0.25 |
| Strict | Routine | 1,722 | 9.55% | 5.87 | 8.28% | 1.72 pts | 0.61 |
The result does not reproduce, in either of the two ways it could have. The level effect is not concentrated in opportunistic trades; it is larger in routine ones under both rules, at 9.55% against 4.35% under the strict rule. And neither class shows a timing premium: at the 180-day shift every paired placebo difference is insignificant, at t-statistics of 1.17, 0.86, 0.25 and 0.61. At the 252-day shift they are 1.32, 2.38, 0.91 and 1.21. The one value above 2 there is the plain rule's routine class, which is the opposite of the class Cohen, Malloy and Pomorski predict, so it does not rescue the construction either.
We would not present this as a refutation of Cohen, Malloy and Pomorski. Their sample spans 1986 to 2007, ours 2020 to 2025 with only three years usable for classification, and their construction has access to a much longer trading history per person than ours does. What we can say is narrower and still useful: on recent data, with the history available, this cut does not rescue the timing premium that the pooled sample lacks.
Robustness
The two benchmarks that are not the S&P 500 agree throughout, at every horizon and in every cut. That is meaningful because they fail differently: the universe median is vulnerable to the composition of the listed universe, and the price-quintile match is vulnerable to entry price being a poor proxy for size.
The placebo is run at two shifts on the main panel and a third on the two-year panel. At 180 trading days the one-year windows overlap the real windows by 72 sessions, which is a genuine weakness for that row alone. At 252 and 504 days no window of the relevant length overlaps, and the conclusions do not change: the level result stays absent, the size result stays present.
Windows containing a corporate-action artefact are excluded from both the real and the shifted measurement by the same rule. We flag any single session in which a listed equity gains more than 300%, which is an unrecorded split rather than a return, and any split whose recorded ratio does not match the observed price jump across its effective date. Together those removed 325 one-year windows out of the 47,784 that had matured.
Every result is computed on a nested panel, so no horizon comparison is contaminated by sample composition, and the two-year section builds its own rather than borrowing this one.
Limitations
Survivorship. This is the most serious limitation, and it runs one way. Our price store holds listed common stocks, and companies that were delisted before the store was built are absent from it rather than present with a terminal value. Every well-known 2020 to 2025 bankruptcy we checked is missing entirely. Only nine price series in six years terminate early. So the level results are biased upward by an unknown amount, and the true abnormal returns are worse than the ones reported here. Shumway (1997) documents the same problem in CRSP and finds the correction material. Our two central claims are both comparisons drawn from within the same biased store, the placebo comparing the same stocks against themselves and the gradient comparing purchases against purchases, so the bias largely cancels there. The level tables are the part to distrust.
Unrecorded reverse splits. The microcap tail of the price store contains corporate actions that were never recorded, which is why 1.05% of 21-day universe returns exceed +100%, the largest reaches +137,999,904%, and 822 stocks show a single session above +300%. We handled this with a median-based benchmark and an explicit exclusion rule rather than by cleaning the underlying data. A study built on means without those defences would produce nonsense, and we produced some before catching it.
Stale quotes. A small mass of the listed universe does not move at all over a measurement window, 4.57% at one month and 1.17% at one year on a representative date. That mass pins the universe median at exactly zero on some dates. The price-quintile benchmark, which does not have this property in the same way, gives materially the same answers.
S&P membership history. The constituent benchmark comes from quarterly SPY holdings rather than an official daily membership file. The list used for an event can therefore lag a change by up to three months, and price-linked coverage rises from 91.7% to 97.4% across the period. Repeating the one-year test with the following quarter's list produces the same 6.55-point median deficit, but this remains a quarterly reconstruction rather than exact daily history.
No risk adjustment. These are raw and relative returns, not alphas. There is no factor model, no beta adjustment, and no volatility scaling. Insider-bought stocks are more volatile than the market, so a risk-adjusted version of the level result would be weaker than the one shown.
Not a trading strategy. Nothing here accounts for transaction costs, the bid-ask spread on the microcaps that dominate the sample, borrow costs on the short leg the size result implies, or the capacity of any of it. The median entry price in decile 1 is 7.79 dollars.
Rule 10b5-1. We report the flagged subset for completeness, 949 events, but the checkbox only exists on filings from April 2023 onward, so the cut is really an era cut and we draw nothing from it.
Period. January 2020 to August 2025 is one regime, and an unusual one. The median listed stock returned nothing over the average one-year window in it. The two-year placebo cohort is narrower still, weighted toward 2020 and 2021. We would not assume any of this generalises.
Reproducible specification
Sample: Form 4 non-derivative transactions with code Purchase and flag Acquired, superseded filings dropped, amendments anchored at the original filing date, aggregated to one event per stock and disclosure date. 131,691 transaction rows filed by 18,604 reporting owners give 56,379 events. Of those, 47,458 are usable at the first four horizons, and that nested panel is 112,647 transaction rows, 16,111 reporting owners, 3,859 stocks, 1,397 disclosure dates, 90.54 billion dollars, median event 47,528 dollars, 6 January 2020 to 15 August 2025. The two-year panel is a separate 38,836 events ending 13 August 2024.
Long-silence cut: events are ordered by company and disclosure date; the prior open-market purchase event must exist and be more than two calendar years earlier. The observed-predecessor rule avoids left-censoring at the 2020 start of the transaction corpus. The cut contains 858 events across 830 companies and 481 disclosure dates, with a median gap of 972 days.
Historical S&P constituent benchmark: the latest complete quarterly SPY holdings snapshot on or before each entry date, restricted to constituents with a resolved company identity and usable prices. Coverage runs from 464 to 491 constituents, or 91.7% to 97.4% of reported holdings. The benchmark is the median constituent buy-and-hold return over the event's identical market window. Constituent windows with a contradictory split or a single-session gain above 300% are excluded. Using the following quarter's membership instead leaves the one-year abnormal return unchanged at -6.55%.
Prices: 10,842,687 daily closes across 8,284 stocks and 1,665 market sessions, primary listing only, split-adjusted.
Entry: first market close strictly after the disclosure date. Horizons: 21, 63, 126, 252 and 504 trading days on the market calendar.
Benchmarks: S&P 500 buy and hold; median point-in-time S&P 500 constituent; median listed stock buy and hold over the identical window; median stock in the same entry-price quintile over the identical window.
Statistic: median abnormal return, with Newey-West t-statistics on per-date medians, lag length equal to the fourth root of the date count.
Placebo: entry shifted forward 180, 252 and 504 trading days, same benchmark, same exclusions, compared as a paired per-event difference on events where both measurements exist.
Conclusion
Three things follow from this that are worth separating.
The first is about the question everyone asks. Does a stock go up after an insider buys it? Over one year the typical purchase gained 2.36% while the typical listed stock gained nothing, which sounds like a yes. But shift the same purchases six months into the future and measure again, and you get the same answer. Whatever these companies had, they had it before the insider bought and they still had it afterwards. Insider purchases identify a type of company, not a moment. The long-silence cut does not change it either. A 0.74% one-year return beat a weak broad-stock median but lagged the typical S&P 500 constituent, and the same relative weakness appears away from the purchase date. That is a company profile, not a long or short signal. Nor does the split the literature says should isolate informed trades.
The second is about the question fewer people ask. When an insider writes an unusually large cheque, what follows is worse, not better, and unlike the first result this one does not appear when you measure the wrong window. The largest decile of purchases trailed the smallest by seven points over six months, by seventeen over a year, and by fifteen over two years, and by nothing at all under the placebo at any of them. Some of that is who writes large cheques, because large outside shareholders dominate the top decile and their companies underperform generally. But removing them entirely leaves most of the effect standing.
The most useful thing here may be the third result, the one about method. The same 47,458 purchases produce a 13% one-year deficit or a 2.6% surplus depending on nothing but what you subtract, and produce a real effect or no effect at all depending on whether you bother to shift the window. A number from this kind of study without a stated benchmark and a stated falsification test is not a weak result. It is not a result.
This note is for informational and educational purposes only. It is not investment advice.
References
Blume, M. E., and Stambaugh, R. F. (1983). Biases in Computed Returns: An Application to the Size Effect. Journal of Financial Economics, 12(3), 387 to 404.
Cohen, L., Malloy, C., and Pomorski, L. (2012). Decoding Inside Information. The Journal of Finance, 67(3), 1009 to 1043.
Jeng, L. A., Metrick, A., and Zeckhauser, R. (2003). Estimating the Returns to Insider Trading: A Performance-Evaluation Perspective. The Review of Economics and Statistics, 85(2), 453 to 471.
Lakonishok, J., and Lee, I. (2001). Are Insider Trades Informative? The Review of Financial Studies, 14(1), 79 to 111.
Newey, W. K., and West, K. D. (1987). A Simple, Positive Semi-Definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix. Econometrica, 55(3), 703 to 708.
Shumway, T. (1997). The Delisting Bias in CRSP Data. The Journal of Finance, 52(1), 327 to 340.
U.S. Securities and Exchange Commission. Form 4, Statement of Changes in Beneficial Ownership of Securities. https://www.sec.gov/files/form4.pdf
Equibles Research (2026). Does Daily Short Volume Predict Short Interest? Evidence From 660,000 Settlement Windows. https://equibles.com/research/does-daily-short-volume-predict-short-interest