Reassessing Guerrilla Firepower in Counterinsurgency Outcomes
Pischedda, Gilli, and Gilli's Weapons of the Weak: Technological Change, Guerrilla Firepower, and Counterinsurgency Outcomes (Journal of Conflict Resolution, 2025) makes a claim that is at once intuitive and underexplored: the absolute lethality of insurgents' weapons, rather than their lethality relative to the government's, is a significant predictor of whether insurgents win. The paper is careful, robustly tested, and fills a genuine gap. It also has identifiable statistical problems. This post explains both. It is also a proposal: these problems are solvable, and solving them together constitutes a new paper.
The Argument
The conventional wisdom in counterinsurgency (COIN) research focuses on incumbents, specifically their strategy, regime type, troop density, and mechanisation, as well as on insurgent attributes like external support and organisational cohesion. Weapons appear only tangentially, and when they do, scholars compare insurgent firepower to government firepower. Pischedda et al. argue this framing is wrong on its face. Guerrillas do not fight pitched battles. They ambush, raid, and disappear. What matters is not whether a guerrilla can match a tank, but whether their weapon meets a threshold of lethal force (e.g., killing a soldier) before they escape into terrain.
The paper traces a 150-year arc of small arms improvement, from smoothbore flintlock muskets (one shot per minute, unreliable in rain) to the Kalashnikov (one hundred rounds per minute, reliable in almost anything), and argues this arc cumulatively shifted the odds of insurgent strategic success. The mechanism is clean: more lethal small arms produce more casualties per hit-and-run attack, which erodes either the incumbent's military capacity or their political will to continue. Both paths to insurgent victory depend on the same tactical foundation: either the Maoist transition to conventional war, or the Mackian attrition of political resolve (Mack, 1975).
To test this, the authors build a new dataset, the Weapons of the Weak (WoW) dataset, covering 275 COIN campaigns from 1800 to 2005. They code what small arms insurgents used in each conflict, construct a composite Firepower index, and show it predicts insurgent success across a battery of regression specifications. They also find that once Firepower is controlled for, the previously influential mechanisation variable from Lyall and Wilson (2009) loses significance entirely.
The Dataset
The WoW dataset is the paper's most durable contribution. For each of 275 COIN campaigns, the authors code insurgent small arms across six categories:
Weapons of the Weak Dataset — Firepower Index (Range: 0–10)
├── Long Guns [Ordinal: 0–5]
│ ├── 0 — Cold weapons (swords, spears, machetes)
│ ├── 1 — Smoothbore flintlock muskets
│ ├── 2 — Single-shot breechloading rifles
│ ├── 3 — Repeating (magazine) rifles
│ ├── 4 — Semiautomatic rifles
│ └── 5 — Automatic (assault) rifles
└── Support Weapons [Binary: 0 or 1 each]
├── Machine guns
├── Submachine guns
├── Mortars
├── Explosives (grenades, landmines, IEDs)
└── Portable antiarmor weapons (RPGs, bazookas, recoilless rifles)
Firepower = Long Guns score + sum of support weapon dummies
Examples:
Navajo (1860s) → Firepower = 0 [cold weapons only]
Arab insurgents (WWI) → Firepower ≈ 4 [repeating rifles + explosives]
Afghan Mujahideen (1980s) → Firepower = 10 [automatic rifles + all support weapons]
Sources range from intelligence reports and academic case studies to participants' memoirs and conflict encyclopaedias. The variable is coded at peak level (i.e., the highest firepower reached at any point during the guerrilla phase) and a practical concession given the sparse historical record for many 19th-century conflicts.
The dependent variable, Insurgent Success, is a binary indicator: 1 for insurgent victory or draw, 0 for insurgent loss. The central finding is that a one-standard-deviation increase in Firepower is associated with approximately a 19% increase in the probability of insurgent strategic success (Pischedda et al., 2025); this effect is comparable in magnitude to external aid and post-WWII anti-colonial norms, both considered major drivers in the literature (Pischedda et al., 2025).
What the Paper Does Well
The theoretical reframing from relative to absolute firepower is both simple and genuinely important. It resolves a conceptual confusion that had persisted in the literature; namely, the incoherence of measuring guerrilla capability against government capability in a fight where guerrillas deliberately avoid confronting government capability directly. That correction alone is worth the paper.
The robustness of the Firepower finding is also notable. It survives controls for Cold War effects, external support, regime type, incumbent capabilities, distance, post-WWII normative change, rebel strength, ideology, troop density, mass media, and literacy. Sensitivity analysis using Cinelli and Hazlett's (2020) framework shows that even a confounder as strong as external aid would be insufficient to eliminate the result. The authors also transparently correct 21 outcome codings from the original Lyall and Wilson (2009) dataset; this level of meticulous replication and self-correction reflects a rare, refreshing commitment to methodological integrity.
Constructive Criticisms
1. The Index Is a Strong Assumption
Firepower is constructed by simple summation — long gun score plus five binary dummies, each contributing equally. This assumes that every weapon category contributes equally to the underlying construct of lethality, and that the scale is linear. Neither assumption is argued for. An automatic rifle almost certainly contributes more to guerrilla effectiveness in a hit-and-run context than a mortar, but the index weights them identically. The jump from musket (1) to single-shot breechloader (2) is treated as equivalent to the jump from repeating rifle (3) to semiautomatic (4), despite the former representing a vastly larger leap in practical lethality.
The solution is to follow what is perhaps a more well-trodden path in modern measurement modeling: treating Firepower as a latent variable estimated from the indicator data rather than a constructed sum. Item Response Theory (IRT) (Wattenberg, 1991) or a Bayesian latent variable model would estimate each weapon category's contribution to the underlying construct, recover a properly scaled score, and, crucially, attach uncertainty to each case's Firepower estimate. That uncertainty is not cosmetic; it propagates into downstream estimates and affects inference.
2. The Secular Trend Problem
This is, arguably, the deepest statistical issue. Firepower trends upward monotonically over 200 years. So does insurgent success. When two series trend together over time, they correlate mechanically, regardless of causal connection. The paper addresses this with decade fixed effects, Cold War dummies, and a post-WWII normative variable. These are discrete categorical controls for what is fundamentally a continuous, slow-moving temporal process. They help but do not resolve.
The diagnostic question is: does the relationship between Firepower and insurgent success survive after removing shared secular trend? Wavelet multiresolution analysis (Percival and Walden, 2000) can decompose both series into trend, cyclical, and idiosyncratic components. If the relationship disappears in the detrended components, the finding is largely spurious. If it persists, the authors have stronger evidence than any set of categorical controls provides. This is not a critique of the paper's conclusions (which I suspect will ultimately hold up), but rather an acknowledgement that a truly thorough verification test has yet to be conducted.
3. Average Effects Conceal Heterogeneity
The paper estimates a single average effect of Firepower across 275 conflicts spanning two centuries. This is almost certainly wrong as a description of reality. The effect of small arms lethality on outcomes should differ substantially across conflict types. In a colonial war where the incumbent is fighting far from home with limited political will and a domestic audience watching casualties, even modest guerrilla firepower may suffice to erode resolve. In a civil war where the government is fighting for survival with no exit option, the same firepower level may be insufficient. The paper tests some of this with subgroup analyses but does not systematically map where Firepower matters most.
Causal forests (Athey and Imbens, 2016), a machine learning method for estimating heterogeneous treatment effects, would estimate how the effect of Firepower varies across the full covariate space without requiring pre-specification of which interactions matter. The output is not a single number but a function: here is the effect of Firepower for conflicts that look like this, and a different effect for conflicts that look like that. This is more theoretically informative and generates novel propositions.
4. The Binary Outcome Discards Information
Collapsing war outcomes into a binary (i.e., insurgents succeed or they do not) throws away two kinds of information. First, it merges insurgent outright victory with negotiated draws, which are arguably different phenomena. Second, and perhaps more importantly, it ignores timing. A conflict that ends in insurgent victory after three years is meaningfully different from one that ends the same way after thirty. The relevant question is not just who wins but whether firepower accelerates insurgent victory, delays incumbent victory, or shifts terminal state distributions toward settlement.
A competing risks survival model (Fine and Gray, 1999), which utilizes separate hazards for incumbent victory, insurgent victory, and negotiated settlement, is more faithful to the data-generating process and recovers richer findings; crucially, it also handles censoring properly, whereas the binary outcome model does not.
5. No Out-of-Sample Validation
Every model in the paper is evaluated in-sample. In-sample fit is necessary but not sufficient evidence that a relationship is real. With 275 observations and a slowly trending predictor, it is possible to achieve good in-sample fit through shared trend alone. The natural test, of course, is to train the model on pre-1945 conflicts and predict post-1945 outcomes: this specific cross-validation has not yet been done. Post-1945 is precisely the period where the secular reversal in COIN outcomes occurs; therefore, a model that learns the pre-1945 pattern and successfully anticipates that reversal provides genuinely different, stronger evidence than a standard in-sample regression.
What Comes Next
These criticisms point toward a coherent research programme rather than a scattershot of methods. The path is:
First, estimate Firepower as a latent variable with uncertainty. Second, decompose the secular trend to test whether the signal survives detrending. Third, use causal forests on the pre-1945 training set to map heterogeneous effects. Fourth, build a predictive model (i.e., appropriate for small N, with kernel structure encoding temporal and network dependence) and evaluate it out-of-sample on post-1945 conflicts. Fifth, reformulate the outcome as a competing risks survival process. Sixth, propagate measurement uncertainty from the first stage through to prediction intervals at the final stage.
Each step addresses a specific identified problem and feeds into the next. Together they constitute a paper whose contribution is distinct from the original: not whether firepower matters on average, but whether the relationship is real enough to survive a predictive test, where it matters most, and what it implies for ongoing conflicts where these quantities can in principle be observed in near-real-time.
Pischedda et al. found a compelling correlation; the question is whether it is robust enough to predict. That is a harder and more consequential test: one that has not yet been applied.
References
Athey, S., & Imbens, G. (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27), 7353–7360.
Cinelli, C., & Hazlett, C. (2020). Making sense of sensitivity: Extending omitted variable bias. Journal of the Royal Statistical Society Series B, 82(1), 39–67.
Fine, J. P., & Gray, R. J. (1999). A proportional hazards model for the subdistribution of a competing risk. Journal of the American Statistical Association, 94(446), 496–509.
Lyall, J., & Wilson, I. (2009). Rage against the machines: Explaining outcomes in counterinsurgency wars. International Organization, 63(1), 67–106.
Mack, A. (1975). Why big nations lose small wars: The politics of asymmetric conflict. World Politics, 27(2), 175–200.
Percival, D. B., & Walden, A. T. (2000). Wavelet methods for time series analysis. Cambridge University Press.
Pischedda, C., Gilli, M., & Gilli, A. (2025). Weapons of the weak: Technological change, guerrilla firepower, and counterinsurgency outcomes. Journal of Conflict Resolution, 69(9), 1553–1579.
Wattenberg, M. F. (1991). Item response theory. Cambridge University Press.