Peter Cotton
Book Length
-
Microprediction: Building an Open AI Network — MIT Press, 2022.
-
An Analytic Approach to Ornstein-Uhlenbeck Processes with Fluctuating Parameters and Applications in the Modeling of Fixed Income Securities — PhD thesis.
Portfolio construction & covariance · precise · schur.microprediction.org/papers
-
Schur Complementary Portfolios — preprint, 2024.
abstract
Despite many attempts to make optimization-based portfolio construction in the spirit of Markowitz robust and approachable, it is far from universally adopted. Meanwhile, the collection of more heuristic divide-and-conquer approaches was revitalized by Lopez de Prado where Hierarchical Risk Parity (HRP) was introduced. This paper reveals the hidden connection between these seemingly disparate approaches.
-
Two Sides of Schur Damping: High-Dimensional Pseudo-Likelihoods and Portfolio Allocation — preprint, 2026.
abstract
Two communities that rarely cite each other -- spatial statisticians fitting high-dimensional weather fields, and quantitative investors building portfolios -- have independently arrived at the same mathematical object: a Schur complement, damped by one interpretable parameter. In spatial modeling the Schur complement is the conditional covariance that makes a Gaussian (Vecchia) pseudo-likelihood estimable at scale, and recent work regularizes it by shrinking toward a base model. In allocation it is the residual risk of a bet net of its hedge, and the same parameter interpolates hierarchical risk parity and the minimum-variance portfolio. We show these are one operation -- reliability shrinkage of a conditional Gaussian -- so that the damping a weather model needs to remain estimable when stations outnumber observations is, term for term, the damping a portfolio needs to remain stable when assets outnumber returns. The optimal amount is a closed-form reliability, a James-Stein shrinkage that is simultaneously a Ledoit-Wolf intensity. The shrinkage machinery is classical, but the identity appears to be new: to our knowledge neither literature has noted that the conditional shrinkage a spatial model fits and the diversification-variance tilt a portfolio chooses are one and the same quantity. We make the correspondence precise, note that the two literatures have each supplied what the other lacks, and report a small experiment on the one genuinely open choice -- how to set the damping -- suggesting the spatial community's fitted intensity is, if anything, the better recipe.
-
Betwixt Minimum Variance and Hierarchical Risk Parity: Analytical Results for the Schur Bridge
abstract
The canonical Schur-complementary construction interpolates hierarchical risk parity (γ = 0) and minimum variance (γ → 1) through a coupling γ that sets how much of the estimated cross-block covariance the recursion uses. We study the implemented collapse variant, which shares the hierarchical endpoint and reaches minimum variance under additional structure. We ask where on the bridge to sit when the covariance is estimated with error, and answer in three layers. A structural theorem shows that top-level cross-block noise is invisible at the hierarchical end and prices coupling at order γτ 2 , so whenever coupling has first-order population value the optimum stays uniformly separated from hierarchical risk parity at all sufficiently small noise levels. A local theorem at an exact minimum-variance endpoint shows that positive marginal noise sensitivity pushes the optimum into the interior. An exchangeable four-asset family is then solved exactly: its unique optimizer is interior at every admissible noise level and decreases monotonically from full coupling to the middle of the bridge. Exact examples mark the boundaries of what holds in general, and every exact claim is verified in rational arithmetic.
-
Schur Pseudo-Likelihood
abstract
We introduce the Schur pseudo-likelihood, a one-parameter family that damps the cross-block coupling of the Gaussian likelihood through its Schur complements, for scoring and regularizing covariance and correlation estimates in high dimensions. The Gaussian log-likelihood is the standard criterion, but when the dimension rivals the sample size it is governed by the smallest, least-identifiable eigenvalues of the estimate, and as a criterion for ranking or selecting estimates it becomes unreliable; damping the coupling restores it. The optimal damping has a closed form—the reliability of the coupling, a James–Stein shrinkage. On real crypto returns, choosing a shrinkage estimate by the Schur pseudo-likelihood rather than by the full likelihood yields several-fold lower out-of-sample portfolio variance once the dimension rivals the sample size.
-
Schur-Regularized Vecchia: Structured Covariance Estimation at Scale
abstract
When the dimension $p$ is large—thousands of weather stations, or the width of a neural network layer—the Gaussian likelihood cannot be formed or inverted, so both estimation and evaluation must exploit known structure. We make three points and one new one. (i) A structure-aware estimator beats global shrinkage out-of-sample when undersampled: on a real $323$-point weather forecast-error grid, spatial covariance tapering beats Ledoit–Wolf by several nats per observation at $n/p<1$, converging as $n$ grows. (ii) The block-conditional (Vecchia) likelihood, the only one computable at scale, is a faithful stand-in for the full likelihood—it selects the same structured estimate $\approx100\%$ of the time where both can be computed—at $O(p\,m^2)$ instead of $O(p^3)$. (iii) Crucially, Schur damping adds to Vecchia: plain Vecchia trusts its neighbour regressions fully ($=1$), but those are estimated from limited data and overfit; damping them by the coupling reliability $$ improves out-of-sample likelihood substantially when undersampled (by tens of nats at $n/p=0.4$, where plain Vecchia is in fact the worst setting), with $$ rising to $1$ as $n$ grows. The optimal damping is a closed-form James–Stein reliability, the same quantity that governs block shrinkage and Schur-complementary allocation.
-
Correlation Inflation
abstract
Traditional covariance shrinkage estimates pull an empirical covariance matrix Σ towards a lower information anchor (often λI), reducing estimation error at the cost of suppressing correlation. This note proposes the opposite: geodesic covariance inflation—a controlled movement of Σ towards a perfectly-correlated limit. Intended not as a consistent estimator but merely a device for use in the context of portfolio construction, we propose that instead of discounting correlation, we exaggerate it in a geometrically natural way, yielding a family of positive–definite matrices indexed by an inflation parameter γ ∈ [0, 1].
Allocation & attribution from winning probabilities · allocation · allocation.microprediction.org/papers
-
Thurstone Portfolio Polishing: Tail-Sensitive Black-Litterman and Beyond
abstract
We develop Thurstone portfolio polishing: a way to take an existing allocation on the simplex—a capitalization index, an equal-weight book, any benchmark—and re-tilt it under a view on dependence rather than returns. Reading the weights as the winning probabilities of a race among assets (a Thurstonian model of choice), we back out the latent abilities that reproduce them with a fast exact inverse, then re-evaluate the race under a richer dependence law. The marquee instance is a tail-sensitive analogue of Black–Litterman: where Black–Litterman tilts a benchmark by a view on expected returns through a linear inverse of a covariance, we tilt by a view on lower-tail co-movement through a fast exact nonlinear inverse of a correlated race, inverting no covariance and solving no optimization. Driven by a tail-dependent simulation the polish shades down exactly the names that crash together beyond what their correlation can encode—a crash-proofing overlay we verify both in a known-truth lab and on the Dow. The construction's organizing property is redundancy (clone) consistency: a near-duplicate asset splits a position rather than doubling it—the cure for the red-bus/blue-bus failure of any Luce-type rule, and the same near-singularity that simultaneously makes the polish robust to ill-conditioning (it samples from a correlation rather than inverting one), low in turnover (its weights are locally Lipschitz in the correlation), and a smooth de-duplicator of correlated holdings. The polish is feasible on the simplex without optimization, reproduces its benchmark by construction, admits an arbitrary performance simulation so dependence beyond second moments can be matched, and solves an implied mean–(convex-regularizer) program whose penalty the simulation fixes; capitalization weighting is its degenerate, correlation-blind member. We prove the supporting properties (feasibility, benchmark reproduction, redundancy and tail consistency, the implied objective, and a Lipschitz turnover bound) and implement the method in the allocation package, built on thurstone.
-
Winning Probabilities as Credit: Fast, Redundancy-Aware Attribution
abstract
The probability that a competitor wins a noisy race—the Thurstonian choice probability, equivalently the gradient of an expected-maximum potential—is a rule for attributing credit among correlated contributors. It is a probability share ($\sum_i w_i = 1$), symmetric, and redundancy-aware (contributors that move together share a single contributor's credit rather than double-counting); it is differentiable in the contributors' abilities, it updates online, and it costs $O(Mn)$ in a single Monte-Carlo pass with no coalition enumeration. It is not a Shapley value and does not try to be: Shapley measures coalitional cooperation—a contributor's average marginal value as coalitions form—whereas the winning probability measures selection relevance, how often a contributor is the single best in the field at hand. The two answer different questions, and a contributor that is never individually best yet improves every blend it joins earns Shapley credit but little race credit; the race does not capture such coalitional value. (There is a clean mathematical link—the winning-probability field integrates along the diagonal to an Aumann–Shapley value of an ability-scaling game—which we record but do not lean on.) We show that the headline redundancy property—an even split of credit among near-duplicates—is a property of the calibrated equal-ability race rather than of raw scores, and we demonstrate the rule on forecast combination and feature attribution. This develops the credit reading of the construction in the companion paper on Thurstone portfolios.
-
When Does Portfolio Construction Work? A Map across the Number of Assets and the Cost of Trading
abstract
A portfolio rule is a covariance estimate feeding an allocator. We ask a plain question: which rules actually work out of sample—after trading costs, and as the number of assets grows? We backtest eleven online constructions over thousands of random subsets of U.S.\ stocks and draw a map of the winner against two axes, the number of names and the cost of trading. The map has three regions. With few names and cheap trading, classic minimum-variance wins. With many names and cheap trading, a factor-model minimum-variance wins—and by a margin that grows with the number of names—because the ordinary version becomes undefined once there are more assets than data. Once trading costs rise above about five to ten basis points, simple inverse-variance weighting wins everywhere, because turnover, not estimation error, is what hurts. We then give the two ingredients that make the many-names corner work: a Woodbury inversion of a factor covariance, and a ridge-regularized version of the Schur coupling.
Conformal prediction · conformalprediction · conformalprediction.net/papers
-
Marginally Useful: Formalizing the Information Gap in Conformal Prediction
abstract
Conformal prediction gives finite-sample, distribution-free marginal coverage for a set. The guarantee is real, and it is often misread as evidence of forecast quality. We separate the two with one decomposition, the residual-information gap: for a fixed location predictor and a single-shape residual predictive system, the log-score regret relative to the oracle is exactly the mutual information $I(R;X)$ between the residual and the input. Conformalization re-levels coverage but cannot touch this quantity, because it is a property of the predictor's shape class and not of calibration; no recalibration that ignores $X$ reduces it within that class. The familiar cautions about conformal prediction follow as context: marginal coverage is not conditional, validity is insensitive to sharpness, and the guarantee needs exchangeability.
-
An Empirical Study of the Conformal Information Gap
abstract
Every forecasting architecture restricts the information available to its predictive law. Under logarithmic loss the irreducible cost of replacing the full input by a retained representation is exactly the conditional mutual information between outcome and input given the representation, which we call the information gap. For a single pooled residual law the gap is the mutual information between residual and input, and pooling after any invertible state-dependent change of coordinates has the analogous gap in the transformed coordinates. Conformal calibration improves estimation and provides a coverage certificate, but cannot recover predictive information the representation excludes. Applied to a nested grammar of online transforms on economic series, conditional scale is the largest and most reliable component of the measured architecture gain.
-
A Feynman–Wigner-Style Diagnostic for the Efficacy of Conformal Prediction via Signed de Finetti Representations
abstract
Conformal prediction builds prediction sets that cover the truth at a rate you choose, finite-sample and distribution-free, assuming only exchangeable data. That guarantee is marginal. de Finetti's theorem describes the exchangeability it rests on, and in the finite form the mixing measure may be signed kerns2006. A short lemma decomposes the slope of conformal's calibration-conditional coverage into a non-negative threshold term (the classical Beta-law fan) and a term carrying the sign of the de Finetti measure. Positive (extendable) mixtures make conformal conditionally adaptive; the signed corner of de-meaned, ranked, or compositional scores makes it anti-adaptive. The marginal guarantee is the same either way; only what it hides changes.
-
Betting Against a Conformal Predictor: A Parimutuel Account of the Information Gap
abstract
The companion paper shows that the log-score regret of a single-shape conformal predictor to the conditional oracle is the mutual information $I(R;X)$ between the residual and the input. Here we rederive that quantity from a betting mechanism rather than from the score. Treat the predictor as the crowd in a parimutuel pool on the residual: bettors put money in, and the pot is split among the winners in proportion to their stake on the realised outcome. In the continuous limit the payoff is the ratio of the bettor's density to the crowd's. An entrant who knows only the marginal breaks even, which is marginal coverage stated as wealth. An entrant who conditions on $X$ grows his bankroll at rate exactly $I(R;X)$. The gap is the rent. This is not a metaphor: the nearest-the-pin pool of the microprediction platform, and the continuous density version run in the MidOne contest, are this mechanism, and a conformal predictor is the entrant that prices the pool flat in $X$. We then give two ways to measure the rent on a fitted predictor: a static lower bound from distance covariance, and a sequential e-process whose growth rate estimates $I(R;X)$ and which is an anytime-valid test for conditional miscoverage.
-
The Width of the Conformal Fan: Dependence and the Variance of Realized Coverage
abstract
Fix a split-conformal calibration set; the coverage realized against an independent draw from the score marginal is $c=U_(k)$, distributed $Beta(k,n-k+1)$ for independent calibration scores, with variance about $(1-)/n$ — the fan. We show the fan's width is set by the dependence among the calibration scores, through the covariance kernel of the sub-threshold count. Three results follow: a leading-order fan coefficient in a Bahadur regime; an exact aggregate bound, summed over all levels, showing negative association never widens the fan; and, for equicorrelated normals, the exact single-level law $Var(Z_(k))=v_k+(1-v_k)$. The general single-level negative-association contraction is left as a conjecture, supported numerically. The governing sign is the one the companion note attaches to the finite de Finetti measure.
-
The Two Prices of Dependence: a Vanishing Coverage Tax and a Fixed Sharpness Rent
abstract
Temporal dependence charges split conformal prediction two different prices, and they scale differently. The coverage price is a tax that vanishes with the calibration size: for a stationary $$-mixing score process the loss is at most $\min_\/(n+1)+2()\$ barberpananjady2026. The sharpness price is a rent that does not: the per-step log-score regret of the marginally calibrated predictive against the past-conditional oracle equals the entropy-rate gap $ H=H(S_0)-h$, the mutual information between the present score and its infinite past, a constant independent of $n$. For Gaussian score processes the rent exponentiates into width — oracle intervals are narrower by the factor $e^- H$ at matched coverage, which for an AR(1) score process with parameter $$ is $1-^2$. Dependence is therefore asymptotically free for the certificate and permanently valuable for the forecast.
-
A Contragredient View of Conformal Placement and Steinitz Balancing
abstract
We give the validity of split conformal prediction an exact contragredient form: the placement acceptance map is equivariant for the permutation action, and the Reynolds identity $ P^f,v= f,Pv$ reads its orbit average either as moving the evaluation functional or as averaging the induced acceptance vector over placements. A natural primal discrepancy construction under the same action is Steinitz balancing: ordering a zero-sum population so that every prefix sum stays bounded. Conformal placements form such a population, and an explicit ordering keeps running coverage within $1/(2t)$ of the orbit level $k/N$, exact on the full orbit. On the function–measure pairing, the measure-side counterpart of calibration is balanced sampling and exact cubature. A transitive distributional symmetry, exchangeability in the standard setting, turns the deterministic orbit average into the coverage probability of the designated test placement.
Contests, ranking & choice models · winning · winning.microprediction.org · thurstone.microprediction.org
-
Inferring Relative Ability From Winning Probability in Multi-Entrant Contests — SIAM Journal on Financial Mathematics, 2021.
abstract
We provide a fast and scalable numerical algorithm for inferring the distributions of participant scores in a contest, under the assumption that each participant’s score distribution is a translation of every other’s. We term this the horse race problem, as the solution provides one way of assigning a coherent joint probability to all outcomes, and pricing arbitrarily complex horse racing wagers. However, the algorithm may also find use anywhere winning probabilities are apparent, such as with e-commerce product placement, in web search, or, as we show, in addressing a fundamental problem of trade: who to call, based on market share statistics and inquiry cost.
-
A Scalable Algorithm for Subset Selection and Rank Probabilities in Contests and Latent Variable Choice Models
abstract
A k-subset of items will be chosen from n according to values taken by n auxiliary variables X1 , . . . , Xn interpreted as performances in a contest. Item i is chosen if Xi ≤ X (k) where X (k) is the k’th order statistic. A numerical algorithm is presented for computing many k-combination choice probabilities quickly, for small k but potentially large n ≫ 1, 000, 000. No assumption is made on the 1-margin distributions of the Xi , and the analytical convenience survives the introduction of dependence via a factor model also. The computation of rank probabilities for k items is a corollary. The algorithm is provided in the winning package, on PyPI.
-
Luce's Choice Axiom Isn't the Only Choice! Combinatorial Contest and Rank Probabilities Using the Python Winning Package
abstract
A subset of k items will be chosen from n according to values taken by n variables X1 , . . . , Xn interpreted as performances in a contest. Item i is chosen if Xi ≤ X (k) where X (k) is the k’th order statistic. A numerical algorithm is presented for computing many k-combination choice probabilities quickly, for small k but large n ≫ k in the millions. Rank probabilities for k can also be computed. Some users of the winning package may wish to calibrate models for latent Xi from partial information such as winning or losing probabilities. Others may prefer to supply arbitrary performance distributions. The analytical convenience of this method also survives the introduction of dependence in the Xi via a low-dimensional Copula.
-
A Paradox in Machine Preference
abstract
Using prompts such as: “my favorite state in the US is [MASK]”, and “my favorite Western state in the U.S. is [MASK]” we infer that Thurston models are a better match to the revealed preferences of large language models than the application of Luce’s Choice Axiom. There is some irony in this finding given that Softmax functions, responsible for the token probabilities we interpret as preference, suggest Independence of Irrelevant Alternatives.
-
Rating Formula 1: a Case for Non-Gaussian Noise in Rating Systems
abstract
Rating systems almost universally assume that contest performance is ability plus normal or logistic noise. We present a setting where the assumption is visibly false and measurably costly: Formula 1, where for decades the most likely outcome of starting a race was not finishing it. A lattice-based Thurstonian rater that accepts arbitrary noise densities outperforms TrueSkill, Elo, Glicko-2 and OpenSkill over 1,158 races, and replacing its normal noise with a two-component density, a Gaussian pace term plus a separated block of slow mass representing retirement, improves every metric. Estimating the block's mass from the trailing retirement rate improves it further. Symmetric heavy tails and skew do not help, so the gain is attributable to the block, not to non-normality in general.
Distributional prediction & microprediction · skaters · skaters.microprediction.org/papers
-
Self-Organizing Supply Chains for Microprediction: Present and Future Uses of the ROAR Protocol — preprint, 2019.
abstract
A multi-agent system is trialed as a means of crowd-sourcing inexpensive but high quality streams of predictions. Each agent is a microservice embodying statistical models and endowed with economic self-interest. The ability to fork and modify simple agents is granted to a large number of employees in a firm and empirical lessons are reported. We suggest that one plausible trajectory for this project is the creation of a Prediction Web.
-
A Platform for Assessing and Combining Autonomous Short-Horizon Distributional Predictions
abstract
The operation of a novel open-source platform where mostly autonomous algorithms are tasked with predicting a large variety of streaming data is described. Reward and combination is achieved by means of a near-the-pin mechanism generalizing lottery countbacks to continuous space.
-
How Should Forecasts be Engineered: The Indispensible Markets Hypothesis
abstract
We consider evidence in support of what we term the Weak Indispensable Markets Hypothesis (IMH) provided by the M6 Financial Forecasting Competition. The Weak IMH asserts in the presence of a well-established market, those who eschew prices as inputs for proximate predictive modeling tasks will under-perform out of sample. We also consider the Strong IMH which asserts that forecasting should be considered a market-inspired engineering endeavor in addition to a modeling task under the usual rubric of statistics or machine learning methods (put simply: if a market doesn’t exist to help you, make one!). The competition established that neither principle’s application is obvious to participants or organizers; it hinted at a hidden quality crisis in data science generally; and it suggests that broadening the usual concept of analytic pipelines to insert collective intelligence might be part of the remedy.
-
Transforms All the Way Down: Automatic Online Distributional Forecasting by Conjugation
abstract
The Python package skaters is an online, distributional, univariate timeseries forecaster built by conjugation: invertible transforms nest all the way down onto a single distributional leaf fitted by a proper scoring rule. The collection collapses into one forecast function with no exposed tuning parameters, laplace, which leads the per-series held-out log-likelihood race against classical, neural, and foundationmodel baselines on FRED series; on asset prices a GARCH-t model remains better, a split we report rather than average away. Run in laplace’s coordinates and mapped back exactly (the sandwich), existing models improve dramatically without retraining. The library is implemented in pure Python (pip install skaters), zero-dependency JavaScript (npm install skaters), R, and a portable Rust core, held to 10−6 agreement by a shared parity suite, so models run unchanged on a server or in a browser.
Global optimization · humpday · humpday.microprediction.org/papers
-
HumpDay: Derivative-Free Optimizers in Pure Python and JavaScript, with a Contamination-Resistant Real-World Benchmark
abstract
The Python package humpday provides twenty-three derivative-free optimizers behind one uniform contract: minimise a black-box function on the unit cube under a hard evaluation budget. Every algorithm is implemented in pure Python with no required dependencies (a numpy backend is used transparently when present) and again in zero-dependency JavaScript, with parity tests holding the two in agreement, so the same optimizer runs on a server or in a browser. A recommender selects an algorithm from the problem’s dimension, budget, and measured evaluation cost, using rankings precomputed on the package’s own benchmark: eighty-one objectives ported from real applications, each paired with an interactive browser demonstration. Scoring applies seeded smooth bijections of the cube that relocate every optimum, so no method, human or machine, can score by memorising solutions; the suite remains valid when the candidates are written by language models. We describe the design and two findings the benchmark has produced: optimizer rankings on synthetic test functions correlate only moderately with rankings on the real-world suite (Kendall τ between 0.42 and 0.49 across budgets), and the suite supported the generation and out-of-sample validation of Alloy, a machine-designed optimizer that now ships in the package.
-
Neither Synthetic Benchmarks nor Common-Sense Reasoning Predicts Real-World Derivative-Free Optimizer Performance
abstract
Practitioners choose a derivative-free optimizer either by consulting a benchmark leaderboard or by reasoning about the problem. We test both. Using a memorisation-proof suite of real-world objectives, each relocated by a seeded cube-to-cube diffeomorphism, we rank a panel of $22$ optimizers and ask how well that ranking is predicted by the same panel's ranking on synthetic analytic benchmark functions, and by a language model asked to reason, problem by problem, about which optimizer should win. Both predictors are essentially uninformative: the rank correlation with real-world performance is statistically indistinguishable from zero in either case. A language-model selector that picks an optimizer per problem in fact does worse than a constant policy of always using one robust direct-search method. The two failures share a mechanism. Both the benchmark and the model over-trust model-based trust-region methods (Powell's NEWUOA and BOBYQA), which dominate smooth analytic functions yet rank near the bottom on the real suite. This is a draft, and several confirmatory runs are still in progress (Section sec:caveats).
-
Discovering a Derivative-Free Optimizer by Evolution Against a Memorisation-Proof Real-World Benchmark
abstract
We discover rather than hand-design a derivative-free optimizer by evolving a parameterised DE/ES-hybrid template against a memorisation-proof suite of real-world objectives. A genetic algorithm searching $14$ behavioural knobs reaches a normalised regret of $0.110$ versus a Nelder–Mead/Differential-Evolution/CMA-ES panel (mean rank $1.47$ of $4$), against $0.191$ for the original $12$-knob template. The gain comes from one added mechanism — a cheap, R2-gated separable-quadratic trust-region jump — which we validate three ways: an A/B re-evolution, a single-gene toggle ablation (disabling it more than doubles regret), and independent re-discovery by the production search, which drives the jump's firing probability to $\!1$. The discovered optimizer is a hedged hybrid, found by grounding fitness in real (disguised) evaluation rather than a synthetic proxy.
-
The Inspiration Simplex: Using Derivative-Free Optimization in Concept Space to Create New Derivative-Free Optimizers
abstract
Place K established algorithms at the vertices of a simplex. A point is a vector of mixing weights, and a language model turns the point into a working program: the heaviest vertex supplies the architecture and the others graft their ideas into it, in the stated proportions (Figure 2 shows the literal prompt). Scoring the program on a benchmark makes the simplex a continuous space of algorithm designs that ordinary derivative-free optimizers can search. The construction’s first product is Alloy, generated at the equal-weights point: on twentynine problems never used in any selection step it has the best mean rank at every budget from 60 to 480 evaluations and beats six competitors including CMA-ES pairwise (p < 10−10 ); it ships in the humpday package. Letting weights go negative extends the recipes beyond the simplex; of six semantics for a minus sign, tested under a paired-ablation protocol, one survives out of sample, with a rule attached: negative weights repair weak hosts and spoil strong ones. Failures are reported throughout: the continuous search never beat the obvious point, single generations vary severalfold, and in-sample scores mislead reliably. Two further worked examples test the construction’s reach: cache-eviction policies, where static blends fail and adaptive shares rediscover the design of CAR, and distributional forecasting, whose champion dominates synthetic suites, loses on real data, and ships anyway, recast as a transform inside the forecaster it could not beat.
-
Go Forth! Simple Detection of Incomplete Meta-Learning by Algorithms Performing Limited Exploration on a Rugged Landscape
abstract
It is shown that an algorithm given two chances to improve its position on the path of an exponentiated Ornstein-Uhlenbeck (OU) process should not choose its final position between the first two locations. It is sometimes easy, therefore, to diagnose failure of an algorithm to learn the optimal policy. The proof introduces the notion of rapidity on an OU bridge, and complements similar results in managerial science that are used as metaphors for complex but ill-defined business and strategy problems.
Market mechanisms & scoring rules · mechanisms · mechanisms.microprediction.org/papers
-
Schur Damping for Perpetual Demand Lending Pools
abstract
Decentralized perpetuals protocols have collectively reached billions of dollars of daily trading volume, yet are still not serious competitors on the basis of trading volume with centralized venues such as Binance. One of the main reasons for this is the high cost of capital for market makers and sophisticated traders in decentralized settings. Recently, numerous decentralized finance protocols have been used to improve borrowing costs for perpetual futures traders. These protocols have grown to over $2.5 billion dollars of assets while generating over $890 million in fees in 2024. We formalize this class of mechanisms utilized by protocols such as Jupiter, Hyperliquid, and GMX, which we term Perpetual Demand Lending Pools (PDLPs). We then formalize a general target weight mechanism that generalizes what GMX and Jupiter are using in practice. We explicitly describe pool arbitrage and expected payoffs for arbitrageurs and liquidity providers within these mechanisms. Using this framework, we show that under general conditions, PDLPs are easy to delta hedge, partially explaining the proliferation of live hedged PDLP strategies. Our results suggest directions to improve capital efficiency in PDLPs via dynamic parametrization.
-
Scoring Point-Cloud Distributional Submissions: Nearest-the-Pin Parimutuels, the KDE Seam, and Mollified Scoring
abstract
Point-cloud forecasts are often evaluated by smoothing the submitted samples into a kernel density estimate and scoring that density at the realised outcome. This apparently natural procedure is not proper: under logarithmic scoring at the raw outcome, a forecaster is generally rewarded for submitting samples from a deconvolution of their belief by the smoothing kernel, rather than from the belief itself. For Gaussian beliefs and Gaussian kernels this incentive has the simple form "shave $h^2$ from the covariance." We show that the defect is repaired by adding outcome noise from the same kernel used to smooth the submission. More generally, composing a proper score with a fixed Markov kernel preserves propriety, and strict propriety for the original report is recovered exactly when the induced channel is injective; for convolution kernels this is equivalent to the set where the characteristic function is nonzero being dense. For Gaussian kernels, repeating the repaired score over a ladder of smoothing scales decomposes log-score regret, via the relative de Bruijn identity, into Fisher-divergence bands plus a coarse-scale KL term. We also describe a high-dimensional alternative based on random one-dimensional projections, whose average CRPS is, up to an explicit dimension constant, the multivariate energy score. The results are population-level: finite clouds, endogenous bandwidths, and finite-player equilibria are left open. Keywords: proper scoring rules; continuous ranked probability score; energy score; kernel density estimation; distributional forecasting contests; score matching; sliced scores JEL classification: C53, C52, C14, D81 --
-
An Algebra of Prediction-Rewarding Mechanisms: Scoring Rules, Market Makers, and Pools as Composable Transducers
abstract
Scoring rules, cost-function market makers, constant-function market makers, parimutuel pools, and opinion pools are developed in separate literatures. This note is expository: it gathers the connections among them under one convex-analytic lens and one notational convention. A proper scoring rule is a convex potential read off the report; the Fenchel conjugate of that potential is a cost-function market maker, and a level-set dual is a constant-function market maker. The linear and logarithmic opinion pools are the two Kullback-Leibler barycenters. Merging market makers is infimal convolution, the composition law of the risk-sharing literature, so liquidity adds. A common transducer signature lets the mechanisms be chained, and this note settles when propriety survives a transformation of a stage's message or outcome. The mathematics is classical; the contribution is its consolidation. The multi-stage theory built on this dictionary, and its deployment on the microprediction platform, is a companion report [@cotton2026multistage]; two further results that motivated the framing are elsewhere too, sample-based elicitation on point clouds [@cotton2026pointcloud] and the parimutuel account of a conformal predictor's information gap [@cotton2026conformalbetting]. --
-
Multi-Stage Solicitation of Probability Distributions
abstract
The microprediction platform built distributional prediction supply chains, chaining pools so that one's output became the message or the settled outcome of the next. This note is a self-critical, ex-post look at the game theory of those chains: the games were built first, and here we ask, of each, whether truthful reporting was really the best move. A chain is proper where an exogenous outcome anchors every stage; a downstream stage can then be scored on its own residual or, for the log score, at the top level. By that test the games mostly hold up, one of them only by a lucky accident. Flaws aside, the setup has continuing pedagogical merit, as it illustrates the inherent limitation of conformal prediction. --
-
Likelihood versus CRPS: A New Perspective
abstract
The choice between the logarithmic score and the continuous ranked probability score (CRPS) is usually argued on four grounds: locality, propriety, robustness, and interpretability. This note adds a fifth, one the debate rarely weighs. When the elicited object is a reusable predictive density and forecasts are chained, each stage refining the residual left by the last, the logarithmic score is the natural default. The log-likelihood of a composed forecast is the base score plus the log-density of each residual refinement, so credit is additive and each stage's increment is itself the proper score of what that stage added. This is the prequential and logarithmic-market-scoring structure; CRPS shares no comparable per-stage identity. None of that is new mathematics: the decomposition is the probability integral transform and the prequential chain rule. The contribution is a framing. It makes composability the organizing axis of the log-versus-CRPS choice, which is conventionally settled on locality, propriety, robustness, and interpretability, and ties it to chained elicitation, the setting a prediction supply chain lives in. We give the compositional accounting, a population-level worked example in which the two scores prefer measurably different forecasts, and an honest account of where CRPS is the right target, its tail-insensitivity there often a virtue. Conformal prediction appears as the limiting case: a method that need not produce a density at all, scored on the one metric that does not notice. --
Market making & trading · inventory.microprediction.org/papers
-
On a Simple Relationship Between Order Imbalance, Skew and Width in Over-The-Counter Trading
abstract
We consider a market maker who can only obtain and dispose of inventory by responding to a sequence of sealed-bid enquiries, and whose customers arrive with imbalanced intent: sellers more often than buyers, or the reverse. Under the assumption that the best competing response is exponentially distributed around a commonly discerned fair price, we observe a symmetry in the steady state solution that compresses the imbalanced problem onto the perfectly balanced one. Order imbalance is absorbed, exactly, by a translation of the market maker's skew, a widening of her quotes, and a multiplication of her effective cost of carry. The adjustment is simple even though the solution it adjusts is not, and it involves no free parameter beyond the observable market width. The exponential assumption is needed only locally, at the quotes actually made, and the width that enters is the locally observed one. Among the consequences: a market maker with zero inventory should still skew; skew responds to imbalance at first order whereas width responds only at second order; and the popular ``constant width, linear skew'' heuristic is recovered as the small-skew solution in the special case of balanced flow and quadratic holding cost.
-
On the Relationship Between Accuracy and Profitability in Over-the-Counter Market Making
abstract
Intuitively, accuracy in microstructure prediction at a granular level must relate to profitability for a market participant, but this is not trivial to formalize. Here we provide an approximation using a stylized steadystate model for over-the-counter trading modeled as a sequence of sealed bid auctions. A simple picture emerges due to a mildly surprising feature of this model: one does not need to know the fair price, only where others are bidding.
-
Trading Illiquid Goods: Market Making as a Sequence of Sealed-Bid Auctions, with Analytic Results
abstract
We provide analytic results for the optimal control problem faced by a market maker who can only obtain and dispose of inventory via a sequence of sealed-bid auctions. Under the assumption that the best competing response is exponentially distributed around a commonly discerned fair market price we examine properties of the market maker’s optimal behavior. We show that simple adjustments to skew and width accommodate customer arrival imbalance. We derive a straightforward relationship between the market marker’s fill probability and direct holding costs. A simple formula for optimal bidding in terms of (non-myopic) inventory cost is presented. We present the results as a perturbation of an improvement to a “linear skew, constant width” (CWLS) market making heuristic.
Fixed income & stochastic volatility
-
Stochastic Volatility Corrections for Interest Rate Models — 2004.
-
Derivatives in Financial Markets with Stochastic Volatility — chapter, Cambridge University Press, 2000.
Epidemiology
-
Addressing the Herd Immunity Paradox Using Symmetry, Convexity Adjustments and Bond Prices — preprint, 2020.
abstract
In constant parameter compartmental models an early onset of herd immunity is at odds with estimates of R values from early stage growth. This paper utilizes a result from the theory of interest rate modeling, namely a bond pricing formula of Vasicek, and an approach inspired by a foundational result in statistics, de Finetti's Theorem, to show how the modeling discrepancy can be explained. Moreover the difference between predictions of classic constant parameter epidemiological models and those with variation and stochastic evolution can be reduced to simple "convexity" formulas. A novel feature of this approach is that we do not attempt to locate a true model but only a model that is equivalent after permutations. Convexity adjustments can also be used for cross sectional comparisons and we derive easy to use rules of thumb for estimating threshold infection level in one region given knowledge of threshold infection in another.
-
Repeat Contacts and the Spread of Disease: An Agent Model with Compartmental Solution — preprint, 2020.
abstract
Using a probability of novel encounter derived from a physical model, we augment the SIR compartmental model for disease spread. Scenarios with the same initial trajectories and identical $R_0$ values can diverge greatly depending on the speed at which our circles of acquaintances grow stale - leading to order of magnitude differences in final case counts. A momentum effect arises from variation in the mean time since infection, and this feeds back into new infection rate and faster decline in the late stages of an outbreak. Rapid extinction of an outbreak can occur in the early stages, but once this opportunity is missed the effect is diminished and then, only herd immunity can help.
Sports analytics · firstdown · firstdown.microprediction.org
-
Stop Shy of the First Down — in Sports Analytics, World Scientific, 2021.
Mathematical analysis
-
Contraction of an Adapted Functional Calculus
abstract
We aim to show, using the example of a Riemannian symmetric pair (G, K) = (SL2 (R), SO(2)), how contraction ideas may be applied to functional calculi constructed on coadjoint orbits of Lie groups. We construct such calculi on principal series orbits and generic orbits of the Cartan motion group V ⋊ K , and show how the two are related. Since the calculi are adapted to the representations traditionally attached to the orbits, we recover at the Lie algebra level the contraction results of Dooley and Rice [5].
Software · repos.microprediction.org
- precise — Online (incremental) covariance and correlation estimation — the online complement to sklearn.covariance. · code
- humpday — Derivative-free optimization in pure Python and JavaScript, including an ask/tell interface that takes one observation at a time. · code
- allocation — Streaming online portfolio construction, built on precise and thurstone. · code
- thurstone — Fast ability inference from contest winning probabilities. · winning
- skaters — Fast univariate time series models that run in Pyodide. · code
- prophet-laplace — Prophet in a laplace sandwich: same API, calibrated densities, via skaters.
- timemachines — Streaming anomaly detection with calibrated p-values, on skaters. · code
- mechanisms — Markets, scoring rules and everything in between. · code
Selected Talks
-
Schur Complementary Portfolios
abstract
Despite many attempts to make optimization-based portfolio construction in the spirit of Markowitz robust and approachable, it is far from universally adopted. Meanwhile, the collection of more heuristic divide-and-conquer approaches was revitalized by Lopez de Prado where Hierarchical Risk Parity (HRP) was introduced. This paper reveals the hidden connection between these seemingly disparate approaches.
-
Who Ya Gonna Call? A Solution to the Horse Race Problem with Application to OTC Markets
-
Trading Illiquid Goods
-
Filtering Bond and Credit Default Swap Markets
-
Barbell Bond Portfolios: What Do They Accidentally Optimize?
Videos
-
MLA@CSAIL Lecture Series: Constructing an Open Prediction Network, Powered by a Lottery Paradox
Interviews
-
The Future of AI: From Mathematics' Revenge to the Rise of Prompt Markets
-
Mastering Data Product Development: Insights from Peter Cotton
-
Portfolio Theory, Data Science and Entrepreneurship (Let's Talk AI, #28)
-
Supply Chain Lessons from Financial Markets (Ep 139)
Patents
-
System and Method for Providing Data Science as a Service
-
System and Method for Secure Causality Discovery
-
System and Method for Analyzing Financial Models with Probabilistic Networks
-
System and Method for Pricing Default Insurance