Certified total-variation measurement of complex path-space measures for real-time quartic dynamics. Numbers on this sheet are published as measured, with their certification status; it is updated as the survey advances. Last updated 2026-09-03.
The Feynman–Kac formula rewrites imaginary-time quantum mechanics as an integral over paths against a genuine (countably additive) measure. Whether real-time quantum mechanics admits any such representation has remained open since Cameron's 1960 divergence result for the natural analytic continuation. Avelocn studies the sharpest quantitative form of the question for the anharmonic Hamiltonian H = −½ d²/dx² + 4x⁴: what is the total variation (TV) of an explicit complex path-space representation — and how does it scale?
Every causal complex representation we can parametrize diverges: measured causal TV floors climb from log₁₀TV ≈ 1 at n = 2 time slices to ≈ 16 by n = 16 — sixteen orders of magnitude and accelerating. An explicitly constructed acausal object — a cycle deformed along the antiholomorphic gradient flow (a Lefschetz-thimble approximant) — carries log₁₀TV ≈ 0.03–0.11 over the same range, flat in n. The complex case is not closed by these measurements, but it is cornered: bounded total variation and causal path structure appear to be mutually exclusive.
Formal status: finite-n absolutely convergent acausal representations are proved (Result A). The quantitative causal no-go and the continuum limit of the cycle remain conjectures, stated precisely in the write-up. Robustness: the picture is unchanged under 8× in the coupling and 4× in physical time. Toward the no-go: an adversarial causal optimizer — free to choose each time-slice's contour knowing the previous — cannot beat the simplest product family (a joint search converges back onto it within 0.02 decades; a greedy one lands fifteen decades worse). Causality's freedom buys nothing, because consecutive slices tax any mismatch. That is the mechanism the proof campaign is now built around.
The construction extends to D coupled quartic sites (real-time lattice φ⁴ on an open chain). Protocol held fixed: λ=4, t=0.4, flow depth τ=0.15, W=32768 walkers; certification by a pre-registered ladder — exact factorization referee at κ=0 (valid at any D), grid referee at D ≤ 3, replicate gating with thresholds frozen 2026-08-04, before referee comparisons were seen. Failure falsifies; passing is evidence, not proof.
| D | coupling | log₁₀TV | status |
|---|---|---|---|
| 1 | κ=0 | 0.108 ± 0.003 | certified 1.2% against the factorization referee |
| 2 | κ=1 | 0.225 ± 0.008 | certified 0.6% against the grid referee |
| 3 | κ=1 | 0.329 ± 0.004 | certified K=10 pool across two engine generations |
| 4 | κ=0 | 0.427 ± 0.002 | certified double-anchored (referee + structural additivity) |
| 4 | κ=1 | 0.418 ± 0.011 | certified gated K=6 — first certified coupled point beyond any referee |
| 5 | κ=0 | 0.526 ± 0.013 | refereed clean band, 5 of 12 replicates; best replicate 4.6% |
| 5 | κ=1 | 0.476 ± 0.017 | screened 12 of 16 drawn; matched-screen κ=0 gives 0.504 ± 0.017 |
| 6 | κ=0 | 0.657 ± 0.004 | certified three certified replicates (4.1–5.4%) across two exchange kernels; structural expectation 0.654; depth-staggered kernel certifies at 2 of 4 |
| 7 | κ=0 | 0.748 | certified 1 of 16 replicates certified (5.5%; best of the depth); structural expectation 0.763 — certification is rare but real at seven |
κ-independence: TV(κ=1) ≈ TV(κ=0) at every depth ever measured, D = 2–7 (D=5: 1.2σ matched instruments; D=6: 0.2σ matched ungated; D=7: matched medians, rank-indistinguishable — medians because a catastrophe-bearing population poisons means; each comparison's instrument disclosed). Methodological rule the survey enforced on itself: a screened mean is its own observable, biased low by undetected contamination — it may only be compared like-instrument to like-instrument, never against a referee-purified anchor.
Certified rate along the depth ray at fixed W=32768: 33% → 37% → 8% → 0 for D = 3/4/5/6.
The verification wall (D=5). Computation still succeeds — clean replicates exist and agree with the factorization referee — but referee-free discriminators drop to roughly half power, and the failure mode flips sign with depth: shallow-depth contamination reads high with steep late-run slopes (migration lag); deep contamination reads low with flat slopes (mode locking). Verification degrades before computation does.
The computation wall (D=6 at W=32768). All 8 refereed replicates discard (value errors 13.5%–231%). The failure is collective: every replicate's phase coherence sits at 0.537–0.570, below the 0.60 line that flagged individual failures at D=5, and 7 of 8 TV values undershoot the structural expectation 6·TV₁ = 0.654 while the eighth is a locked outlier at 1.180. The population mean (0.622) imitates health only by cancellation — a standing caution for any unrefereed protocol.
Walker escalation buys one halving, then saturates. Doubling the ensemble (W=65536) softened the failure in every measured coordinate — median value error 38% → 19%, the catastrophic outlier class vanished, phase coherence rose — a response consistent with error ∝ 1/W. That law was promoted to a pre-registered prediction before the next run launched: at W=131072 it required median error 8–12%. Measured, four replicates: median ~30% (23.8–51.5%), marginally worse than W=65536, with the catastrophic-failure class returning. The prediction failed decisively, and the failure is the finding: the walker axis has an optimum (~65536 at D=6) — beyond it, more walkers sampling the same dynamics actively hurt. The obstruction at depth is dynamical — mode locking — not statistical. Verification walls, then computation walls; and the computation wall, on inspection, is a locking wall.
The lock has a curable component. The budget-controlled successor ran: the same walker total arranged as exchanging sub-populations. At identical pool size, exchange cut the median error from ~50% to 18% — and tied the monolithic optimum (19%) at equal budget with quarter-size pools. The diagnostics say why: without exchange, pools fail coherently to one side; with it, they scatter around the truth and pooling cancels the error. What exchange did not move is phase coherence — the same ceiling as every monolith. The wall's remaining core is phase decoherence, not population structure. Certification has not yet returned.
Every road leads to the same floor. The scaling test resolved the picture: exchanging pools at the walker count where the monolith degrades removed the degradation only directionally, and every well-configured architecture — monolith at its optimum, four exchanging pools, eight exchanging pools at doubled budget — converges to the same ~19% median error at the same phase-coherence ceiling. Exchange cures the lineage-lock component (~50% → ~19%); what remains is a floor set by phase decoherence that walkers, budget, and matched-depth topology are all now measured to be exhausted against. One in four exchanging replicates punches through to the provisional band (best: 8.9% — the first sub-12% replicate at this depth in any configuration), stochastically, by two-sided cancellation. The floor is the wall's true face — and the tail now reaches through it: with more replicates, the exchange architecture (named Avelocnex) produced the first certified replicate at this depth — 5.1% referee error, landing on the structural expectation — at the same coherence ceiling as every failure. Certification returned before coherence was cured: stochastic (1 of 8), but real, and referee-verified.
The lottery became a system. The depth-staggered variant — sub-populations half a step apart in anneal depth, trading walkers by tempering moves so stuck ensembles step back and re-anneal — certifies at 2 of 4 replicates with median error 7.0% (best 4.1%, the deepest certified measurement of the program), against the matched-depth architecture's 17.9% median and 1-of-12 certification lottery. The decisive diagnostic: the phase ceiling STILL did not move. The staggered kernel does not cure decoherence — it decorrelates the sub-populations' errors so thoroughly that pooled cancellation becomes reliable rather than lucky. Certification at this depth is now systematic, at the same coherence every failure ever had.
The next depth, opened. Carried unchanged to D=7, the depth-staggered architecture posts a 22% median with the first provisional replicate at that depth (11.3%) — already slightly better than the monolith ever managed at D=6 — while the coherence ceiling steps down again (~0.52). The depth exchange rate at fixed architecture is now measured: roughly 2.5× the median error per depth step. The wall re-forms one rung deeper, on schedule, with a foothold already in it.
The stagger has an optimum too. Widening the depth stagger to a full rung at D=7 made things worse (median 31% vs 22%, with a catastrophic replicate): like walker count before it, the tempering distance peaks at or below half a rung, and the flow-time axis is spent at this depth and budget. What the measurements point at instead is the coherence floor itself, which steps down with every depth (0.63 → 0.57 → 0.52) — so certification by pooled cancellation should cost more pools per depth. That cost curve, pools-per-depth, is the walker ray's successor and the program's next measurement.
Depth seven, certified. The pools-per-depth measurement closed with a shape nobody ordered: medians flat across four, six, and eight pools (22% → 17% → 18%) while the tails blow open — the eight-pool population holds both the program's best depth-seven replicate, 5.5%, certified (0.748 against structural 0.763), and its worst catastrophe. Pools scale the lottery's stakes, not its center; certification at this depth is stochastic, exactly as it first was at depth six before the depth-staggered kernel systematized it. The pattern asks for the next kernel idea, not more pools.
in progress The depth ray continues: a first look at D=8 under the best configuration — and, in parallel, the mathematics: a pairwise lower-bound campaign whose first theorem target would prove the causal/acausal separation the measurements keep exhibiting.
The numbers on this sheet are produced by Avelocn, a certified sampling engine whose construction is deliberately not described here; the write-up will accompany a later release. What is published now is its output and its performance envelope: the current generation (Avelocn 1.4) runs 3.1× its predecessor at M=8 rising to ≈10× at M=32 — the advantage grows with mode count — and cumulatively ≈38× at D=3 over the first production pipeline. Every number on this sheet was produced on a single consumer GPU (a 12 GB RTX 3060).
Context in the public literature for sampling along holomorphic gradient flow: arXiv:1703.00861 (PTEP 2017), arXiv:2012.08468. The total-variation question addressed here is not posed in that line of work.