Already in a 2012 table
TL;DR
- Every whole number from 1 to 7.0·10^13 ends in one of the three known loops of the 3x − 1 problem, checked by at least two independently written programs. Others report or imply ranges 32 to about 13,000 times larger. This is a reproduction of a small corner.
- My headline result, that the loop through 17 collects the most numbers, was in a 2012 table. I found that out on day three.
- Not found in my limited search: a pre-registered test showing the lead is fixed among numbers below about 10,000. On 30 other rules, counting only small numbers predicted the shares at 10^15 to 10^16 to within about 0.01 points on average.
- Two patterns I can't place: a leading-digit "mode law" that passed three pre-registered tests with caveats, and a neighbour effect still growing at 2^55.
- We wrote our own GPU code and hand-optimised what the compiler missed. A rebuilt version ran 8.2 times faster on a test window but hasn't verified anything yet.
- The whole project used close to a full week of my Claude Max 20x plan. None of it is progress on the conjecture.
Every whole number from 1 to 70,000,000,000,000 falls into one of three loops in the 3x − 1 problem. The rule: halve the number if it's even, triple it and subtract 1 if it's odd, and repeat. The known loops are 1 → 2 → 1, the five numbers 5, 14, 7, 20, 10, and an 18-number loop through 17. Whether every positive number ends in one of them is an open problem. It is the Collatz conjecture asked about the negative integers, step for step.
That range is the safe part of this post. The part I got wrong was the result I had treated as the headline. Somebody had written it down in 2012.
I directed the work and set the rules (no uploads, no search engines). Claude, Anthropic's model, wrote every line of code and ran the computations through Claude Code on my PC; the model itself runs on Anthropic's servers. It also drafted the analysis and this post. It ran mainly as Opus 5.5 in Claude Code's ultracode mode, with a little Sonnet 5.5 and Fable 5.1 near the end, and separate Claude agents checked the work by recomputing results with their own code. No human mathematician has reviewed any of it. It was a 3 day or so detour on this mathematical conjecture: the main work ran from 29 September to 1 October 2026, with follow-up work on 2 and 4 October.
The finding that was already in a table
Call the loops A (through 1), B (through 5) and C (through 17). There is no obvious reason for one loop to collect more numbers than another. On the evening of day one we counted every x below 2^34, about 17.2 billion numbers, exactly. The next morning a second, independently written program matched every count.
| Loop | Berg & Opfer, x ≤ 10^8 (2012) | Ours, x < 2^34 |
|---|---|---|
| A | 32.697 % | 32.677 % |
| B | 32.471 % | 32.505 % |
| C | 34.832 % | 34.817 % |
On day two, about three hours of agent runs and 2.75 million subagent tokens went into explaining C's lead. On day three I asked whether any of it was new. A December 2012 Hamburg preprint by Berg and Opfer, "An analytic approach to the Collatz 3n+1 problem for negative start values with an appendix of tables", has exact counts to 10^8 in its Table 1.3, and our own count to 10^8 matches it number for number. They note that the third case "is a little more frequent than the other two cases." A September 2026 Zenodo preprint by M. Dominik reports the same lead from an exact census to 10^7, and our count to 10^7 matches that one number for number too.
Worse, Dominik's preprint was already in our day-two reading notes, where the agent that found it called it prior art for exactly this lead. By day three our own status summary still listed the shares as possibly new.
What survives is precision, about 172 times as many numbers as the 2012 table, and a partial answer to why. Much of C's surplus sits in one family of predecessors: 7.69 % of all numbers pass through 175, against the 1.37 % an average model expects, and about 30 % of the traffic through 175 has come down from 257 = 2^8 + 1, which first climbs eight odd steps to 6562 = 3^8 + 1. It isn't the only cause. The lead is a small difference between big pieces, and several branches are each necessary. My own question about "doors", the numbers through which orbits enter a loop, found one of them: without the door at 23, C would come last.
Then the law-or-luck test. If the lead is settled by how the small numbers are wired, counting small numbers should predict the shares at any size. We fixed the protocol and hashed it first, chose 30 rules of the form 3x + d that we had never looked at, and predicted each rule's shares from numbers below 2^14·d only. Two separately written samplers then measured the real shares at 10^15 to 10^16. The plain small-number count was off by 0.0115 points on average and named the winning loop for all 30 rules. Nothing here explains why the small numbers are wired the way they are. Dominik's preprint gives a different kind of answer: a backward-tree advantage for C, which has seven odd loop members against one and two, that the forward orbits mostly wash out. It names no particular families or doors and tests no other rules, so the two accounts sit side by side.
Anyone can re-derive the shares: follow each number until it repeats and note which loop it lands in. A short program does that for every number below 2^28 in about 70 seconds on my i7-7700K.
Everything, in one table, and what was borrowed
| Result | Status | Known before? |
|---|---|---|
| Every start from 1 to 7.0·10^13 ends in A, B or C | verified by two programs, AI-checked | yes, to far larger ranges |
| An unknown loop needs at least 55,245,847 steps | proved, given the range | method and table: Sinisalo 2003; Shimizu (2026, unrefereed) claims 357,638,239 |
| No unknown loop with 59 or fewer up-and-down blocks | argued, conditional; lemmas checked only by AI | implied by unrefereed claims: Cochin 61, Shimizu 63 |
| C collects the most: 34.817 % below 2^34 | exact count, two programs | yes, to 10^8 (2012) |
| The lead is fixed among small numbers (30 rules, 10^15 to 10^16) | pre-registered test, passed | not found |
| One predecessor family, through 175, holds much of C's surplus | verified by four programs (±0.01 points) | not found |
| Shares swing with leading digits, at m where 3^m is close to a power of 2 | mode law: heuristic, 3 tests passed with caveats | not found |
| Odd x and x + 1 share a loop far above chance, still rising at 2^55 | measured; limit unknown | not searched; Berg & Opfer show runs of consecutive numbers in one loop |
| x and 10x agree slightly above chance, fading with size | exact to 10^6, sampled to 2^50 | not searched |
| Loop lists for 3x ± d, d up to 999 (odd, not divisible by 3) | all 1,596 loops match published lists | yes |
| About 70 outside numbers and claims looked at | about 35 reproduced or matching our data, 5 contradicted | n/a |
"Not found" means not found in a limited search: a local browser, no search engines, zbMATH behind a bot check we didn't bypass, several papers paywalled or unread. It doesn't mean new. Two of the five contradicted claims said the shares are about equal thirds: an informal remark in one paper and the conclusion of a self-published one.
Most of the machinery is other people's. Checking each start only until it drops below itself is the standard method. Sieves on the last binary digits are in Barina's papers; the published forms are for 3x + 1, so ours is a 3x − 1 version with its own percentages. What we checked against the literature was Roosendaal's published count for the join sieve, 1,720 survivors out of 65,536 at sixteen binary digits, which our code reproduced exactly for 3x + 1 — and which for 3x − 1 removes nothing at all. The merge rules are Angeltveit's and Hercher's, and step tables are Oliveira e Silva's and Barina's. The loop-length method is Sinisalo's, and the m-cycle approach runs from Steiner through Simons and de Weger to Hercher and Cochin. My own ideas were mostly questions that turned into tools: "ending digits" became the binary sieve, "skip what we've seen" became the modulo-9 rule, the x + (x >> 1) step, the doors, and the 10x question below.
What I computed, and where it sits
Every part of the range up to 7.0·10^13 was covered by at least two independently written programs. Above 1.5·10^12 that meant the main GPU verifier and a deliberately simpler GPU checker agreeing on 20 summary numbers. In the newest part each file also had to contain exactly the number of starts an independent count predicted, and ten blocks of 10^10, picked by a seed fixed before the run, were recomputed on the CPU. Reviewer agents planted 11 bugs in the main verifier. Ten stopped the run, and the eleventh was provably harmless. A proven sieve on the last 24 binary digits meant only 1,195,725,599,972 starts, about 1.7 %, needed computing at all.
| Source | Range | Status |
|---|---|---|
| This work | 7.0·10^13 | two programs on every part, AI-checked only |
| Cochin, Zenodo, v1.1.1 (25 Sep 2026) | 2^51 ≈ 2.25·10^15 | unrefereed preprint; own CPU run to 2^44, one GPU sweep to 2^51 |
| Shimizu, Zenodo/GitHub, v0.4.0 (30 Sep 2026) | 4.52·10^15, just above 2^52 | unrefereed software release; single CPU runs above 2^50 |
| Angeltveit, arXiv, v1 (11 Feb 2026) | at least 9.29·10^17, my inference from its last path record (it states no range) | unrefereed preprint |
I checked none of those beyond the part that overlaps mine.
On one GTX 1070 the main GPU program covered 1.5·10^12 to 7.0·10^13 in about 2,040 seconds. The GPU was never the expensive part. Agents were: 32 workflow runs and 170 agent launches in the first three days, 41 runs and 230 launches by the end, which used close to a full week of my Claude Max 20x plan. One nine-agent run produced nothing. It launched right after I'd asked an unrelated question about agent limits, so each agent was told that question was the request behind the run, and all nine correctly declined.
The GPU code we wrote
None of the GPU code came from a library. Claude wrote the kernels that run on the card and the host program that feeds them and collects the results. It talks to the card through the driver's own OpenCL support and is compiled with the C# compiler that ships with Windows, so nothing was downloaded. The main verifier and the checker were written independently, so every range was computed twice by different code.
I suspected the GPU compiler wasn't making the obvious rewrites by itself, so we checked. Two agents read the compiler's intermediate output. It still multiplied by 3 where x + (x >> 1) gives (3x − 1)/2 for odd x, kept the odd/even branch, and couldn't invent a table jump. Hand-writing them paid off: 1.5 times for x + (x >> 1), 1.2 to 1.5 in a benchmark for dropping the branch (in the rebuilt checker it made things slower, so it was left out), and 2.0 to 3.0 for a two-phase loop that uses table jumps while they are provably safe. Bigger batches per GPU thread, with the results summed on the card, gave 3.6 times in a benchmark.
The rebuilt program also checks fewer numbers. A skip rule modulo 9 came from my idea of skipping numbers we had already seen: drop a start when a smaller start provably passes through it, which removes 44 % of the odd starts. Others use rules of this kind too: Shimizu's verifier skips modulo 3^10. Claude combined ours with merge rules from Angeltveit's and Hercher's papers. Rates below count every number in the range, including the 98.3 % the sieve skips.
| What ran | Range | Numbers per second |
|---|---|---|
| My first verifier, JavaScript, 1 CPU core | 5.5·10^11 to 10^12 | 5.3·10^7 |
| Main GPU program | same | 3.6·10^10 (668×) |
| Main program, then the independent checker | same | 8.3·10^9 (155×) |
| Main GPU program | 10^14 to 1.01·10^14 | 3.5·10^10 |
| Rebuilt main program, never used for a verified range | same | 2.9·10^11 (8.2×) |
| Independent checker | same | 1.2·10^10 |
| Checker on the shortened start list, 64-bit fast path | same | 5.4·10^10 (4.6×) |
The gains don't multiply, because the rebuilt program contains most of them. The 8.2 times is GPU time on a window the old program needs about 28 seconds for. Wall-clock it is 5.8 times, and near 2^50 the GPU-time gain falls to 6.5. The rebuilt code also runs hot. Five runs in ten minutes, with only short cool-downs between them, took the card to 73 °C, two degrees under our stop limit, so it never runs back to back. Pushing the verified range to 2^50 would take about 37 hours with the old programs and about 8 with the rebuilt ones. That is a model, and I haven't run it, on this card or on newer hardware.
What I can call proved
Not much, and none of it touches the hard part.
If a fourth loop exists, every member is above 7.0·10^13, which forces it to be at least 55,245,847 steps long, 21,372,011 of them odd. That is a standard consequence of the verification. The method and the step counts are in a 2003 preprint by Sinisalo (Lemma 6 and Table 2). Only the verified bound is mine, and Shimizu's note claims 357,638,239 from its larger range.
The second result I can only call argued. It is conditional. Count "triple, subtract 1, halve" as one up-step, and split a loop into blocks, each a run of up-steps followed by a run of halvings. Apart from the trivial loop A, only B and C exist among loops with 59 or fewer blocks. It leans on the verified range, on a theorem about how close a power of 2 can come to a power of 3 (a lower bound for linear forms in logarithms, Laurent 2008, Corollary 2, which we checked against the original), and on extra lemmas the AI derived for this map from an idea Hercher used for 3x + 1. Only AI checkers have read those lemmas, and they may not be new. Two programs with different refinements both reach 59, and the deciding cases cleared by only 0.005 and 0.006 bit. It isn't new as a statement either: a 2026 preprint by Cochin claims 61, resting on the author's own run to 2^51, and a 30 September note by Shimizu claims 63 from the author's own range of about 4.5·10^15. Neither is refereed. We reproduced Cochin's method and all 30 numbers in its two tables, but not the run, and we haven't checked Shimizu's note.
The patterns I can't place
Split the numbers with 34 binary digits (2^33 to 2^34) into 64 narrow windows by their leading binary digits, and C's share runs from 30.4 % to 38.6 %. In 27 of the 64 windows C isn't the largest. Treat the shares as a wave over the position of x inside its doubling range. The wave is made of a few harmonics, at the m for which 3^m is unusually close to a power of 2: 12, 41, 53 and so on. It's the same kind of near-coincidence that limits which loop shapes can exist, such as 3^12 = 531,441 just above 2^19 = 524,288.
Claude built a "mode law" that predicts how each harmonic shrinks and shifts every time the numbers double, with nothing fitted in that step. We pre-registered tests by hashing the predictions before any results were opened. The first was weak. The second measured 16 windows from about 6·10^13 to 2.3·10^18 and passed with χ² 23.1 on 32 degrees of freedom, well under the 1 % threshold of 53.5. A value that low means the error bars were generous, and they were calibrated on data the model had already seen. Against sampling noise alone it scores 219.8.
The third test grew out of a question of mine: does 10x end in the same loop as x? For x up to a million it does 34.9 % of the time, against 33.4 % by chance. Before the counts were opened, the mode law predicted that 7x, 11x and 13x would agree with x less often than chance, and at 2^24 and 2^25 all three did, though at those sizes the prediction is built from the same table it is tested on. It over-predicted the 10x effect, which fades to about 0.1 points between 2^36 and 2^50. And the registration failed open: the file meant to show the predictions were locked before measuring was never written, and only hashes the measuring agent had saved proved it. The law stays a heuristic.
The neighbour pattern is the one I find hardest to leave. For odd x, x and x + 1 end in the same loop far more often than chance. An exact case analysis, proved and machine-checked between 2^24 and 2^30, leaves one hard case: x one less than a multiple of 8. Pairs with x three more than a multiple of 8 always agree, and the remaining cases track the hard one closely, though that part is heuristic. The hard pair agrees 52.0 % of the time around 2^34, 55.4 % around 2^45 and 57.8 % around 2^55, against 33.4 % by chance, from 2^23 random pairs per doubling. The rise is slowing, from 0.6 points per doubling near 2^21 to 0.2 near 2^55. None of the six curves registered before measuring stayed within its registered tolerance. The two closest bracket the data from 2^45 on: one levels off near 62 %, the other never does. The measuring agent could see the curves' values, so it wasn't fully blind.
What this does not show
It doesn't show the conjecture, or progress toward it. The problem has two halves: no other loops, and no orbit that climbs forever. The verification stops at 7.0·10^13 and the block result at 59. Nothing here touches runaway orbits. The percentages are finite counts, and I know of no proof that limits exist. Apart from where our numbers match published tables, every check was AI checking AI.
I'm archiving the project with work left over. A script counts the project's task list, and these are the parts I might come back to one day:
| Left on the backlog | Items |
|---|---|
| Experiments stopped half-way | 8 |
| Long jobs, such as the run to 2^50 | 2 |
| Research leads for another day | 69 |
I don't know why the small numbers are wired in C's favour. I don't know where the neighbour curve ends. The mode law has passed the three tests we set it, the third with a miss in size. Nobody else has set it one yet.