Vue normale

Il y a de nouveaux articles disponibles, cliquez pour rafraîchir la page.
À partir d’avant-hierFlux principal

Ruliology of the “Forgotten” Code 10

1 juin 2024 à 17:21

My All-Time Favorite Science Discovery

June 1, 1984—forty years ago today—is when it would be fair to say I made my all-time favorite science discovery. Like with basically all significant science discoveries (despite the way histories often present them) it didn’t happen without several long years of buildup. But June 1, 1984, was when I finally had my “aha” moment—even though in retrospect the discovery had actually been hiding in plain sight for more than two years.

My diary from 1984 has a cryptic note that shows what happened on June 1, 1984:

Ruliology of the "Forgotten" Code 10

There’s a part that says “BA 9 pm → LDN”, recording the fact that at 9pm that day I took a (British Airways) flight to London (from New York; I lived in Princeton at that time). “Sent vega monitor → SUN” indicates that I had sent the broken display of a computer I called “vega” to Sun Microsystems. But what’s important for our purposes here is the little “side” note:
Take C10 pict.
R30
R110

What did that mean? C10, R30 and R110 were my shorthand designations for particular, very simple programs of types I’d been studying: “code 10”, “rule 30” and “rule 110”. And my note reminded me that I wanted to take pictures of those programs with me that evening, making them on the laser printer I’d just got (laser printers were rare and expensive devices at the time).

I’d actually made (and even published) pictures of all these programs before, but at least for rule 30 and rule 110 those pictures were very low resolution:

Click to enlarge

But on June 1, 1984, my picture was much better:

Click to enlarge

For several years I’d been studying the question of “where complexity comes from”, for example in nature. I’d realized there was something very computational about it (and that had even led me to the concept of computational irreducibility—a term I coined just a few days before June 1, 1984). But somehow I had imagined that “true complexity” must come from something already complex or at least random. Yet here in this picture, plain as anything, complexity was just being “created”, basically from nothing. And all it took was following a very simple rule, starting from a single black cell.

Our usual intuition that to make something complex required “complex effort” was, I realized, simply wrong. In the computational universe one needed a new intuition. And the picture of rule 30 I generated that day was what finally made me understand that. Still, although I hadn’t internalized it before, several years of work had prepared me for this. And just days later I was at a conference already talking confidently about the implications of what I’d seen in rule 30.

Over the years that followed, rule 30 became basically the face of the phenomenon I had discovered. By 1985 I devoted a whole paper to it, in A New Kind of Science it was my initial and quintessential example, for the past quarter century a picture of rule 30 has adorned my personal business cards, and in 2019 we launched the Rule 30 Prizes to promote the rich basic science of rule 30:

Rule 30 Prize

But what about “C10”—the first item in my cryptic note? What was that? And what became of it?

First Sightings of Code 10

Well, C10 was “code 10”, or, more fully, “k = 2, r = 2 totalistic code 10 cellular automaton”. (I used the term “code” as a way to indicate a totalistic, rather than general, “rule”.) And, actually, I had looked at code 10 several times before, never really paying much attention to it.

The first explicit mention I find in my archives is from February 1983 (apparently reporting on something I’d done in January of that year). I had been doing all sorts of computer experiments on cellular automata, recording the results in a lab notebook. One page has observations about what I then called “summational rules” (I would soon rename these “totalistic”). And there’s code 10:

Click to enlarge

Mostly I had been studying the behavior starting from random initial conditions, but for code 10 I noted: “very irregular, even from simple initial state”. Within a couple of months I had even made (on an electrostatic printer) a high-resolution picture of code 10 starting from a single black cell—and here it is, prepared for publication, Scotch tape and all:

Click to enlarge

It appeared in a paper I wrote in May 1983. But the paper (entitled “Universality and Complexity in Cellular Automata”) was mostly about other things (for example, introducing my four general classes of cellular automaton behavior and talking quite a lot about code 20 as an example of class 4 rule), and it contained only a passing comment about code 10:

Click to enlarge

Code 10 is a range-2 rule, which means that the patterns it generates can grow by 2 cells on each side at each step. And the result is that the patterns quickly get quite wide, so that if one cuts them off when they “hit the edge of the page” (as my early programs “conveniently” tended to do) they don’t go very far, and one doesn’t get to see much of code 10’s behavior.

And it was this piece of “ergonomics” that caused me to basically ignore code 10—and not to recognize the “rule 30 phenomenon” until I happened to produce that high-resolution image of rule 30 on June 1, 1984.

I didn’t entirely forget code 10, for example mentioning it in a note to “Why These Discoveries Were Not Made Before” in my 2002 book A New Kind of Science:

Click to enlarge

But now that forty years have passed since I made—and basically ignored—that “C10” picture, I thought it would be nice to go back and see what I missed, and to use our modern Wolfram Language tools to spend a few hours checking out the story of code 10.

It’s an exercise in what I now call “ruliology”—the basic science of studying what simple rules do. And whenever one does ruliology there are certain standard things one can look at—that I showed many examples of in A New Kind of Science. But in a quintessential reflection of computational irreducibility there are also always “surprises”, and special phenomena one did not expect. And so it is with code 10.

Code 10: The Basic Story

Note: Click any diagram to get Wolfram Language code to reproduce it.

Code 10 is a cellular automaton operating on a line of black and white cells, at each step adding up the values of the 5 cells up to distance 2 from any given cell (black is 1, white is 0). If the total is 1 or 3, the cell is black on the next step; otherwise it’s white (in base 2 the number 10 is 001010):

And, yes, in many ways this rule is even simpler to describe—at least in words—than rule 30. And if one thinks of it in terms of Boolean expressions, it can also be written in a very simple form:

(By the way, as a general k = 2, r = 2 rule, code 10 is rule 376007062.)

So what does code 10 do? Here are a few steps of its evolution starting from a single black cell:

And here are 2000 steps:

And, yes, even though it’s a simple rule, its behavior looks highly complex, and in many ways quite random. One immediate observation is that—unlike rule 30—code 10 is symmetric, so the pattern it generates is left-right symmetric. The center column isn’t interesting: after having black cells for 2 steps, it’s white thereafter. (And by substituting values yx0xy into the Boolean expression above, it’s easy to prove this.)

Filling the white region around the center column with red we get:

There doesn’t seem to be any long-range regularity to the way the width of this region changes:

And indeed the (even) widths seem at least close to exponentially distributed:

What if one goes one column to the left or right of the center? Here’s the beginning of the sequence one gets:

And, yes, every other cell is white. Picking only “even-numbered positions” we get:

Looking at the accumulated mean for 100,000 steps suggests that this sequence isn’t “uniformly random”, and that slightly fewer than 50% of the cells end up being black:

Going away from the center line, every other column has white cells every two steps. Sampling the pattern only at “odd positions” in both “space and time” we get a pattern that looks similar—though not identical—to our original one:

Looking at every cell, the overall density of the pattern seems to approach about 0.361. Looking only at “odd positions” the overall density seems to be about 0.49. And, yes, the fact that it doesn’t seem to become exactly 1/2 is one of those typical “not-quite-as-expected” things that one routinely finds in doing ruliology.

There are some aspects of the code 10 pattern, though, that inevitably work in particular ways. For example, if we “rotate” the pattern so that its boundary is vertical, we can see that close to the boundary the pattern is periodic:

The period progressively doubles at depths separated by 1, 1, 4, 6, 8, 14, 124, …—yielding what may perhaps ultimately be a logarithmic growth of period with depth:

Other Initial Conditions, and a Surprise

We’ve looked at what happens with an initial condition consisting of a single black cell. But what about other initial conditions? Here are a few examples:

We might have thought that the “strength of randomness” would be large enough that we’d get patterns that look basically the same in all cases. But so what’s going on in the case? Running twice and five times as long reveals it’s actually nothing special; there just happen to be a few large triangles near the top:

So will nothing else notable happen with larger initial conditions?

What about ? Let’s run that a little longer:

And OMG! It’s not random and unpredictable at all. It’s a nested pattern!

Even in the midst of all that randomness and computational irreducibility, here is a dash of computational reducibility—and a reminder that there are always pockets of reducibility to be found in any computationally irreducible system, though there’s no guarantee how difficult they will be to find in any given case.

The particular nested pattern we get here is a bit like the one from the additive elementary rule 150, that simply computes the total mod 2 of the three cells in each neighborhood:

And it turns out to be almost exactly the r = 2 analog of this—the additive rule (code 42) that takes the total mod 2 of the five cells in the r = 2 neighborhood:

The limiting fractal dimension of this pattern is:

Is unique, or does this same phenomenon happen with other “seeds”? It turns out to happen again for:

So what’s going on here? Comparing the detailed pattern in the code 10 case with the additive rule case, there’s no immediate obvious correspondence:

But if we look at the rules for code 10 and code 42 respectively:

We notice that there’s really only one difference. In code 10, gives while in code 42, it gives . In other words, if code 10 avoids ever generating any block, it will inevitably behave just like code 42—and shows nested before. And that’s what happens for the initial conditions above; they can for example lead to blocks, but never .

Another notable and at first unexpected phenomenon concerns the overall density of black cells in patterns from different initial conditions:

And what we find is that for even-length initial blocks the density is about 0.47, while for odd ones it’s about 0.36. At first it might seem very strange that something as global as overall density could be affected by the initial conditions. But once again, it’s a story of what blocks can occur: in the odd-length case, there’s a checkerboard of guaranteed-white cells, which just doesn’t exist in the even-length case.

Other Things to Study

We’ve been looking at what code 10 does with specific, simple initial conditions. What about with random initial conditions? Well, it’s not terribly exciting. It basically just looks random all the way through—which, by the way, is part of the reason I didn’t pay much attention to code 10 back in 1983:

But even though this looks quite random, it’s for example not the case that every single possible block of values can occur. Though it’s very close. Let’s say we start from all possible sequences of 0s and 1s in the initial conditions. Then—using methods I developed in 1984 based on finite automata—it’s possible to determine that even after 1 step there are some blocks of values that can’t occur. But it turns out that one has to go all the way to blocks of length 36 before one finds an example:

Although the patterns generated by code 10 generally look quite random, if we look closely we can see at least patches that are fairly regular. The most obvious examples are white triangles. But there are other examples, most notably associated with regions consisting of repetitions of blocks with periodic behavior:

Complementary to this is the question of what code 10 does in regions of limited size—say with cyclic boundary conditions, starting from a single black cell. The result is quite different for regions of different sizes:

For a region of size n, a symmetric rule like code 10 must repeat with a period of at most 2n/2. Here are the actual repetition periods as a function of size, shown on a log plot:

These are results specifically for the single-cell initial condition. We can also generate state transition diagrams for all 2n possible states in a size n code 10 system:

And mostly what we see is highly contractive behavior, with many different initial states evolving to the same final state—even though “eventually” we should start seeing larger cycles of the kind we picked up above when we looked at evolution from a single-cell initial condition.

And, yes, I could go on, for example repeating analyses I’ve done in the past for rule 30. A lot of what we’d see would be at least qualitatively much the same as for rule 30—and essentially the result of the appearance of computational irreducibility in both cases. But it’s a feature of the computational universe—and indeed one of the many consequences of computational irreducibility—that different computational systems will inevitably have different “idiosyncrasies”. And so it is for rule 30 and code 10. Rule 30 has an “Xor on one side” which gives it special surjectivity properties. Code 10 on the other hand has its block emulations, which lead, for example, to the surprise of nesting.

I’ve now spent many years studying the ruliology of simple programs, and if there’s one thing that still amazes me after all that time it’s that there are always surprises. Even with very simple underlying rules one can never be sure what will happen; there’s no choice but to just do the experiments and see. And, in my experience, pretty much whenever one thinks one’s “got to the end” and “seen everything there is to see”, something completely unexpected will pop out—a reminder that, as the Principle of Computational Equivalence tells us, these simple computational systems are in some sense a microcosm of everything that’s possible.

Ruliology is in many ways the ultimate foundational science—a science concerned with pure abstract rules not set up with any particular reference either to nature or to human choice. In a sense ruliology is our best path to ultimate pure abstraction—and unfettered exploration of the ruliad. And at least for me, it’s also something very satisfying to do. These days, with modern Wolfram Language, it’s all very streamlined and fast. Sitting at one’s computer, one can immediately start visiting vast uncharted areas of the computational universe, seeing things—and often very beautiful things—that have never been seen before, and discovering new but everlasting things anchored in the bedrock of computation and of simple programs.

It’s been fun spending a few hours studying the ruliology of code 10. Essentially everything I’ve done here I could have done (though not nearly as efficiently) back in 1983 when I first came up with code 10. But as it was, code 10 in a sense had to “wait patiently” for someone to come and look at it. The form of the rule 30 pattern is in some ways more “human-scaled” than code 10. But, as we’ve seen here, code 10 still manifests the same core phenomenon as rule 30. And now, forty years after printing that “C10” picture, I’m happy to be able to say that I think I’ve finally gotten at least a passing acquaintance with another remarkable “computational world” out there in the computational universe: the world of code 10.

When Exactly Will the Eclipse Happen? A Multimillennium Tale of Computation

29 mars 2024 à 19:32

When Exactly Will the Eclipse Happen? A Multimillennium Tale of Computation

Updated and expanded from a post for the eclipse of August 21, 2017.

IWhen Exactly Will the Eclipse Happen? A Multimillennium Tale of Computation

Preparing for April 8, 2024

On April 8, 2024, there’s going to be a total eclipse of the Sun visible on a line across the US. But when exactly will the eclipse occur at a given location? Being able to predict astronomical events has historically been one of the great triumphs of exact science. But how well can it actually be done now?

The answer is well enough that even though the edge of totality moves at just over 1000 miles per hour, it’s possible to predict when it will arrive at a given location to within perhaps a second. And as a demonstration of this, for the total eclipse back in 2017 we created a website to let anyone enter their geo location (or address) and then immediately compute when the eclipse would reach them—as well as generate many pages of other information.

Click to enlarge

It’s an Old Business

These days it’s easy to find out when the next solar eclipse will be; indeed built right into the Wolfram Language there’s just a function that tells you (in this form the output is the “time of greatest eclipse”):

It’s also easy to find out, and plot, where the region of totality will be:

Or to determine that the whole area of totality (including lots of ocean and some of Canada) will be about a third of the area of the US:

But computing eclipses is not exactly a new business. In fact, the Antikythera device from 2000 years ago even tried to do it—using 37 metal gears to approximate the motion of the Sun and Moon (yes, with the Earth at the center). To me there’s something unsettling—and cautionary—about the fact that the Antikythera device stands as such a solitary piece of technology, forgotten but not surpassed for more than 1600 years.

But right there on the bottom of the device there’s an arm that moves around, and when it points to an Η or Σ marking, it indicates a possible Sun or Moon eclipse. The way of setting dates on the device is a bit funky (after all, the modern calendar wouldn’t be invented for another 1500 years), but if one takes the simulation on the Wolfram Demonstrations Project (which was calibrated back in 2012 when the Demonstration was created), and turns the crank to set the device for April 8, 2024, here’s what one gets:

Click to enlarge

Antikythera device in action

And, yes, all those gears move so as to line the Moon indicator up with the Sun—and to make the arm on the bottom point right at an H—just as it should for a solar eclipse. It’s amazing to see this computation successfully happen on a device designed 2000 years ago.

Of course the results are a lot more accurate today. Though, strangely, despite all the theoretical science that’s been done, the way we actually compute the position of the Sun and Moon is conceptually very much like the gears—and effectively epicycles—of the Antikythera device. It’s just that now we have the digital equivalent of hundreds of thousands of gears.

Why Do Eclipses Happen?

A total solar eclipse occurs when the Moon gets in front of the Sun from the point of view of a particular location on the Earth. And it so happens that at this point in the Earth’s history the Moon can just block the Sun because it has almost exactly the same angular diameter in the sky as the Sun (about 0.5° or 30 arc-minutes).

So when does the Moon get between the Sun and the Earth? Well, basically every time there’s a new moon (i.e. once every lunar month). But we know there isn’t an eclipse every month. So how come?

Well, actually, in the analogous situation of Ganymede and Jupiter, there is an eclipse every time Ganymede goes around Jupiter (which happens to be about once per week). Like the Earth, Jupiter’s orbit around the Sun lies in a particular plane (the “plane of the ecliptic”). And it turns out that Ganymede’s orbit around Jupiter also lies in essentially the same plane. So every time Ganymede reaches the “new moon” position (or, in official astronomy parlance, when it’s aligned “in syzygy”—pronounced sizz-ee-gee), it’s in the right place to cast its shadow onto Jupiter, and to eclipse the Sun wherever that shadow lands. (From Jupiter, Ganymede appears about 3 times the size of the Sun.)

But our Moon is different. Its orbit doesn’t lie in the plane of the ecliptic. Instead, it’s inclined at about 5°. (How it got that way is unknown, but it’s presumably related to how the Moon was formed.) But that 5° is what makes eclipses so comparatively rare: they can only happen when there’s a “new moon configuration” (syzygy) right at a time when the Moon’s orbit passes through the plane of the ecliptic.

To show what’s going on, let’s draw an exaggerated version of everything. Here’s the Moon going around the Earth, colored red whenever it’s close to the plane of the ecliptic:

Now let’s look at what happens over the course of about a year. We’re showing a dot for where the Moon is each day. And the dot is redder if the Moon is closer to the plane of the ecliptic that day. (Note that if this was drawn to scale, you’d barely be able to see the Moon’s orbit, and it wouldn’t ever seem to go backwards like it does here.)

Now we can start to see how eclipses work. The basic point is that there’s a solar eclipse whenever the Moon is both positioned between the Earth and the Sun, and it’s in the plane of the ecliptic. In the picture, those two conditions correspond to the Moon being as far as possible towards the center, and as red as possible. So far we’re only showing the position of the (exaggerated) Moon once per day. But to make things clearer, let’s show it four times a day—and now prune out cases where the Moon isn’t at least roughly lined up with the Sun:

And now we can see that at least in this particular case, there are two points (indicated by arrows) where the Moon is lined up and in the plane of the ecliptic (so shown in red)—and these points will then correspond to solar eclipses.

In different years, the picture will look slightly different, essentially because the Moon is starting at a different place in its orbit at the beginning of the year. Here are schematic pictures for a few successive years:

It’s not so easy to see exactly when eclipses occur here—and it’s also not possible to tell which are total eclipses where the Moon is exactly lined up, and which are only partial eclipses. But there’s at least an indication, for example, that there are “eclipse seasons” in different parts of the year where eclipses happen.

OK, so what does the real data look like? Here’s a plot for 20 years in the past and 20 years in the future, showing the actual days in each year when total and partial solar eclipses occur (the small dots everywhere indicate new moons):

The reason for the “drift” between successive years is just that the lunar month (29.53 days) doesn’t line up with the year, so the Moon doesn’t go through a whole number of orbits in the course of a year, with the result that at the beginning of a new year, the Moon is in a different phase. But as the picture makes clear, there’s quite a lot of regularity in the general times at which eclipses occur—and for example there are usually 2 eclipses in a given year—though there can be more (and in 0.2% of years there can be as many as 5, as there last were in 1935).

To see more detail about eclipses, let’s plot the time differences (in fractional years) between all successive solar eclipses for 100 years in the past and 100 years in the future:

And now let’s plot the same time differences, but just for total solar eclipses:

There’s obviously a fair amount of overall regularity here, but there are also lots of little fine structure and irregularities. And being able to correctly predict all these details has basically taken science the better part of a few thousand years.

Ancient History

It’s hard not to notice an eclipse, and presumably even from the earliest times people did. But were eclipses just reflections—or omens—associated with random goings-on in the heavens, perhaps in some kind of soap opera among the gods? Or were they things that could somehow be predicted?

A few thousand years ago, it wouldn’t have been clear what people like astrologers could conceivably predict. When will the Moon be at a certain place in the sky? Will it rain tomorrow? What will the price of barley be? Who will win a battle? Even now, we’re not sure how predictable all of these are. But the one clear case where prediction and exact science have triumphed is astronomy.

At least as far as the Western tradition is concerned, it all seems to have started in ancient Babylon—where for many hundreds of years, careful observations were made, and, in keeping with the ways of that civilization, detailed records were kept. And even today we still have thousands of daily official diary entries written in what look like tiny chicken scratches preserved on little clay tablets. “Night of the 14th: Cold north wind. Moon was in front of α Leonis. From 15th to 20th river rose 1/2 cubit. Barley was 1 kur 5 siit. 25th, last part of night, moon was 1 cubit 8 fingers behind ε Leonis. 28th, 74° after sunrise, solar eclipse…”

Click to enlarge

If one looks at what happens on a particular day, one probably can’t tell much. But by putting observations together over years or even hundreds of years, it’s possible to see all sorts of repetitions and regularities. And back in Babylonian times the idea arose of using these to construct an ephemeris—a systematic table that said where a particular heavenly body such as the Moon was expected to be at any particular time.

(Needless to say, reconstructing Babylonian astronomy is a complicated exercise in decoding what’s by now basically an alien culture. A key figure in this effort was a certain Otto Neugebauer, who happened to work down the hall from me at the Institute for Advanced Study in Princeton in the early 1980s. I would see him almost every day—a quiet white-haired chap, with a twinkle in his eye—and just sometimes I’d glimpse his huge filing system of index cards which I now realize was at the center of understanding Babylonian astronomy.)

One thing the Babylonians did was to measure surprisingly accurately the repetition period for the phases of the Moon—the so-called synodic month (or “lunation period”) of about 29.53 days. And they noticed that 235 synodic months was very close to 19 years—so that about every 19 years, dates and phases of the Moon repeat their alignment, forming a so-called Metonic cycle (named after Meton of Athens, who described it in 432 BC).

It probably helps that the random constellations in the sky form a good pattern against which to measure the precise position of the Moon (it reminds me of the modern fashion of wearing fractals to make motion capture for movies easier). But the Babylonians noticed all sorts of details of the motion of the Moon. They knew about its “anomaly”: its periodic speeding up and slowing down in the sky (now known to be a consequence of its slightly elliptical orbit). And they measured the average period of this—the so-called anomalistic month—to be about 27.55 days. They also noticed that the Moon went above and below the plane of the ecliptic (now known to be because of the inclination of its orbit)—with an average period (the so-called draconic month) that they measured as about 27.21 days.

And by 400 BC they’d noticed that every so-called saros of about 18 years 11 days all these different periods essentially line up (223 synodic months, 239 anomalistic months and 242 draconic months)—with the result that the Moon ends up at about the same position relative to the Sun. And this means that if there was an eclipse at one saros, then one can make the prediction that there’s going to be an eclipse at the next saros too.

When one’s absolutely precise about it, there are all sorts of effects that prevent precise repetition at each saros. But over timescales of more than 1300 years, there are in fact still strings of eclipses separated from each other by one saros. (Over the course of such a saros series, the locations of the eclipses effectively scan across the Earth; the upcoming eclipse is number 30 in a series of 71 that began in 1501 AD with an eclipse near the North Pole and will end in 2763 AD with an eclipse near the South Pole.)

Any given moment in time will be in the middle of quite a few saros series (right now it’s 40)—and successive eclipses will always come from different series. But knowing about the saros cycle is a great first step in predicting eclipses—and it’s for example what the Antikythera device uses. In a sense, it’s a quintessential piece of science: take many observations, then synthesize a theory from them, or at least a scheme for computation.

It’s not clear what the Babylonians thought about abstract, formal systems. But the Greeks were definitely into them. And by 300 BC Euclid had defined his abstract system for geometry. So when someone like Ptolemy did astronomy, they did it a bit like Euclid—effectively taking things like the saros cycle as axioms, and then proving from them often surprisingly elaborate geometrical theorems, such as that there must be at least two solar eclipses in a given year.

Ptolemy’s Almagest from around 150 AD is an impressive piece of work, containing among many other things some quite elaborate procedures—and explicit tables—for predicting eclipses. (Yes, even in the later printed version, numbers are still represented confusingly by letters, as they always were in ancient Greek.)

Click to enlarge

In Ptolemy’s astronomy, Earth was assumed to be at the center of everything. But in modern terms that just meant he was choosing to use a different coordinate system—which didn’t affect most of the things he wanted to do, like working out the geometry of eclipses. And unlike the mainline Greek philosophers he wasn’t trying to make a fundamental theory of the world; he just wanted whatever epicycles and so on he needed to fit what he observed.

The Dawn of Modern Science

For more than a thousand years Ptolemy’s theory of the Moon defined the state of the art. In the 1300s Ibn al-Shatir revised Ptolemy’s models, achieving somewhat better accuracy. In 1472 Regiomontanus (Johannes Müller), systematizer of trigonometry, published more complete tables as part of his launch of what was essentially the first-ever scientific publishing company. But even in 1543 when Nicolaus Copernicus introduced his Sun-centered model of the solar system, the results he got were basically the same as Ptolemy’s, even though his underlying description of what was going on was quite different.

It’s said that Tycho Brahe got interested in astronomy in 1560 at age 13 when he saw a solar eclipse that had been predicted—and over the next several decades his careful observations uncovered several effects in the motion of the Moon (such as speeding up just before a full moon)—that eventually resulted in perhaps a factor 5 improvement in the prediction of its position. To Tycho eclipses were key tests, and he measured them carefully, and worked hard to be able to predict their timing more accurately than to within a few hours. (He himself never saw a total solar eclipse, only partial ones.)

Click to enlarge

Armed with Tycho’s observations, Johannes Kepler developed his description of orbits as ellipses—introducing concepts like inclination and eccentric anomaly—and in 1627 finally produced his Rudolphine Tables, which got right a lot of things that had been got wrong before, and included all sorts of detailed tables of lunar positions, as well as vastly better predictions for eclipses.

Click to enlarge

Using Kepler’s Rudolphine Tables (and a couple of pages of calculations) the first known actual map of a solar eclipse was published in 1654. And while there are some charming inaccuracies in overall geography, the geometry of the eclipse isn’t too bad.

Whether it was Ptolemy’s epicycles or Kepler’s ellipses, there were plenty of calculations to do in determining the motions of heavenly bodies (and indeed the first known mechanical calculator—excepting the Antikythera device—was developed by a friend of Kepler’s, presumably for the purpose). But there wasn’t really a coherent underlying theory; it was more a matter of describing effects in ways that could be used to make predictions.

So it was a big step forward in 1687 when Isaac Newton published his Principia, and claimed that with his laws for motion and gravity it should be possible—essentially from first principles—to calculate everything about the motion of the Moon. (Charmingly, in his “Theory of the World” section he simply asserts as his Proposition XXII “That all the motions of the Moon… follow from the principles which we have laid down.”)

Newton was proud of the fact that he could explain all sorts of known effects on the basis of his new theory. But when it came to actually calculating the detailed motion of the Moon he had a frustrating time. And even after several years he still couldn’t get the right answer—in later editions of the Principia adding the admission that actually “The apse of the Moon is about twice as swift” (i.e. his answer was wrong by a factor of 2).

Still, in 1702 Newton was happy enough with his results that he allowed them to be published, in the form of a 20-page booklet on the “Theory of the Moon”, which proclaimed that “By this Theory, what by all Astronomers was thought most difficult and almost impossible to be done, the Excellent Mr. Newton hath now effected, viz. to determine the Moon’s Place even in her Quadratures, and all other Parts of her Orbit, besides the Syzygys, so accurately by Calculation, that the Difference between that and her true Place in the Heavens shall scarce be two Minutes…”

Click to enlarge

Newton didn’t explain his methods (and actually it’s still not clear exactly what he did, or how mathematically rigorous it was or wasn’t). But his booklet effectively gave a step-by-step algorithm to compute the position of the Moon. He didn’t claim it worked “at the syzygys” (i.e. when the Sun, Moon and Earth are lined up for an eclipse)—though his advertised error of two arc-minutes was still much smaller than the angular size of the Moon in the sky.

But it wasn’t eclipses that were the focus then; it was a very practical problem of his day: knowing the location of a ship out in the open ocean. It’s possible to determine what latitude you’re at just by measuring how high the Sun gets in the sky. But to determine longitude you have to correct for the rotation of the Earth—and to do that you have to accurately keep track of time. But back in Newton’s day, the clocks that existed simply weren’t accurate enough, especially when they were being tossed around on a ship.

But particularly after various naval accidents, the problem of longitude was deemed important enough that the British government in 1714 established a “Board of Longitude” to offer prizes to help get it solved. One early suggestion was to use the regularity of the moons of Jupiter discovered by Galileo as a way to tell time. But it seemed that a simpler solution (not requiring a powerful telescope) might just be to measure the position of our Moon, say relative to certain fixed stars—and then to back-compute the time from this.

But to do this one had to have an accurate way to predict the motion of the Moon—which is what Newton was trying to provide. In reality, though, it took until the 1760s before tables were produced that were accurate enough to be able to determine time to within a minute (and thus distance to within 15 miles or so). And it so happens that right around the same time a marine chronometer was invented that was directly able to keep good time.

The Three-Body Problem

One of Newton’s great achievements in the Principia was to solve the so-called two-body problem, and to show that with an inverse square law of gravity the orbit of one body around another must always be what Kepler had said: an ellipse.

In a first approximation, one can think of the Moon as just orbiting the Earth in a simple elliptical orbit. But what makes everything difficult is that that’s just an approximation, because in reality there’s also a gravitational pull on the Moon from the Sun. And because of this, the Moon’s orbit is no longer a simple fixed ellipse—and in fact it ends up being much more complicated. There are a few definite effects one can describe and reason about. The ellipse gets stretched when the Earth is closer to the Sun in its own orbit. The orientation of the ellipse precesses like a top as a result of the influence of the Sun. But there’s no way in the end to work out the orbit by pure reasoning—so there’s no choice but to go into the mathematics and start solving the equations of the three-body problem.

In many ways this represented a new situation for science. In the past, one hadn’t ever been able to go far without having to figure out new laws of nature. But here the underlying laws were supposedly known, courtesy of Newton. Yet even given these laws, there was difficult mathematics involved in working out the behavior they implied.

Over the course of the 1700s and 1800s the effort to try to solve the three-body problem and determine the orbit of the Moon was at the center of mathematical physics—and attracted a veritable who’s who of mathematicians and physicists.

An early entrant was Leonhard Euler, who developed methods based on trigonometric series (including much of our current notation for such things), and whose works contain many immediately recognizable formulas:

Click to enlarge

In the mid-1740s there was a brief flap—also involving Euler’s “competitors” Clairaut and d’Alembert—about the possibility that the inverse-square law for gravity might be wrong. But the problem turned out to be with the calculations, and by 1748 Euler was using sums of about 20 trigonometric terms and proudly proclaiming that the tables he’d produced for the three-body problem had predicted the time of a total solar eclipse to within minutes. (Actually, he had said there’d be 5 minutes of totality, whereas in reality there was only 1—but he blamed this error on incorrect coordinates he’d been given for Berlin.)

Mathematical physics moved rapidly over the next few decades, with all sorts of now-famous methods being developed, notably by people like Lagrange. And by the 1770s, for example, Lagrange’s work was looking just like it could have come from a modern calculus book (or from a Wolfram|Alpha step-by-step solution):

Click to enlarge

Particularly in the hands of Laplace there was increasingly obvious success in deriving the observed phenomena of what he called “celestial mechanics” from mathematics—and in establishing the idea that mathematics alone could indeed generate new results in science.

At a practical level, measurements of things like the position of the Moon had always been much more accurate than calculations. But now they were becoming more comparable—driving advances in both. Meanwhile, there was increasing systematization in the production of ephemeris tables. And in 1767 the annual publication began of what was for many years the standard: the British Nautical Almanac.

The almanac quoted the position of the Moon to the arc-second, and systematically achieved at least arc-minute accuracy. The primary use of the almanac was for navigation (and it was what started the convention of using Greenwich as the “prime meridian” for measuring time). But right at the front of each year’s edition were the predicted times of the eclipses for that year—in 1767 just two solar eclipses:

Click to enlarge

The Math Gets More Serious

At a mathematical level, the three-body problem is about solving a system of three ordinary differential equations that give the positions of the three bodies as a function of time. If the positions are represented in standard 3D Cartesian coordinates , the equations can be stated in the form:

3D Cartesian coordinates equations

The {x, y, z} coordinates here aren’t, however, what traditionally show up in astronomy. For example, in describing the position of the Moon one might use longitude and latitude on a sphere around the Earth. Or, given that one knows the Moon has a roughly elliptical orbit, one might instead choose to describe its motions by variables that are based on deviations from such an orbit. In principle it’s just a matter of algebraic manipulation to restate the equations with any given choice of variables. But in practice what comes out is often long and complex—and can lead to formulas that fill many pages.

But, OK, so what are the best kinds of variables to use for the three-body problem? Maybe they should involve relative positions of pairs of bodies. Or relative angles. Or maybe positions in various kinds of rotating coordinate systems. Or maybe quantities that would be constant in a pure two-body problem. Over the course of the 1700s and 1800s many treatises were written exploring different possibilities.

But in essentially all cases the ultimate approach to the three-body problem was the same. Set up the problem with the chosen variables. Identify parameters that, if set to zero, would make the problem collapse to some easy-to-solve form. Then do a series expansion in powers of these parameters, keeping just some number of terms.

The calculations were difficult, and people’s results often didn’t agree. And for example in 1843 Ada Lovelace noted that “In the solution of the famous problem of the Three Bodies, there are, out of about 295 coefficients of lunar perturbations [that had recently been computed]… only 101… agree precisely both in signs and in amount [with previous works]” (going on to say, rather farsightedly, that this was something the Analytical Engine would be able to solve).

By the 1860s Charles Delaunay had, however, spent 20 years developing the most extensive theory of the Moon using series expansions. He’d identified five parameters with respect to which to do his expansions (eccentricities, inclinations, and ratios of orbit sizes)—and in the end he generated about 1800 pages like this (yes, he really needed Mathematica!):

Click to enlarge

But the sad fact was that despite all this effort, he didn’t get terribly good answers. And eventually it became clear why. The basic problem was that Delaunay wanted to represent his results in terms of functions like sin and cos. But in his computations, he often wanted to do series expansions with respect to the frequencies of those functions. Here’s a minimal case:

And here’s the problem. Take a look even at the second term. Yes, the δ parameter may be small. But how about the parameter, standing for time? If you don’t want to make predictions very far out, that’ll stay small. But what if you want to figure out what will happen further in the future?

Well, eventually that term will get big. And higher-order terms will get even bigger. But unless the Moon is going to escape its orbit or something, the final mathematical expressions that represent its position can’t have values that are too big. So in these expressions the so-called secular terms that increase with must somehow cancel out.

But the problem is that at any given order in the series expansion, there’s no guarantee that will happen in a numerically useful way. And in Delaunay’s case—even though with immense effort he often went to 7th order or beyond—it didn’t.

One nice feature of Delaunay’s computation was that it was in a sense entirely algebraic: everything was done symbolically, and only at the very end were actual numerical values of parameters substituted in.

But even before Delaunay, Peter Hansen had taken a different approach—substituting numbers as soon as he could, and dropping terms based on their numerical size rather than their symbolic form. His presentations look less pure (notice things like all those , where is the time in years), and it’s more difficult to tell what’s going on. But as a practical matter, his results were much better, and in fact were used for many national almanacs from about 1862 to 1922, achieving errors as small as 1 or 2 arc-seconds at least over periods of a decade or so. (Over longer periods, the errors could rapidly increase because of the lack of terms that had been dropped as a result of what amounted to numerical accidents.)

Click to enlarge

Both Delaunay and Hansen tried to represent orbits as series of powers and trigonometric functions (so-called Poisson series). But in the 1870s, George Hill in the US Nautical Almanac Office proposed instead using as a basis numerically computed functions that came from solving an equation for two-body motion with a periodic driving force of roughly the kind the Sun exerts on the Moon’s orbit. A large-scale effort was mounted, and starting in 1892 Ernest W. Brown (who had moved to the US, but had been a student of George Darwin, Charles Darwin’s physicist son) took charge of the project and in 1918 produced what would stand for many years as the definitive “Tables of the Motion of the Moon”.

Brown’s tables consist of hundreds of pages like this—ultimately representing the position of the Moon as a combination of about 1400 terms with very precise coefficients:

Click to enlarge

He says right at the beginning that the tables aren’t particularly intended for unique events like eclipses, but then goes ahead and does a “worked example” of computing an eclipse from 381 BC, reported by Ptolemy:

Click to enlarge

It was an impressive indication of how far things had come. But ironically enough the final presentation of Brown’s tables had the same sum-of-trigonometric-functions form that one would get from having lots of epicycles. At some level it’s not surprising, because any function can ultimately be represented by epicycles, just as it can be represented by a Fourier or other series. But it’s a strange quirk of history that such similar forms were used.

Can the Three-Body Problem Be Solved?

It’s all well and good that one can find approximations to the three-body problem, but what about just finding an outright solution—like as a mathematical formula? Even in the 1700s, there’d been some specific solutions found—like Euler’s collinear configuration, and Lagrange’s equilateral triangle. But a century later, no further solutions had been found—and finding a complete solution to the three-body problem was beginning to seem as hopeless as trisecting an angle, solving the quintic, or making a perpetual motion machine. (That sentiment was reflected for example in a letter Charles Babbage wrote Ada Lovelace in 1843 mentioning the “horrible problem [of] the three bodies”—even though this letter was later misinterpreted by Ada’s biographers to be about a romantic triangle, not the three-body problem of celestial mechanics.)

In contrast to the three-body problem, what seemed to make the two-body problem tractable was that its solutions could be completely characterized by “constants of the motion”—quantities that stay constant with time (in this case notably the direction of the axis of the ellipse). So for many years one of the big goals with the three-body problem was to find constants of the motion.

In 1887, though, Heinrich Bruns showed that there couldn’t be any such constants of the motion, at least expressible as algebraic functions of the standard {x, y, z} position and velocity coordinates of the three bodies. Then in the mid-1890s Henri Poincaré showed that actually there couldn’t be any constants of the motion that were expressible as any analytic functions of the positions, velocities and mass ratios.

One reason that was particularly disappointing at the time was that it had been hoped that somehow constants of the motion would be found in n-body problems that would lead to a mathematical proof of the long-term stability of the solar system. And as part of his work, Poincaré also saw something else: that at least in particular cases of the three-body problem, there was arbitrarily sensitive dependence on initial conditions—implying that even tiny errors in measurement could be amplified to arbitrarily large changes in predicted behavior (the classic “chaos theory” phenomenon).

But having discovered that particular solutions to the three-body problem could have this kind of instability, Poincaré took a different approach that would actually be characteristic of much of pure mathematics going forward: he decided to look not at individual solutions, but at the space of all possible solutions. And needless to say, he found that for the three-body problem, this was very complicated—though in his efforts to analyze it he invented the field of topology.

Poincaré’s work all but ended efforts to find complete solutions to the three-body problem. It also seemed to some to explain why the series expansions of Delaunay and others hadn’t worked out—though in 1912 Karl Sundman did show that at least in principle the three-body problem could be solved in terms of an infinite series, albeit one that converges outrageously slowly.

But what does it mean to say that there can’t be a solution to the three-body problem? Galois had shown that there couldn’t be a solution to the generic quintic equation, at least in terms of radicals. But actually it’s still perfectly possible to express the solution in terms of elliptic or hypergeometric functions. So why can’t there be some more sophisticated class of functions that can be used to just “solve the three-body problem”?

Here are some pictures of what can actually happen in the three-body problem, with various initial conditions:

Click to enlarge

And looking at these immediately gives some indication of why it’s not easy to just “solve the three-body problem”. Yes, there are cases where what happens is fairly simple. But there are also cases where it’s not, and where the trajectories of the three bodies continue to be complicated and tangled for a long time.

So what’s fundamentally going on here? I don’t think traditional mathematics is the place to look. But I think what we’re seeing is actually an example of a general phenomenon I call computational irreducibility that I discovered in the 1980s in studying the computational universe of possible programs.

Many programs, like many instances of the three-body problem, behave in quite simple ways. But if you just start looking at all possible simple programs, it doesn’t take long before you start seeing behavior like this:

How can one tell what’s going to happen? Well, one can just keep explicitly running each program and seeing what it does. But the question is: is there some systematic way to jump ahead, and to predict what will happen without tracing through all the steps?

The answer is that in general there isn’t. And what I call the Principle of Computational Equivalence suggests that pretty much whenever one sees complex behavior, there won’t be.

Here’s the way to think about this. The system one’s studying is effectively doing a computation to work out what its behavior will be. So to jump ahead we’d in a sense have to do a more sophisticated computation. But what the Principle of Computational Equivalence says is that actually we can’t—and that whether we’re using our brains or our mathematics or a Turing machine or anything else, we’re always stuck with computations of the same sophistication.

So what about the three-body problem? Well, I strongly suspect that it’s an example of computational irreducibility: that in effect the computations it’s doing are as sophisticated as any computations that we can do, so there’s no way we can ever expect to systematically jump ahead and solve the problem. (We also can’t expect to just define some new finite class of functions that can just be evaluated to give the solution.)

I’m hoping that one day someone will rigorously prove this. There’s some technical difficulty, because the three-body problem is usually formulated in terms of real numbers that immediately have an infinite number of digits—but to compare with ordinary computation one has to require finite processes to set up initial conditions. (Ultimately one wants to show for example that there’s a “compiler” that can go from any program, say for a Turing machine, and can generate instructions to set up initial conditions for a three-body problem so that the evolution of the three-body problem will give the same results as running that program—implying that the three-body problem is capable of universal computation.)

I have to say that I consider Newton in a sense very lucky. It could have been that it wouldn’t have been possible to work out anything interesting from his theory without encountering the kind of difficulties he had with the motion of the Moon—because one would always be running into computational irreducibility. But in fact, there was enough computational reducibility and enough that could be computed easily that one could see that the theory was useful in predicting features of the world (and not getting wrong answers, like with the apse of the Moon)—even if there were some parts that might take two centuries to work out, or never be possible at all.

Newton himself was certainly aware of the potential issue, saying that at least if one was dealing with gravitational interactions between many planets then “to define these motions by exact laws admitting of easy calculation exceeds, if I am not mistaken, the force of any human mind”.
And even today it’s extremely difficult to know what the long-term evolution of the solar system will be.

It’s not particularly that there’s sensitive dependence on initial conditions: we actually have measurements that should be precise enough to determine what will happen for a long time. The problem is that we just have to do the computation—a bit like computing the digits of π—to work out the behavior of the -body problem that is our solar system.

Existing simulations show that for perhaps a few tens of millions of years, nothing too dramatic can happen. But after that we don’t know. Planets could change their order. Maybe they could even collide, or be ejected from the solar system. Computational irreducibility implies that at least after an infinite time it’s actually formally undecidable (in the sense of Gödel’s theorem or the halting problem) what can happen.

One of my children, when they were very young, asked me whether when dinosaurs existed the Earth could have had two moons. For years when I ran into celestial mechanics experts I would ask them that question—and it was notable how difficult they found it. Most now say that at least at the time of the dinosaurs we couldn’t have had an extra moon—though a billion years earlier it’s not clear.

We used to only have one system of planets to study. And the fact that there were (then) 9 of them used to be a classic philosopher’s example of a truth about the world that just happens to be the way it is, and isn’t “necessarily true” (like 2 + 2 = 4). But now of course we know about lots of exoplanets. And it’s beginning to look as if there might be a theory for things like how many planets a solar system is likely to have.

At some level there’s presumably a process like natural selection: some configurations of planets aren’t “fit enough” to be stable—and only those that are survive. In biology it’s traditionally been assumed that natural selection and adaptation is somehow what’s led to the complexity we see. But actually I suspect much of it is instead just a reflection of what generally happens in the computational universe—both in biology and in celestial mechanics. Now in celestial mechanics, we haven’t yet seen in the wild any particularly complex forms (beyond a few complicated gap structures in rings, and tumbling moons and asteroids). But perhaps elsewhere we’ll see things like those obviously tangled solutions to the three-body problem—that come closer to what we’re used to in biology.

It’s remarkable how similar the issues are across so many different fields. For example, the whole idea of using “perturbation theory” and series expansions that has existed since the 1700s in celestial mechanics is now also core to quantum field theory. But just like in celestial mechanics there’s trouble with convergence (maybe one should try renormalization or resummation in celestial mechanics). And in the end one begins to realize that there are phenomena—no doubt like turbulence or the three-body problem—that inevitably involve more sophisticated computations, and that need to be studied not with traditional mathematics of the kind that was so successful for Newton and his followers but with the kind of science that comes from exploring the computational universe.

Approaching Modern Times

But let’s get back to the story of the motion of the Moon. Between Brown’s tables, and Poincaré’s theoretical work, by the beginning of the 1900s the general impression was that whatever could reasonably be computed about the motion of the Moon had been computed.

Occasionally there were tests. Like for example in 1925, when there was a total solar eclipse visible in New York City, and the New York Times perhaps overdramatically said that “scientists [were] tense… [wondering] whether they or Moon is wrong as eclipse lags five seconds behind”. The fact is that a prediction accurate to 5 seconds was remarkably good, and we can’t do all that much better even today. (By the way, the actual article talks extensively about “Professor Brown”—as well as about how the eclipse might “disprove Einstein” and corroborate the existence of “coronium”—but doesn’t elaborate on the supposed prediction error.)

Click to enlarge

As a practical matter, Brown’s tables were not exactly easy to use: to find the position of the Moon from them required lots of mechanical desk calculator work, as well as careful transcription of numbers. And this led Leslie Comrie in 1932 to propose using a punch-card-based IBM Hollerith automatic tabulator—and with the help of Thomas Watson, CEO of IBM, what was probably the first “scientific computing laboratory” was established—to automate computations from Brown’s tables.

Click to enlarge

(When I was in elementary school in England in the late 1960s—before electronic calculators—I always carried around, along with my slide rule, a little book of “4-figure mathematical tables”. I think I found it odd that such a book would have an author—and perhaps for that reason I still remember the name: “L. J. Comrie”.)

By the 1950s, the calculations in Brown’s tables were slowly being rearranged and improved to make them more suitable for computers. But then with John F. Kennedy’s 1962 “We choose to go to the Moon”, there was suddenly urgent interest in getting the most accurate computations of the Moon’s position. As it turned out, though, it was basically just a tweaked version of Brown’s tables, running on a mainframe computer, that did the computations for the Apollo program.

At first, computers were used in celestial mechanics purely for numerical computation. But by the mid-1960s there were also experiments in using them for algebraic computation, and particularly to automate the generation of series expansions. Wallace Eckert at IBM started using FORMAC to redo Brown’s tables, while in Cambridge David Barton and Steve Bourne (later the creator of the “Bourne shell” (sh) in Unix) built their own CAMAL computer algebra system to try extending the kind of thing Delaunay had done. (And by 1970, Delaunay’s 7th-order calculations had been extended to 20th order.)

When I myself started to work on computer algebra in 1976 (primarily for computations in particle physics), I’d certainly heard about CAMAL, but I didn’t know what it had been used for (beyond vaguely “celestial mechanics”). And as a practicing theoretical physicist in the late 1970s, I have to say that the “problem of the Moon” that had been so prominent in the 1700s and 1800s had by then fallen into complete obscurity.

I remember for example in 1984 asking a certain Martin Gutzwiller, who was talking about quantum chaos, what his main interest actually was. And when he said “the problem of the Moon”, I was floored; I didn’t know there still was any problem with the Moon. As it turns out, in writing this post I found out that Gutzwiller was actually the person who took over from Eckert and spent nearly two decades working on trying to improve the computations of the position of the Moon.

Why Not Just Solve It?

Traditional approaches to the three-body problem come very much from a mathematical way of thinking. But modern computational thinking immediately suggests a different approach. Given the differential equations for the three-body problem, why not just directly solve them? And indeed in the Wolfram Language there’s a built-in function NDSolve for numerically solving systems of differential equations.

So what happens if one just feeds in equations for a three-body problem? Well, here are the equations:

Now as an example let’s set the masses to random values:

And let’s define the initial position and velocity for each body to be random as well:

Now we can just use NDSolve to get the solutions (it gives them as implicit approximate numerical functions of ):

And now we can plot them. And now we’ve got a solution to a three-body problem, just like that!

Well, obviously this is using the Wolfram Language and a huge tower of modern technology. But would it have been possible even right from the beginning for people to generate direct numerical solutions to the three-body problem, rather than doing all that algebra? Back in the 1700s, Euler already knew what’s now called Euler’s method for finding approximate numerical solutions to differential equations. So what if he’d just used that method to calculate the motion of the Moon?

The method relies on taking a sequence of discrete steps in time. And if he’d used, say, a step size of a minute, then he’d have had to take 40,000 steps to get results for a month, but he should have been able to successfully reproduce the position of the Moon to about a percent. If he’d tried to extend to 3 months, however, then he would already have had at least a 10% error.

Any numerical scheme for solving differential equations in practice eventually builds up some kind of error—but the more one knows about the equations one’s solving, and their expected solutions, the more one’s able to preprocess and adapt things to minimize the error. NDSolve has enough automatic adaptivity built into it that it’ll do pretty well for a surprisingly long time on a typical three-body problem. (It helps that the Wolfram Language and NDSolve can handle numbers with arbitrary precision, not just machine precision.)

But if one looks, say, at the total energy of the three-body system—which one can prove from the equations should stay constant—then one will typically see an error slowly build up in it. One can avoid this if one effectively does a change of variables in the equations to “factor out” energy. And one can imagine doing a whole hierarchy of algebraic transformations that in a sense give the numerical scheme as much help as possible.

And indeed since at least the 1980s that’s exactly what’s been done in practical work on the three-body problem, and the Earth-Moon-Sun system. So in effect it’s a mixture of the traditional algebraic approach from the 1700s and 1800s, together with modern numerical computation.

The Real Earth-Moon-Sun Problem

OK, so what’s involved in solving the real problem of the Earth-Moon-Sun system? The standard three-body problem gives a remarkably good approximation to the physics of what’s happening. But it’s obviously not the whole story.

For a start, the Earth isn’t the only planet in the solar system. And if one’s trying to get sufficiently accurate answers, one’s going to have to take into account the gravitational effect of other planets. The most important is Jupiter, and its typical effect on the orbit of the Moon is at about the level—sufficiently large that for example Brown had to take it into account in his tables.

The next effect is that the Earth isn’t just a point mass, or even a precise sphere. Its rotation makes it bulge at the equator, and that affects the orbit of the Moon at the level.

Orbits around the Earth ultimately depend on the full mass distribution and gravitational field of the Earth (which is what Sputnik-1 was nominally launched to map)—and both this, and the reverse effect from the Moon, come in at the level. At the level there are then effects from tidal deformations (“solid tides”) on the Earth and Moon, as well as from gravitational redshift and other general relativistic phenomena.

To predict the position of the Moon as accurately as possible one ultimately has to have at least some model for these various effects.

But there’s a much more immediate issue to deal with: one has to know the initial conditions for the Earth, Sun and Moon, or in other words, one has to know as accurately as possible what their positions and velocities were at some particular time.

And conveniently enough, there’s now a really good way to do that, because Apollo 11, 14 and 15 all left laser retroreflectors on the Moon. And by precisely timing how long it takes a laser pulse from the Earth to round-trip to these retroreflectors, it’s now possible in effect to measure the position of the Moon to millimeter accuracy.

OK, so how do modern analogs of the Babylonian ephemerides actually work? Internally they’re dealing with the equations for all the significant bodies in the solar system. They do symbolic preprocessing to make their numerical work as easy as possible. And then they directly solve the differential equations for the system, appropriately inserting models for things like the mass distribution in the Earth.

They start from particular measured initial conditions, but then they repeatedly insert new measurements, trying to correct the parameters of the model so as to optimally reproduce all the measurements they have. It’s very much like a typical machine learning task—with the training data here being observations of the solar system (and typically fitting just being least squares).

But, OK, so there’s a model one can run to figure out something like the position of the Moon. But one doesn’t want to have to explicitly do that every time one needs to get a result; instead one wants in effect just to store a big table of pre-computed results, and then to do something like interpolation to get any particular result one needs. And indeed that’s how it’s done today.

How It’s Really Done

Back in the 1960s NASA started directly solving differential equations for the motion of planets. The Moon was more difficult to deal with, but by the 1980s that too was being handled in a similar way. Ongoing data from things like the lunar retroreflectors was added, and all available historical data was inserted as well.

The result of all this was the JPL Development Ephemeris (JPL DE). In addition to new observations being used, the underlying system gets updated every few years, for example to get what’s needed for some spacecraft going to some new place in the solar system. (The latest is DE441—that follows DE432, which was built for going to Pluto.)

But so how is the actual ephemeris delivered? Well, for every thousand years covered, the ephemeris has about 100 megabytes of results, given as coefficients for Chebyshev polynomials, which are convenient for interpolation. And for any given quantity in any given coordinate system over a particular period of time, one accesses the appropriate parts of these results.

In Wolfram Language, it’s all packaged up into the function AstroPosition—which here gives the position of the Moon in coordinates relative to the equator of the Earth right now:

OK, but so how does one find an eclipse? Well, it’s an iterative process. Start with an approximation, perhaps from the saros cycle. Then interpolate the ephemeris and look at the result. Then keep iterating until one finds out just when the Moon will be in the appropriate position.

But actually there’s some more to do. Because what’s originally computed are the positions of the barycenters (centers of mass) of the various bodies. But now one has to figure out how the bodies are oriented.

The Earth rotates, and we know its rate quite precisely. But the Moon is basically locked with the same face pointing to the Earth, except that in practice there are small “librations” where the Moon wobbles a little back and forth—and these turn out to be particularly troublesome to predict.

Where Will the Eclipse Be?

OK, so let’s say one knows where the Earth, Moon and Sun are. How does one then figure out what type of eclipse will happen, and where on the Earth the eclipse will actually hit? Well, there’s some further geometry to do. And here’s the beginning of what’s involved:

Basically the Moon generates a cone of shadow, and then the question is how this cone intersects the Earth. If the tip of the cone is inside the Earth, that means there’ll be a region of total shadow (“umbra”) on the Earth—and a total eclipse. (If the tip is above the Earth, there’ll be an annular eclipse, in which there’s a “ring of sun” visible around the Moon.)

By the way, the more complete geometry is like this (again not to scale)

where now we’ve included the penumbra in which only part of the Sun is shadowed by the Moon. In the particular case shown, the umbra cone “misses the Earth”, so there’s no total eclipse, but there’s still a partial eclipse where part of the Sun is shadowed.

OK, but let’s say there’s going to be a total eclipse. To see where on the Earth the region of totality will be, we have to work out where the “cone of total shadow” (i.e. umbra) will intersect the surface of the Earth. It’s a somewhat complicated 3D geometry problem:

It’s easiest to understand what happens by looking at things from the position of the Sun. The light gray region is the penumbra, and the little black dot is the region of totality (i.e. the umbra):

As the Earth and Moon move in their orbits, the region of shadow will move relative to the Earth:

But now there’s another part of the story, which is the rotation of the Earth. And if we include that, we’ll see that the region of totality (at least in this case) traces out a kind of S-shaped curve on the surface of the Earth:

All this tricky geometry got figured out in 1824 by Friedrich Bessel, who introduced what are now called the Besselian elements—eight variables that specify the location, orientation and aperture of the umbra and penumbra cones, as well as the orientation of the Earth, as a function of time. And for any given eclipse, the whole story of its appearance and trajectory on the surface of the Earth is determined by its Besselian elements.

When Will the Eclipse Arrive?

OK, so now we know what the trajectory of an eclipse will be. But how do we figure out at what time the eclipse will actually reach a given point on Earth? Well, first we have to be clear on our definition of time. And there’s an immediate issue with the speed of light and special relativity. What does it mean to say that the positions of the Earth and Sun are such-and-such at such-and-such a time? Because it takes light about 8 minutes to get to the Earth from the Sun, we only get to see where the Sun was 8 minutes ago, not where it is now.

And what we need is really a classic special relativity setup. We essentially imagine that the solar system is filled with a grid of clocks that have been synchronized by light pulses. And what a modern ephemeris does is to quote the results for positions of bodies in the solar system relative to the times on those clocks. (General relativity implies that in different gravitational fields the clocks will run at different rates, but for our purposes this is a tiny effect. But what isn’t a tiny effect is including retardation in the equations for the -body problem—making them become delay differential equations.)

But now there’s another issue. If one’s observing the eclipse, one’s going to be using some timepiece (phone?) to figure out what time it is. And if it’s working properly that timepiece should show official “civil time” that’s based on UTC—which is what NTP internet time is synchronized to. But the issue is that UTC has a complicated relationship to the time used in the astronomical ephemeris.

The starting point is what’s called UT1: a definition of time in which one day is the average time it takes the Earth to rotate once relative to the Sun. But the point is that this average time isn’t constant, because the rotation of the Earth is gradually slowing down, primarily as a result of interactions with the Moon. But meanwhile, UTC is defined by an atomic clock whose timekeeping is independent of any issues about the rotation of the Earth.

There’s a convention for keeping UT1 aligned with UTC: if UT1 is going to get more than 0.9 seconds away from UTC, then a leap second is added to UTC. One might think this would be a tiny effect, but actually, since 1972, a total of 27 leap seconds have been added (as specified in the Wolfram Language by GeoOrientationData):

Exactly when a new leap second will be needed is unpredictable; it depends on things like what earthquakes have occurred. But we need to account for leap seconds if we’re going to get the time of the eclipse correct to the second relative to UTC or internet time.

There are a few other effects that are also important in the precise observed timing of the eclipse. The most obvious is geo elevation. In doing astronomical computations, the Earth is assumed to be an ellipsoid. (There are many different definitions, corresponding to different geodetic “datums”—and that’s an issue in defining things like “sea level”, but it’s not relevant here.) But if you’re at a different height above the ellipsoid, the cone of shadow from the eclipse will reach you at a different time. And the size of this effect can be as much as 0.3 seconds for every 1000 feet of height.

All of the effects we’ve talked about we’re readily able to account for. But there is one remaining effect that’s a bit more difficult. Right at the beginning or end of totality one typically sees points of light on the rim of the Moon. Known as Baily’s beads, these are the result of rays of light that make it to us between mountains on the Moon. Figuring out exactly when all these rays are extinguished requires taking geo elevation data for the Moon, and effectively doing full 3D ray tracing. And in doing this, one ends with the rather bizarre conclusion that the region of shadow on the earth isn’t a perfect circle; instead it’s roughly a polygon, each of whose edges is associated with a particular Baily’s bead. The effect of this can last as long as a second, and can cause the precise edge of totality to move by as much as a mile. (One can also imagine effects having to do with the corona of the Sun, which is constantly changing.)

But in the end, even though the shadow of the Moon on the Earth moves at more than 1000 mph, modern science successfully makes it possible to compute when the shadow will reach a particular point on Earth to an accuracy of perhaps a second. And that’s what our precisioneclipse.com website is set up to do.

Eclipse Experiences

Written August 15, 2017

I saw my first partial solar eclipse more than 50 years ago. And I’ve seen one total solar eclipse before in my life—in 1991. It was the longest eclipse (6 minutes 53 seconds) that’ll happen for more than a century.

There was a certain irony to my experience, though, especially in view of our efforts now to predict the exact arrival time of next week’s eclipse. I’d chartered a plane and flown to a small airport in Mexico (yes, that’s me on the left with the silly hat)—and my friends and I had walked to a beautiful deserted beach, and were waiting under a cloudless sky for the total eclipse to begin.

Click to enlarge

I felt proud of how prepared I was—with maps marking to the minute when the eclipse should arrive. But then I realized: there we were, out on a beach with no obvious signs of modern civilization—and nobody had brought any properly set timekeeping device (and in those days my cellphone was just a phone, and didn’t even have signal there).

And so it was that I missed seeing a demonstration of an impressive achievement of science. And instead I got to experience the eclipse pretty much the way people throughout history have experienced eclipses—even if I did know that the Moon would continue gradually eating into the Sun and eventually cover it, and that it wouldn’t make the world end.

There’s always something sobering about astronomical events, and about realizing just how tiny human scales are compared to them. Billions of eclipses have happened over the course of the Earth’s history. Recorded history has covered only a few thousand of them. On average, there’s an eclipse at any given place on Earth roughly every 400 years; in Jackson, WY, where I’m planning to see next week’s eclipse, it turns out the next total eclipse will be 727 years from now—in 2744.

In earlier times, civilizations built giant monuments to celebrate the motions of the Sun and Moon. Today, for the eclipse next week, what we’re making is a website. But that website builds on one of the great epics of human intellectual history—stretching back to the earliest times of systematic science, and encompassing contributions from a remarkable cross-section of the most celebrated scientists and mathematicians from past centuries.

It’ll be about 9538 days since the eclipse I saw in 1991. The Moon will have traveled some 500 million miles around the Earth, and the Earth some 15 billion miles around the Sun. But now—in a remarkable triumph of science—we’re computing to the second when they’ll be lined up again.

Written in Anticipation of April 8, 2024

In the days leading up to August 21, 2017, millions of people accessed our precisioneclipse.com website—with their geoIPs increasingly concentrating near the path of totality. I had traveled to Wyoming and—with a couple of hours to spare—found a place with a clear view across a valley. And this now being 2017, I tweeted:

Click to enlarge

But would our carefully done computations actually be accurate? The eclipse was going to make landfall on the Oregon coast, and, conveniently, we had a spotter right there. From where I was, the partial eclipse was well underway. And then I got a text: yes, totality in Oregon had come at the predicted second. It would take the shadow of the Moon a little under 20 minutes to reach me.

Unlike in 1991, I and everyone else had a cellphone with a precise clock—and was able to access our precisioneclipse.com site:

Click to enlarge

While I was waiting I was making little images of the crescent Sun with my fingers—repeating something I’d first noticed more than 50 years earlier when I saw my first partial eclipse at the age of 6:

Click to enlarge

A few minutes before the eclipse, I started to see a strange shimmering (invisible on any video I took): shadow bands, a strange and poorly understood eclipse phenomenon. And then, there it came, sweeping across the valley: totality. Arriving right at the predicted second:

Click to enlarge

I had set up a camera to capture a video of the eclipse, and a little later that day I did an analysis of it—and, since earlier in 2017 I’d started routinely doing livestreams, I livestreamed it:

Click to enlarge

At the end, I published my notebook to the Wolfram Cloud, and it’s still there:

Click to enlarge

And now six years have passed. Much has happened in our human world. But the Moon has just inexorably continued in its orbit. And 2422 days later, it will once again line up to create a total eclipse…

PHRC : 68 projets financés en psychiatrie entre 2012 à 2019

Par : SanteMentale
1 février 2024 à 16:04

A l’occasion des 30 ans du Programme hospitalier de recherche clinique (PHRC), le ministère du Travail, de la Santé et des Solidarités a présenté, le ...

Lire la suite

L’article PHRC : 68 projets financés en psychiatrie entre 2012 à 2019 est apparu en premier sur Santé Mentale.

How to Think Computationally about AI, the Universe and Everything

27 octobre 2023 à 21:47

Transcript of a talk at TED AI on October 17, 2023, in San Francisco

Human language. Mathematics. Logic. These are all ways to formalize the world. And in our century there’s a new and yet more powerful one: computation.

And for nearly 50 years I’ve had the great privilege of building an ever taller tower of science and technology based on that idea of computation. And today I want to tell you some of what that’s led to.

There’s a lot to talk about—so I’m going to go quickly… sometimes with just a sentence summarizing what I’ve written a whole book about.

You know, I last gave a TED talk thirteen years ago—in February 2010—soon after Wolfram|Alpha launched.

TED Talk 2010

And I ended that talk with a question: is computation ultimately what’s underneath everything in our universe?

I gave myself a decade to find out. And actually it could have needed a century. But in April 2020—just after the decade mark—we were thrilled to be able to announce what seems to be the ultimate “machine code” of the universe.

Wolfram Physics Project

And, yes, it’s computational. So computation isn’t just a possible formalization; it’s the ultimate one for our universe.

It all starts from the idea that space—like matter—is made of discrete elements. And that the structure of space and everything in it is just defined by the network of relations between these elements—that we might call atoms of space. It’s very elegant—but deeply abstract.

But here’s a humanized representation:

A version of the very beginning of the universe. And what we’re seeing here is the emergence of space and everything in it by the successive application of very simple computational rules. And, remember, those dots are not atoms in any existing space. They’re atoms of space—that are getting put together to make space. And, yes, if we kept going long enough, we could build our whole universe this way.

Eons later here’s a chunk of space with two little black holes, that eventually merge, radiating ripples of gravitational radiation:

And remember—all this is built from pure computation. But like fluid mechanics emerging from molecules, what emerges here is spacetime—and Einstein’s equations for gravity. Though there are deviations that we just might be able to detect. Like that the dimensionality of space won’t always be precisely 3.

And there’s something else. Our computational rules can inevitably be applied in many ways, each defining a different thread of time—a different path of history—that can branch and merge:

But as observers embedded in this universe, we’re branching and merging too. And it turns out that quantum mechanics emerges as the story of how branching minds perceive a branching universe.

The little pink lines here show the structure of what we call branchial space—the space of quantum branches. And one of the stunningly beautiful things—at least for a physicist like me—is that the same phenomenon that in physical space gives us gravity, in branchial space gives us quantum mechanics.

In the history of science so far, I think we can identify four broad paradigms for making models of the world—that can be distinguished by how they deal with time.

4 paradigms

In antiquity—and in plenty of areas of science even today—it’s all about “what things are made of”, and time doesn’t really enter. But in the 1600s came the idea of modeling things with mathematical formulas—in which time enters, but basically just as a coordinate value.

Then in the 1980s—and this is something in which I was deeply involved—came the idea of making models by starting with simple computational rules and then just letting them run:

Can one predict what will happen? No, there’s what I call computational irreducibility: in effect the passage of time corresponds to an irreducible computation that we have to run to know how it will turn out.

But now there’s something even more: in our Physics Project things become multicomputational, with many threads of time, that can only be knitted together by an observer.

It’s a new paradigm—that actually seems to unlock things not only in fundamental physics, but also in the foundations of mathematics and computer science, and possibly in areas like biology and economics too.

You know, I talked about building up the universe by repeatedly applying a computational rule. But how is that rule picked? Well, actually, it isn’t. Because all possible rules are used. And we’re building up what I call the ruliad: the deeply abstract but unique object that is the entangled limit of all possible computational processes. Here’s a tiny fragment of it shown in terms of Turing machines:

OK, so the ruliad is everything. And we as observers are necessarily part of it. In the ruliad as a whole, everything computationally possible can happen. But observers like us can just sample specific slices of the ruliad.

And there are two crucial facts about us. First, we’re computationally bounded—our minds are limited. And second, we believe we’re persistent in time—even though we’re made of different atoms of space at every moment.

So then here’s the big result. What observers with those characteristics perceive in the ruliad necessarily follows certain laws. And those laws turn out to be precisely the three key theories of 20th-century physics: general relativity, quantum mechanics, and statistical mechanics and the Second Law.

It’s because we’re observers like us that we perceive the laws of physics we do.

We can think of different minds as being at different places in rulial space. Human minds who think alike are nearby. Animals further away. And further out we get to alien minds where it’s hard to make a translation.

How can we get intuition for all this? We can use generative AI to take what amounts to an incredibly tiny slice of the ruliad—aligned with images we humans have produced.

We can think of this as a place in the ruliad described using the concept of a cat in a party hat:

Zooming out, we see what we might call “cat island”. But pretty soon we’re in interconcept space. Occasionally things will look familiar, but mostly we’ll see things we humans don’t have words for.

In physical space we explore more of the universe by sending out spacecraft. In rulial space we explore more by expanding our concepts and our paradigms.

We can get a sense of what’s out there by sampling possible rules—doing what I call ruliology:

Even with incredibly simple rules there’s incredible richness. But the issue is that most of it doesn’t yet connect with things we humans understand or care about. It’s like when we look at the natural world and only gradually realize we can use features of it for technology. Even after everything our civilization has achieved, we’re just at the very, very beginning of exploring rulial space.

But what about AIs? Just like we can do ruliology, AIs can in principle go out and explore rulial space. But left to their own devices, they’ll mostly be doing things we humans don’t connect with, or care about.

The big achievements of AI in recent times have been about making systems that are closely aligned with us humans. We train LLMs on billions of webpages so they can produce text that’s typical of what we humans write. And, yes, the fact that this works is undoubtedly telling us some deep scientific things about the semantic grammar of language—and generalizations of things like logic—that perhaps we should have known centuries ago.

You know, for much of human history we were kind of like LLMs, figuring things out by matching patterns in our minds. But then came more systematic formalization—and eventually computation. And with that we got a whole other level of power—to create truly new things, and in effect to go wherever we want in the ruliad.

But the challenge is to do that in a way that connects with what we humans—and our AIs—understand.

And in fact I’ve devoted a large part of my life to building that bridge. It’s all been about creating a language for expressing ourselves computationally: a language for computational thinking.

The goal is to formalize what we know about the world—in computational terms. To have computational ways to represent cities and chemicals and movies and formulas—and our knowledge about them.

It’s been a vast undertaking—that’s spanned more than four decades of my life. It’s something very unique and different. But I’m happy to report that in what has been Mathematica and is now the Wolfram Language I think we have now firmly succeeded in creating a truly full-scale computational language.

In effect, every one of the functions here can be thought of as formalizing—and encapsulating in computational terms—some facet of the intellectual achievements of our civilization:

It’s the most concentrated form of intellectual expression I know: finding the essence of everything and coherently expressing it in the design of our computational language. For me personally it’s been an amazing journey, year after year building the tower of ideas and technology that’s needed—and nowadays sharing that process with the world on open livestreams.

A few centuries ago the development of mathematical notation, and what amounts to the “language of mathematics”, gave a systematic way to express math—and made possible algebra, and calculus, and ultimately all of modern mathematical science. And computational language now provides a similar path—letting us ultimately create a “computational X” for all imaginable fields X.

We’ve seen the growth of computer science—CS. But computational language opens up something ultimately much bigger and broader: CX. For 70 years we’ve had programming languages—which are about telling computers in their terms what to do. But computational language is about something intellectually much bigger: it’s about taking everything we can think about and operationalizing it in computational terms.

You know, I built the Wolfram Language first and foremost because I wanted to use it myself. And now when I use it, I feel like it’s giving me a superpower:

I just have to imagine something in computational terms and then the language almost magically lets me bring it into reality, see its consequences and then build on them. And, yes, that’s the superpower that’s let me do things like our Physics Project.

And over the past 35 years it’s been my great privilege to share this superpower with many other people—and by doing so to have enabled such an incredible number of advances across so many fields. It’s a wonderful thing to see people—researchers, CEOs, kids—using our language to fluently think in computational terms, crispening up their own thinking and then in effect automatically calling in computational superpowers.

And now it’s not just people who can do that. AIs can use our computational language as a tool too. Yes, to get their facts straight, but even more importantly, to compute new facts. There are already some integrations of our technology into LLMs—and there’s a lot more you’ll be seeing soon. And, you know, when it comes to building new things, a very powerful emerging workflow is basically to start by telling the LLM roughly what you want, then have it try to express that in precise Wolfram Language. Then—and this is a critical feature of our computational language compared to a programming language—you as a human can “read the code”. And if it does what you want, you can use it as a dependable component to build on.

OK, but let’s say we use more and more AI—and more and more computation. What’s the world going to be like? From the Industrial Revolution on, we’ve been used to doing engineering where we can in effect “see how the gears mesh” to “understand” how things work. But computational irreducibility now shows that won’t always be possible. We won’t always be able to make a simple human—or, say, mathematical—narrative to explain or predict what a system will do.

And, yes, this is science in effect eating itself from the inside. From all the successes of mathematical science we’ve come to believe that somehow—if only we could find them—there’d be formulas to predict everything. But now computational irreducibility shows that isn’t true. And that in effect to find out what a system will do, we have to go through the same irreducible computational steps as the system itself.

Yes, it’s a weakness of science. But it’s also why the passage of time is significant—and meaningful. We can’t just jump ahead and get the answer; we have to “live the steps”.

It’s going to be a great societal dilemma of the future. If we let our AIs achieve their full computational potential, they’ll have lots of computational irreducibility, and we won’t be able to predict what they’ll do. But if we put constraints on them to make them predictable, we’ll limit what they can do for us.

So what will it feel like if our world is full of computational irreducibility? Well, it’s really nothing new—because that’s the story with much of nature. And what’s happened there is that we’ve found ways to operate within nature—even though nature can still surprise us.

And so it will be with the AIs. We might give them a constitution, but there will always be consequences we can’t predict. Of course, even figuring out societally what we want from the AIs is hard. Maybe we need a promptocracy where people write prompts instead of just voting. But basically every control-the-outcome scheme seems full of both political philosophy and computational irreducibility gotchas.

You know, if we look at the whole arc of human history, the one thing that’s systematically changed is that more and more gets automated. And LLMs just gave us a dramatic and unexpected example of that. So does that mean that in the end we humans will have nothing to do? Well, if you look at history, what seems to happen is that when one thing gets automated away, it opens up lots of new things to do. And as economies develop, the pie chart of occupations seems to get more and more fragmented.

And now we’re back to the ruliad. Because at a foundational level what’s happening is that automation is opening up more directions to go in the ruliad. And there’s no abstract way to choose between them. It’s just a question of what we humans want—and it requires humans “doing work” to define that.

A society of AIs untethered by human input would effectively go off and explore the whole ruliad. But most of what they’d do would seem to us random and pointless. Much like now most of nature doesn’t seem like it’s “achieving a purpose”.

One used to imagine that to build things that are useful to us, we’d have to do it step by step. But AI and the whole phenomenon of computation tell us that really what we need is more just to define what we want. Then computation, AI, automation can make it happen.

And, yes, I think the key to defining in a clear way what we want is computational language. You know—even after 35 years—for many people the Wolfram Language is still an artifact from the future. If your job is to program it seems like a cheat: how come you can do in an hour what would usually take a week? But it can also be daunting, because having dashed off that one thing, you now have to conceptualize the next thing. Of course, it’s great for CEOs and CTOs and intellectual leaders who are ready to race onto the next thing. And indeed it’s impressively popular in that set.

In a sense, what’s happening is that Wolfram Language shifts from concentrating on mechanics to concentrating on conceptualization. And the key to that conceptualization is broad computational thinking. So how can one learn to do that? It’s not really a story of CS. It’s really a story of CX. And as a kind of education, it’s more like liberal arts than STEM. It’s part of a trend that when you automate technical execution, what becomes important is not figuring out how to do things—but what to do. And that’s more a story of broad knowledge and general thinking than any kind of narrow specialization.

You know, there’s an unexpected human-centeredness to all of this. We might have thought that with the advance of science and technology, the particulars of us humans would become ever less relevant. But we’ve discovered that that’s not true. And that in fact everything—even our physics—depends on how we humans happen to have sampled the ruliad.

Before our Physics Project we didn’t know if our universe really was computational. But now it’s pretty clear that it is. And from that we’re inexorably led to the ruliad—with all its vastness, so hugely greater than all the physical space in our universe.

So where will we go in the ruliad? Computational language is what lets us chart our path. It lets us humans define our goals and our journeys. And what’s amazing is that all the power and depth of what’s out there in the ruliad is accessible to everyone. One just has to learn to harness those computational superpowers. Which starts here. Our portal to the ruliad:

Remembering Doug Lenat (1950–2023) and His Quest to Capture the World with Logic

6 septembre 2023 à 00:23

Logic, Math and AI

In many ways the great quest of Doug Lenat’s life was an attempt to follow on directly from the work of Aristotle and Leibniz. For what Doug was fundamentally trying to do over the forty years he spent developing his CYC system was to use the framework of logic—in more or less the same form that Aristotle and Leibniz had it—to capture what happens in the world. It was a noble effort and an impressive example of long-term intellectual tenacity. And while I never managed to actually use CYC myself, I consider it a magnificent experiment—that if nothing else ultimately served to demonstrate the importance of building frameworks beyond logic alone in usefully representing and reasoning about the world.

Doug Lenat started working on artificial intelligence at a time when nobody really knew what might be possible—or even easy—to do. Was AI (whatever that might mean) just a clever algorithm—or a new type of computer—away? Or was it all just an “engineering problem” that simply required pulling together a bigger and better “expert system”? There was all sorts of mystery—and quite a lot of hocus pocus—around AI. Did the demo one was seeing actually prove something, or was it really just a trivial (if perhaps unwitting) cheat?

I first met Doug Lenat at the beginning of the 1980s. I had just developed my SMP (“Symbolic Manipulation Program”) system, that was the forerunner of Mathematica and the modern Wolfram Language. And I had been quite exposed to commercial efforts to “do AI” (and indeed our VCs had even pushed my first company to take on the dubious name “Inference Corporation”, complete with a “=>” logo). And I have to say that when I first met Doug I was quite dismissive. He told me he had a program (that he called “AM” for “Automated Mathematician”, and that had been the subject of his Stanford CS PhD thesis) that could discover—and in fact had discovered—nontrivial mathematical theorems.

“What theorems?” I asked. “What did you put in? What did you get out?” I suppose to many people the concept of searching for theorems would have seemed like something remarkable, and immediately exciting. But not only had I myself just built a system for systematically representing mathematics in computational form, I had also been enumerating large collections of simple programs like cellular automata. I poked at what Doug said he’d done, and came away unconvinced. Right around the same time I happened to be visiting a leading university AI group, who told me they had a system for translating stories from Spanish into English. “Can I try it?” I asked, suspending for a moment my feeling that this sounded like science fiction. “I don’t really know Spanish”, I said, “Can I start with just a few words?” “No”, they said, “the system works only with stories.” “How long does a story have to be?” I asked. “Actually it has to be a particular kind of story”, they said. “What kind?” I asked. There were a few more iterations, but eventually it came out: the “system” translated one particular story from Spanish into English! I’m not sure if my response included an expletive, but I wondered what kind of science, technology, or anything else this was supposed to be. And when Doug told me about his “Automated Mathematician”, this was the kind of thing I was afraid I was going to find.

Years later, I might say, I think there’s something AM could have been trying to do that’s valid, and interesting, if not obviously possible. Given a particular axiom system it’s easy to mechanically generate infinite collections of “true theorems”—that in effect fill metamathematical space. But now the question is: which of these theorems will human mathematicians find “interesting”? It’s not clear how much of the answer has to do with the “social history of mathematics”, and how much is more about “abstract principles”. I’ve been studying this quite a bit in recent years (not least because I think it could be useful in practice)—and have some rather deep conclusions about its relation to the nature of mathematics. But I now do wonder to what extent Doug’s work from all those years ago might (or might not) contain heuristics that would be worth trying to pursue even now.

CYC

I ran into Doug quite a few times in the early to mid-1980s, both around a company called Thinking Machines (to which I was a consultant) and at various events that somehow touched on AI. There was a fairly small and somewhat fragmented AI community in those days, with the academic part in the US concentrated around MIT, Stanford and CMU. I had the impression that Doug was never quite at the center of that community, but was somehow nevertheless a “notable member”, who—particularly with his work being connected to math—was seen as “doing upscale things” around AI.

In 1984 I wrote an article for a special issue of Scientific American on “computer software” (yes, software was trendy then). My article was entitled “Computer Software in Science and Mathematics”, and the very next article was by Doug, entitled “Computer Software for Intelligent Systems”. The summary at the top of my article read: “Computation offers a new means of describing and investigating scientific and mathematical systems. Simulation by computer may be the only way to predict how certain complicated systems evolve.” And the summary for Doug’s article read: “The key to intelligent problem solving lies in reducing the random search for solutions. To do so intelligent computer programs must tap the same underlying ‘sources of power’ as human beings”. And I suppose in many ways both of us spent most of our next four decades essentially trying to fill out the promise of these summaries.

A key point in Doug’s article—with which I wholeheartedly agree—is that to create something one can usefully identify as “AI”, it’s essential to somehow have lots of knowledge of the world built in. But how should that be done? How should the knowledge be encoded? And how should it be used?

Doug’s article in Scientific American illustrated his basic idea:

Click to enlarge

Encode knowledge about the world in the form of statements of logic. Then find ways to piece together these statements to derive conclusions. It was, in a sense, a very classic approach to formalizing the world—and one that would at least in concept be familiar to Aristotle and Leibniz. Of course it was now using computers—both as a way to store the logical statements, and as a way to find inferences from them.

At first, I think Doug felt the main problem was how to “search for correct inferences”. Given a whole collection of logical statements, he was asking how these could be knitted together to answer some particular question. In essence it was just like mathematical theorem proving: how could one knit together axioms to make a proof of a particular theorem? And especially with the computers and algorithms of the time, this seemed like a daunting problem in almost any realistic case.

But then how did humans ever manage to do it? What Doug imagined was that the critical element was heuristics: strategies for guessing how one might “jump ahead” and not have to do the kind of painstaking searches that systematic methods seemed to imply would be needed. Doug developed a system he called EURISKO that implemented a range of heuristics—that Doug expected could be used not only for math, but basically for anything, or at least anything where human-like thinking was effective. And, yes, EURISKO included not only heuristics, but also at least some kinds of heuristics for making new heuristics, etc.

But OK, so Doug imagined that EURISKO could be used to “reason about” anything. So if it had the kind of knowledge humans do, then—Doug believed—it should be able to reason just like humans. In other words, it should be able to deliver some kind of “genuine artificial intelligence” capable of matching human thinking.

There were all sorts of specific domains of knowledge to consider. But Doug particularly wanted to push in what seemed like the most broadly impactful direction—and tackle the problem of commonsense knowledge and commonsense reasoning. And so it was that Doug began what would become a lifelong project to encode as much knowledge as possible in the form of statements of logic.

In 1984 Doug’s project—now named CYC—became a flagship part of MCC (Microelectronics and Computer Technology Corporation) in Austin, TX—an industry-government consortium that had just been created to counter the perceived threat from the Japanese “Fifth Generation Computer Project”, that had shocked the US research establishment by putting immense resources into “solving AI” (and was actually emphasizing many of the same underlying rule-based techniques as Doug). And at MCC Doug had the resources to hire scores of people to embark on what was expected to be a few thousand person-years of effort.

I didn’t hear much about CYC for quite a while, though shortly after Mathematica was released in 1988 Marvin Minsky mused to me about how it seemed like we were doing for math-like knowledge what CYC was hoping to do for commonsense knowledge. I think Marvin wasn’t convinced that Doug had the technical parts of CYC right (and, yes, they weren’t using Marvin’s theories as much as they might). But in those years Marvin seemed to feel that CYC was one of the few AI projects going on that actually made any sense. And indeed in my archives I find a rather charming email from Marvin in 1992, attaching a draft of a science fiction novel (entitled The Turing Option) that he was writing with Harry Harrison, which contained mention of CYC:

June 19, 2024

When Brian and Ben reached the lab, the computer was running
but the tree-robot was folded and motionless. “Robin,
activate.”

“Robin will have to use different concepts of progress for
different kinds of problems. And different kinds of subgoals
for reducing those different kinds of differences.”

“Won’t that require enormous amounts of knowledge?”

“It will indeed—and that’s one reason human education takes
so long. But Robin should already contain a massive amount of
just that kind of information—as part of his CYC-9 knowledge-
base.”

“There now exists a procedural model for the behavior of a
human individual, based on the prototype human described in
section 6.001 of the CYC-9 knowledge base. Now customizing
parameters on the basis of the example person Brian Delaney
described in the employment, health, and security records of
Megalobe Corporation.”

A brief silence ensued. Then the voice continued.

“The Delaney model is judged as incomplete as compared to those
of other persons such as President Abraham Lincoln, who has
3596.6 megabytes of descriptive text, or Commander James
Bond, who has 16.9 megabytes.”

Later, one of the novel’s characters observes: “Even if we started with nothing but the
old Lenat–Haase representation-languages, we’d still be far ahead of what any animal ever evolved.” (Ken Haase was a student of Marvin’s who critiqued and extended Doug’s work on heuristics.)

I was exposed to CYC again in 1996 in connection with a book called HAL’s Legacy—to which both Doug and I contributed—published in honor of the fictional birthday of the AI in the movie 2001. But mostly AI as a whole was in the doldrums, and almost nobody seemed to be taking it seriously. Sometimes I would hear murmurs about CYC, mostly from government and military contacts. Among academics, Doug would occasionally come up, but rather cruelly he was most notable for his name being used for a unit of “bogosity”—the lenat—of which it was said that “Like the farad it is considered far too large a unit for practical use, so bogosity is usually expressed in microlenats”.

Doug Meets Wolfram|Alpha

Many years passed. I certainly hadn’t forgotten Doug, or CYC. And a few times people suggested connecting CYC in some way to our technology. But nothing ever happened. Then in the spring of 2009 we were nearing the first release of Wolfram|Alpha, and it seemed like I finally had something that I might meaningfully be able to talk to Doug about.

I sent a rather tentative email:

Subject: something you might find interesting…
Date: Thu, 05 Mar 2009 11:15:04 -0500
From: Stephen Wolfram
To: Doug Lenat

We’re in the final stages of a rather large project that I think relates to
some of your interests.

I just made a small blog post about it:

http://blog.wolfram.com/2009/03/05/wolframalpha-is-coming/

I’d be pleased to give you a webconference demo if you’re interested.

I hope you’ve been well all these years.

— Stephen

Doug quickly responded:

Subject: Re: something you might find interesting…
Date: Thu, 5 Mar 2009 13:23:31 -0600
From: Doug Lenat
To: Stephen Wolfram

Hi, Stephen.

You have become a master of understatement! This certainly
does relate to the 1000 person-years we’ve spent building Cyc’s ontology,
knowledge base, and inference engines, over the last 25 years. I’d very
much like to see a webconference demo, so we identify the opportunities for
synergy.

Regards
Doug

It was definitely a “you’re on my turf” kind of response. And I wasn’t sure what to expect from Doug. But a few days later we had a long call with Doug and some of the senior members of what was now the Cycorp team. And Doug did something that deeply impressed me. Rather than for example nitpicking that Wolfram|Alpha was “not AI” he basically just said “We’ve been trying to do something like this for years, and now you’ve succeeded”. It was a great—and even inspirational—show of intellectual integrity. And whatever I might think of CYC and Doug’s other work (and I’d never formed a terribly clear opinion), this for me put Doug firmly in the category of people to respect.

Doug wrote a blog post entitled “I was positively impressed with Wolfram Alpha”, and immediately started inviting us to various AI and industry-pooh-bah events to which he was connected.

Doug seemed genuinely pleased that we had made such progress in something so close to his longtime objectives. I talked to him about the comparison between our approaches. He was just working with “pure human-like reasoning”, I said, like one would have had to do in the Middle Ages. But, I said, “In a sense we cheated”. Because we used all the things that got invented in modern times in science and math and so on. If he wanted to work out how some mechanical system would behave, he would have to reason through it: “If you push this down, that pulls up, then this rolls”, etc. But with what we’re doing, we just have to turn everything into math (or something like it), then systematically solve it using equations and so on.

And there was something else too: we weren’t trying to use just logic to represent the world, we were using the full power and richness of computation. In talking about the Solar System, we didn’t just say that “Mars is a planet contained in the Solar System”; we had an algorithm for computing its detailed motion, and so on.

Doug and CYC had also emphasized the scraps of knowledge that seem to appear in our “common sense”. But we were interested in systematic, computable knowledge. We didn’t just want a few scattered “common facts” about animals. We wanted systematic tables of properties of millions of species. And we had very general computational ways to represent things: not just words or tags for things, but systematic ways to capture computational structures, whether they were entities, graphs, formulas, images, time series, or geometrical forms, or whatever.

I think Doug viewed CYC as some kind of formalized idealization of how he imagined human minds work: providing a framework into which a large collection of (fairly undifferentiated) knowledge about the world could be “poured”. At some level it was a very “pure AI” concept: set up a generic brain-like thing, then “it’ll just do the rest”. But Doug still felt that the thing had to operate according to logic, and that what was fed into it also had to consist of knowledge packaged up in the form of logic.

But while Doug’s starting points were AI and logic, mine were something different—in effect computation writ large. I always viewed logic as something not terribly special: a particular formal system that described certain kinds of things, but didn’t have any great generality. To me the truly general concept was computation. And that’s what I’ve always used as my foundation. And it’s what’s now led to the modern Wolfram Language, with its character as a full-scale computational language.

There is a principled foundation. But it’s not logic. It’s something much more general, and structural: arbitrary symbolic expressions and transformations of them. And I’ve spent much of the past forty years building up coherent computational representations of the whole range of concepts and constructs that we encounter in the world and in our thinking about it. The goal is to have a language—in effect, a notation—that can represent things in a precise, computational way. But then to actually have the built-in capability to compute with that representation. Not to figure out how to string together logical statements, but rather to do whatever computation might need to be done to get an answer.

But beyond their technical visions and architectures, there is a certain parallelism between CYC and the Wolfram Language. Both have been huge projects. Both have been in development for more than forty years. And both have been led by a single person all that time. Yes, the Wolfram Language is certainly the larger of the two. But in the spectrum of technical projects, CYC is still a highly exceptional example of longevity and persistence of vision—and a truly impressive achievement.

Later Years

After Wolfram|Alpha came on the scene I started interacting more with Doug, not least because I often came to the SXSW conference in Austin, and would usually make a point of reaching out to Doug when I did. Could CYC use Wolfram|Alpha and the Wolfram Language? Could we somehow usefully connect our technology to CYC?

When I talked to Doug he tended to downplay the commonsense aspects of CYC, instead talking about defense, intelligence analysis, healthcare, etc. applications. He’d enthusiastically tell me about particular kinds of knowledge that had been put into CYC. But time and time again I’d have to tell him that actually we already had systematic data and algorithms in those areas. Often I felt a bit bad about it. It was as if he’d been painstakingly planting crops one by one, and we’d come through with a giant industrial machine.

In 2010 we made a big “Timeline of Systematic Data and the Development of Computable Knowledge” poster—and CYC was on it as one of the six entries that began in the 1980s (alongside, for example, the web). Doug and I continued to talk about somehow working together, but nothing ever happened. One problem was the asymmetry: Doug could play with Wolfram|Alpha and Wolfram Language any time. But I’d never once actually been able to try CYC. Several times Doug had promised API keys, but none had ever materialized.

Eventually Doug said to me: “Look, I’m worried you’re going to think it’s bogus”. And particularly knowing Doug’s history with alleged “bogosity” I tried to assure him my goal wasn’t to judge. Or, as I put it in a 2014 email: “Please don’t worry that we’ll think it’s ‘bogus’. I’m interested in finding the good stuff in what you’ve done, not criticizing its flaws.”

But when I was at SXSW the next year Doug had something else he wanted to show me. It was a math education game. And Doug seemed incredibly excited about its videogame setup, complete with 3D spacecraft scenery. My son Christopher was there and politely asked if this was the default Unity scenery. I kept on saying, “Doug, I’ve seen videogames before; show me the AI!” But Doug didn’t seem interested in that anymore, eventually saying that the game wasn’t using CYC—though did still (somewhat) use “rule-based AI”.

I’d already been talking to Doug, though, about what I saw as being an obvious, powerful application of CYC in the context of Wolfram|Alpha: solving math word problems. Given a problem, say, in the form of equations, we could solve pretty much anything thrown at us. But with a word problem like “If Mary has 7 marbles and 3 fall down a drain, how many does she now have?” we didn’t stand a chance. Because to solve this requires commonsense knowledge of the world, which isn’t what Wolfram|Alpha is about. But it is what CYC is supposed to be about. Sadly, though, despite many reminders, we never got to try this out. (And, yes, we built various simple linguistic templates for this kind of thing into Wolfram|Alpha, and now there are LLMs.)

Independent of anything else, it was impressive that Doug had kept CYC and Cycorp running all those years. But when I saw him in 2015 he was enthusiastically telling me about what I told him seemed to me to be a too-good-to-be-true deal he was making around CYC. A little later there was a strange attempt to sell us the technology of CYC, and I don’t think our teams interacted again after that.

I personally continued to interact with Doug, though. I sent him things I wrote about the formalization of math. He responded pointing me to things he’d done on AM. On the tenth anniversary of Wolfram|Alpha Doug sent me a nice note, offering that “If you want to team up on, e.g., knocking the Winograd sentence pairs out of the park, let me know.” I have to say I wondered what a “Winograd sentence pair” was. It felt like some kind of challenge from an age of AI long past (apparently it has to do with identifying pronoun reference, which of course has become even more difficult in modern English usage).

And as I write this today, I realize a mistake I made back in 2016. I had for years been thinking about what I’ve come to call “symbolic discourse language”—an extension of computational language that can represent “everyday discourse”. And—stimulated by blockchain and the idea of computational contracts—I finally wrote something about this in 2016, and I now realize that I overlooked sending Doug a link to it. Which is a shame, because maybe it would have finally been the thing that got us to connect our systems.

And Now There Are LLMs

Doug was a person who believed in formalism, particularly logic. And I have the impression that he always considered approaches like neural nets not really to have a chance of “solving the problem of AI”. But now we have LLMs. So how do they fit in with things like the ideas of CYC?

One of the surprises of LLMs is that they often seem, in effect, to use logic, even though there’s nothing in their setup that explicitly involves logic. But (as I’ve described elsewhere) I’m pretty sure what’s happened is that LLMs have “discovered” logic much as Aristotle did—by looking at lots of examples of statements people make and identifying patterns in them. And in a similar way LLMs have “discovered” lots of commonsense knowledge, and reasoning. They’re just following patterns they’ve seen, but—probably in effect organized into what I’ve called a “semantic grammar” that determines “laws of semantic motion”—that’s enough to often achieve some fairly impressive commonsense-like results.

I suspect that a great many of the statements that were fed into CYC could now be generated fairly successfully with LLMs. And perhaps one day there’ll be good enough “LLM science” to be able to identify mechanisms behind what LLMs can do in the commonsense arena—and maybe they’ll even look a bit like what’s in CYC, and how it uses logic. But in a sense the very success of LLMs in the commonsense arena strongly suggests that you don’t fundamentally need deep “structured logic” for that. Though, yes, the LLM may be immensely less efficient—and perhaps less reliable—than a direct symbolic approach.

It’s a very different story, by the way, with computational language and computation. LLMs are through and through based on language and patterns to be found through it. But computation—as it can be accessed through structured computational language—is something very different. It’s about processes that are in a sense thoroughly non-human, and that involve much deeper following of general formal rules, as well as much more structured kinds of data, etc. An LLM might be able to do basic logic, as humans have. But it doesn’t stand a chance on things where humans have had to systematically use formal tools that do serious computation. Insofar as LLMs represent “statistical AI”, CYC represents a certain level of “symbolic AI”. But computational language and computation go much further—to a place where LLMs can’t and shouldn’t follow, and should just call them as tools.

Doug always seemed to have a very optimistic view of the promise of AI. In 2013 he wrote to me:

Of course you are coming at this from the opposite end of the Chunnel than
we are, but you’re proceeding, frankly, much more rapidly toward us than we
are toward you. I probably appreciate the significance of what you’ve
accomplished more than almost anyone else: when your and our approaches do
meet up, the combination will be the existence of real AI on Earth. I
think that’s the main motivation in your life, as it is in mine: to live to
see real AI, with the obvious sweeping change in all aspects of life when
there is (i) cradle-to-grave 24×7 Aristotle mentoring and advising for
every human being and, in effect, (ii) a Land of Faerie intelligence
effectively present [e.g., that one can converse with] in every door, floor
tile,…every tangible object above a certain microscopic size.) And to
live to see and be users ourselves in an era of massively amplified human
intelligence …

The last mail I received from Doug was on January 10, 2023—telling me that he thought it was great that I was talking about connecting our tech to ChatGPT. He said, though, that he found it “increasingly worrisome that these models train on CONVINCINGNESS rather than CORRECTNESS”, then gave an example of ChatGPT getting a math word problem wrong.
His email ended:

Yes, let’s chat again at your convenience… it bothers both of us, I
believe, that our systems aren’t leveraging each other! That just bothers
me more and more as I get old (not just older).

Sadly we never did chat again. We now have a team actively working on symbolic discourse language, and just last week I mentioned CYC to them—and lamented that I’d never been able to try it. And then on Friday I heard that Doug had died. A remarkable pioneer of AI who steadfastly pursued his vision over the whole course of his career, and was taken far too soon.

Remembering the Improbable Life of Ed Fredkin (1934–2023) and His World of Ideas and Stories

Par : Mark Long
22 août 2023 à 22:03

Programmer of the Universe

Click to enlarge

“OK, so let me tell you…” And so it would begin. A long and colorful story. An elaborate description of a wild idea. In the forty years I knew Ed Fredkin I heard countless wild ideas and colorful stories from him. He always radiated a certain adventurous joy—together with supreme, almost-childlike confidence. Ed was someone who wanted to independently figure things out for himself, and delighted in presenting his often somewhat-outlandish conclusions—whether about technology, science, business or the world—with dramatic showman-like panache.

In all the years I knew Ed, I’m not sure he ever really listened to anything I said (though he did use tools I built). He used to like to tell people I’d learned a lot from him. And indeed we had intellectual interests that should have overlapped. But in actuality our ways of thinking about them mostly didn’t connect much at all. But at a personal and social level it was still always a lot of fun being around Ed and being exposed to his unique intense opportunistic energy—with its repeating themes but ever-changing directions.

And there was one way in which Ed and I were very much aligned: both of our lives were deeply influenced by computers and computing. Ed had started with computers in 1956—as part of one of the very first cohorts of programmers. And perhaps on the basis of that experience, he would still, even at the end of his life, matter-of-factly refer to himself as “the world’s best programmer”. Indeed, so confident was he of his programming prowess that he became convinced that he should in effect be able to write a program for the universe—and make all of physics into a programming problem. It didn’t help that his knowledge of physics was at best spotty (and, for example, I don’t think he ever really learned calculus). But his almost lifelong desire to “program physics” did successfully lead him to the concept of reversible logic, and to what’s now called the “Fredkin gate”. But it also led him to the idea that the universe must be a giant cellular automaton—whose program he could invent.

I first met Ed in 1982—on an island in the Caribbean he had bought with money from taking public a tech company he’d founded. The year before, I had started studying cellular automata, but, unlike Ed, I wasn’t trying to “program” them—to be the universe or anything else. Instead, I was mostly doing what amounted to empirical science, running computer experiments to see what they did, and treating them as part of a computational universe of possible programs “out there to explore”. It wasn’t a methodology I think Ed ever really understood—or cared about. He was a programmer (and inventor), not an empirical scientist. And he was convinced—like a modern analog of an ancient Greek philosopher—that by pure thought he could come up with the whole “clockwork” of the universe.

Central to his picture was the idea that at the bottom of everything was a cellular automaton, with its grid of cells somehow laid out in space. I told Ed countless times that what was known from twentieth-century physics implied this really couldn’t be how things worked at a fundamental level. I tried to interest Ed in my way of using cellular automata. But Ed wasn’t interested. He was going for what he saw as the big prize: using them to “construct the universe”.

Every few years Ed would tell me he’d made progress—and rather dramatically say things like that he’d “found the electron”. I’d politely ask for details. Then start pointing out that it couldn’t work that way. But soon Ed would be telling a story or talking about some completely different idea—about technology, business or something else.

By the mid-1980s I’d discovered a lot about cellular automata. And I always felt a bit embarrassed by Ed’s attempt to use them in what seemed to me like a very naive way for fundamental physics—and I worried (as did happen a few times) that people would dismiss my efforts by identifying them with his.

My own career had begun in the 1970s with traditional fundamental physics. And while I didn’t think cellular automata as such could be directly applied to fundamental physics, I did think that the core computational phenomena I’d discovered through studying cellular automata might be very relevant. And then in the early 1990s I had an idea. In a cellular automaton, space has a fixed grid-like structure. But what if the structure of space is in fact dynamic, and everything in the universe emerges just from the dynamics of that structure? Finally I felt as if there might be a plausible computational foundation for fundamental physics.

I wrote about this in one chapter of my 2002 book A New Kind of Science. I don’t know if Ed ever read what I wrote, but in any case it didn’t seem to affect his idea that the universe was a cellular automaton—and to confuse things further, he told quite a few people that was what I was saying too. At first I found this frustrating—and upsetting—but eventually I realized it was just “Ed being Ed”, and there were still plenty of things to like about Ed.

Nearly twenty years passed. I would see Ed with some regularity. And sometimes I would mention physics. But Ed would just keep talking about his idea that the universe is a cellular automaton. And when we finally made the breakthrough that led in 2020 to our Physics Project it made me a little sad that I didn’t even try to explain it to Ed. The universe isn’t a cellular automaton. But it is computational. And I think that knowing this would have brought a certain intellectual closure to Ed’s long journey and aspirations around physics.

Ed might have considered physics his single most important quest. But Ed’s life as a whole was filled with a remarkably rich assortment of activities and interests. Computers. Inventions. Companies. Airplanes. MIT. His island. The Soviet Union. Not to mention people, like Marvin Minsky, John McCarthy and Richard Feynman (as well as Tom Watson, Richard Branson, and many more). And he would tell stories about all these people and things, and more. Sometimes (particularly later in his life) the stories would repeat. But with remarkable regularity Ed would surprise me with yet another—often at first hard-to-believe—story about a situation or topic that I had no idea he’d ever been involved in.

But what was the “whole Ed story”? I knew a lot of fragments, often quite colorful. But they didn’t seem to fit together into the narrative of a life. And now that Ed is sadly no longer with us, I decided I should really try to “understand Ed” and his story. A few times over the years I had made efforts to ask Ed for systematic historical accounts—and in 2014 I even recorded many hours of oral history with him. But there was clearly much more. And in writing this piece I found myself going through lots of documents and archives—and having quite a few conversations— and unearthing even yet more stories than I already knew. And in the end there’s a lot to say—and indeed this has turned into the most difficult and complicated biographical piece I’ve ever written. But I hope that everything I’ve assembled will help tell the often so-wild-you-can’t-make-this-stuff-up story of that most singular individual who I knew all those years.

The Beginning of the Story

Ed never said much to me about his early life. And in fact I think it was only in writing this piece that I even learned he’d grown up in Los Angeles (specifically, East Hollywood). His parents were both (Jewish) Russian immigrants (his father was born in St. Petersburg; his mother in Odessa; they met in LA). His father’s university engineering studies had been cut short by the Russian Revolution, and he now had a one-man wholesale electronic parts business. His mother had in her youth been trained as a concert pianist, and died when Ed was 11, leaving a somewhat fragmented family situation. Ed had a half-sister, 14 years older than him, a brother 6 years older, and a sister a year older. As he told it in later oral histories, he got interested in both machines and money very early, repairing appliances for a fee even as a tween, and soon learning about the idea of owning stock in companies.

But Ed Fredkin’s first piece of public visibility seems to have come in 1948, when he was 13 years old—and it reminds me so much of many of Ed’s later “self-imposed” adventures. There was at that time an exhibition of historic US documents traveling around the country on a train named the Freedom Train. And when the train came to Los Angeles, the young Ed Fredkin decided he had to be the first person to see it:

Click to enlarge

The Los Angeles Times published his account of his adventure—a younger but “quintessentially Ed” story:

Click to enlarge

Ed’s record in high school was at best spotty. But as he tells it, he figured out very early a system for improving the odds in multiple-choice tests, and for example in 9th grade got a top score on a newly instituted (multiple-choice) California-wide IQ test. At the end of high school, Ed applied to Caltech (which was only 13 miles away from where he lived), and largely on the basis of his test scores, was admitted. He ended up spending time working various jobs to support himself, didn’t do much homework, and by his sophomore year—before having to pick a major—dropped out. In 2015 Ed told me a nice story about his time at Caltech:

In 1952–53, I was a student in Linus Pauling’s class where he lectured Freshman Chemistry at Caltech. After class, one day, I asked Pauling “What is a superconductor at the highest known temperature?” Pauling immediately replied “Niobium Nitride, 18 Kelvin”. I was puzzled because I had never heard of Niobium, so I looked it up and, with some difficulty found a reference that defined it as a European name for the metal Columbium.

Later that same day, reading a Pasadena newspaper, I saw an article about Pauling: It announced that Pauling had just returned from Europe (London is what I recall) where Pauling, as Chairman of the International Committee on the naming of the elements, had decided that henceforth the metal Columbium would be renamed Niobium.

I recently looked into that matter and discovered that evidently that renaming was part of a USA–Europe Compromise… In Europe it had been Wolfram and Niobium, in the USA it had been Tungsten and Columbium.

Europe got its way re Niobium and the USA got its way re Tungsten… Perhaps it was a flip of a coin? Someone might know.

As a Wolfram, I thought you might be interested (and, of course, perhaps all this is old hat to you…).

(For what it’s worth, I actually didn’t know this “Wolfram story”, though the details weren’t quite as dramatic as Ed said: the “niobium” decision was actually made in 1949, without Pauling specifically involved, though Pauling did indeed travel to London just before the beginning of the 1952 school year.)

With his interest in machinery, Ed had always been keen on cars, and in his freshman year at Caltech, he also decided to learn to fly a plane. Ed’s older brother, Norman, had joined the Air Force five years earlier. And when he left Caltech—in 1954 at age 19—Ed joined the Air Force too. (If he hadn’t done that, he would have been drafted into the Army.) Ed’s brother Norman (who would spend his whole career in aviation) had been involved in the Korean War, particularly doing aerial reconnaissance—here pictured with his plane (and, no, there don’t seem to be any Air Force pictures of Ed himself):

Click to enlarge

By the time Ed joined the Air Force, the Korean War was over. Ed was assigned to an airbase in Arizona, and by the summer of 1955 he had qualified as a fighter pilot. Ed was never officially a “test pilot”, but he told me stories about figuring out how to take his plane higher than anyone else—and achieving weightlessness by flying his plane in a perfect free-fall trajectory by maintaining an eraser floating in midair in front of him.

By 1956 Ed had been grounded from flying as a result of asthma, and was now at an airbase in Florida as an “intercept controller”—essentially an air traffic controller responsible for guiding fighters to intercept bombers. It was a time when the Air Force was developing the SAGE (Semi-Automatic Ground Environment) air defense system—a huge project whose concept was to use computers to coordinate data from many radars so as to be able to intercept Soviet bombers that might attack the US (cf. Dr. Strangelove, etc.). The center of SAGE development was Lincoln Lab (then part of MIT) in Lexington, MA—with IBM providing computers, Bell (AT&T) providing telecommunications, RAND providing algorithms, etc. And in mid-1956 the Air Force sent a group—including Ed—to test the next phase of SAGE. But as Ed tells it, they were soon informed that actually there would be a one-year delay.

At the time, the SAGE project was busily trying to train people about computers, and some people from the Air Force stayed in the Boston area to participate in this. As Ed tells it, however, he was the only one who didn’t drop out of the training—and over the course of a year it taught him “much of what was then known about computer programming and computer hardware design”. There were at the time only a few hundred people in the world who could call themselves programmers. And Ed was now one of them. (Perhaps he was even “the world’s best”.)

Computers!

Having learned to program, Ed remained at Lincoln Lab, paid by the Air Force, doing what amounted to computational “odd jobs”. Often this had to do with connecting systems together, or coming up with “clever hacks” to overcome particular system limitations. Occasionally it was a little more algorithmic—like when Sputnik was launched in 1957, and Ed got pulled into a piece of “emergency programming” for orbit calculations.

Ed told many stories about “hacking” the bureaucracy at the Air Force (being given a “Secret” stamp so he could read his own documents; avoiding being sent for a year to the Canadian Arctic by finding a loophole associated with his wife being pregnant, etc.)—and in 1958 he left the Air Force (though he would remain a captain in the reserves for many years), but stayed on at Lincoln Lab. Officially he was there as an “administrative assistant”, because—without a degree—that was all they could offer him. But by then he was becoming known as a “computer person”—with lots of ideas. He wanted to start his own company. And (as he tells it) the very first potential customer he visited was an MIT-spinoff acoustics firm called Bolt Beranek & Newman (BBN). And the person he saw there was their “vice president of engineering psychology”—a certain J. C. R. “Lick” Licklider—who persuaded Ed to join BBN to “teach them about computers”.

It didn’t really come to light until he was at BBN, but while at Lincoln Lab Ed had made what would eventually become his first lasting contribution to computer science. He thought of it as a new way of storing textual information in a computer, and he called it “TRIE memory” (after “reTRIEval”). Nowadays we’d call it the trie (or prefix tree) data structure. Here it is for some common words in English made from the letters of “wolf”:

Licklider persuaded Ed to write a paper about tries—which appeared in 1960, and for a couple of decades was essentially Ed’s only academic-style publication:

Click to enlarge

The paper has a pretty clear description of tries, even with some nice diagrams:

Click to enlarge

Even in analyzing the performance of tries, there was only the faintest hint of math in the paper—though Ed realized (probably with input from Licklider) that the efficiency of tries would depend on the Shannon-style redundancy of what they were storing, and he ran Monte Carlo simulations to investigate this:

Click to enlarge

(He explains: “The test program was written in FORTRAN for the IBM 709. The program is composed of 42 subroutines, of which 19 were coded specially for this program and 23 were taken from the library.”)

Tries didn’t make a splash when Ed first introduced them—not least because computers didn’t really have the memory then to make use of them. I think I first heard about them in the late 1970s in connection with spellchecking, and nowadays they’re widely used in lots of text search, bioinformatics and other applications.

Ed had apparently first started talking about tries when he was still in the Air Force. As he explained it to me in 2014:

The Air Force [people] had no idea [what I was talking about]. But I kept on [saying] “I need to find someone who knows something about this that can critique it for me.” And someone says to me, “There’s a guy at MIT who deals in something similar, he calls it lists”. And that was John McCarthy. So, I call up, I get a secretary and, you know, I make a date, and I go to MIT and in building 56 with the computation center, I go to his office and the secretary says he’s somewhere out in the hall. I see some guy wandering back and forth. I go up and say, “You John McCarthy?” He says, “Yes.” So, I say, “I’ve had this idea—” I can’t remember if I was in uniform or not; I might’ve been. I said, “I had this idea, and I’ve written a program and tested it. And might you take a look?” Then he takes this thing, and he starts to read it.

Then he did something that struck me as very weird. He turned around slowly and started walking away, he’s reading and walk, walk, walk, walk, stop. Turns around, walk, walk, walk, walk, back slowly, you know. Finally, he comes back and he stops and he reads and reads. And he’s obviously angry. And I thought, “This is weird.” I said “Does it make sense or anything?” He says, “Yes, it makes sense.” And I said, “Well, what’s up?” He says, “Well, I’ve had the same idea.” And I said, “Oh.” He says, “But I’ve never written it down.” And I said, “Oh, okay. So, do you think I ought to work on it or do something?” He says, “Yeah”. So, that’s how I met John McCarthy.

Ed remained friends with McCarthy for the rest of McCarthy’s life, and involved him in many of his endeavors. In 1956 McCarthy had been one of the organizers of the conference that coined the term “artificial intelligence”, and in 1958 McCarthy began the development of LISP (which was based on linked lists). I have to say I wish I’d known Ed’s story with McCarthy much earlier; I would have handled my own interactions with McCarthy differently—because, as it was, over the course of various encounters from 1981 to 2003 I never persisted very far beyond the curmudgeon stage.

Back around 1958, the circle of “serious computer people” in the Boston area wasn’t very large—and another was Marvin Minsky (who I knew for many years). Between Ed and Licklider, both McCarthy and Minsky became consultants at BBN, and all of them would have many interactions in the years to come.

But in late 1959 there was another entrant in the Boston computer scene: the PDP-1 computer, designed by a certain Ben Gurley for a new company named Digital Equipment Corporation (DEC) that had essentially spun off from Lincoln Lab and MIT. BBN was the first customer for the PDP-1, and Ed was its anchor user:

Click to enlarge

John McCarthy had had the “theoretical” idea of timesharing, whereby multiple users could work on a single computer. Ed figured out how to make it practical on the PDP-1, in the process inventing what would now be called asynchronous interrupts (then the “sequence break system”). And so began a process which led BBN to become a significant force in computing, the creation of the internet, etc.

But in 1961, Ed and a certain Roland Silver, who also worked at BBN, decided to quit BBN—and, strangely enough, to move to Brazil, where they were enamored of the recently elected new president. But when that new president unexpectedly resigned, they abandoned their plan. And when BBN didn’t want them back, Ed decided to start a company, initially doing consulting for DEC. As Ed tells it, he and Roland Silver were such good friends and had so much they talked about that together they couldn’t get anything done, so they decided they’d better split up.

As I was writing this piece, I decided to look up more about Roland Silver—who I found out had been a college roommate of Marvin Minsky’s at Harvard, and had had a long career in math, etc. at MITRE (the holding company for Lincoln Lab). But I also remembered that many years ago I’d received letters and a rather new-age newsletter from a certain “Rollo Silver”:

Click to enlarge

Could it be the same person? Yes! And in my archives I also found an ad:

Click to enlarge

Some time after my work on cellular automata in the 1980s, Roland Silver—together with my longtime friend Rudy Rucker—started a newsletter about cellular automata, notably not mentioning Ed, but including a colorful bio for Silver:

Click to enlarge

“Triple-I” (III)

But back to Ed and his story. It was 1961, and Ed had quit his job at BBN. In 1957, he’d met on a Cape Cod beach a woman from Western Massachusetts named Dorothy Abair (who was at the time working at a beauty salon)—and six weeks later they’d married, and now had a 3-year-old daughter. Ed had already lined up some consulting with DEC, and as Ed tells it, with a little “hacking” of bank loans, etc. he was able to officially start Information International Incorporated (III)—with a tiny office in Maynard, MA (home of DEC). But then, one day he gets a call from the Woods Hole Oceanographic Institute. He drives down to Woods Hole with a certain Henry Stommel—an oceanography professor at Harvard—who tells him about a “vortex ocean model”, and asks Ed if he can program it on a PDP-1 so that it displays ocean currents on a screen. And the result is that III soon has a contract for $10k (about $100k today) to do this.

I might add a small footnote here. Years later I was talking to Ed about the origins of cellular automata, and he tells me that a certain Henry Stommel had told him that there were cellular automaton models of sand dunes from the 1930s. At the time—before the web—I couldn’t easily track down who Henry Stommel was (and I had no idea how Ed knew him), and to this day I don’t know what those sand dune models might have been.

But in any case, Ed’s interaction with Woods Hole led to what became III’s first major business: digital reading of film. As Ed tells it:

At Woods Hole … they had these meters which would measure how fast the ocean current was going and which way—and recorded it on 16 mm film with little tiny lights and a little fiber optic thing. And they had built a machine to read that film. I looked at the machine and said “That’ll never work”. And they said “Who are you? Of course it’ll work”, and so on, so forth. OK, so some months later they call me up and say it didn’t work.

I have to tell you this but this is insanely funny. So I decide I’m going to make a film reader and here’s how I’m going to do it. I knew there was a 16 mm projector you could rent from a company and you could stop it and then say “Advance one frame” by clicking and it would just advance one frame at a time. So I thought: say I take the lightbulb out and put a photomultiplier in and point it at the screen of the computer. Then light will come from the screen, go through the lens and be focused on the film, and some would go through the film to the photomultiplier and I would be able to tell how much light got through. And we could write a program to do the rest.

That was my idea, OK.

So not having any money, we rented that projector and I got Digital (DEC) to let me use their milling machine and I bought the photomultiplier tube, and I got Ben Gurley to design the circuitry and connect it to the computer. But there was one more thing. The photomultiplier tube was like a vacuum tube but it had like 16 pins and a very odd connector that no one had. But I thought “Lincoln Labs has parts for everything in their electronics warehouse”. So I called someone I used to work with there, and said “Look, do me a favor and sneak into the parts area, take that part and just give it to me. I’ve ordered one but I’m not going to get it for a while and when I get it I’ll give it to you and you can put it back so it’s not actually a theft.” And he said “OK, I’ll do it” but he asked me why I wanted it and I told him “Well, I’m doing this stuff for Woods Hole to read some film with a computer”.

OK, so he gave me the part and we get it going right away and we’re reading the film, and that solved the problem. But meanwhile this very funny thing happened. Someone from Lincoln Labs found out about all this and said “Hey, you’re reading some kind of film. Is that what you used that thing for?” And I said “Yeah”. And they said “Well, we tried to read some films so we built a gadget and did the same thing you did: we pointed it at the screen of the computer, but we can’t make the software work”. And I said “OK, well, come down and tell me about it”. So they come down and what happens is this. There’s some army people and they have a radar that’s looking at a missile coming in and records on film from an oscilloscope. And they asked could we read this. And to make a long story short they signed another contract….

The whole setup was eventually captured in a patent entitled simply “High-Speed Film Reading”:

Click to enlarge

And actually this wasn’t Ed’s first patent. That had been filed in 1960, while Ed was at BBN—and it was for a mechanical punched card sorter, with arrays of metal pins and the like, and no computer in evidence:

Click to enlarge

III ended up discovering that there were many applications—military and otherwise—for film readers. But their Woods Hole relationship led in another direction as well: computer graphics and data visualization. By 1963 there were perhaps 300,000 oceanographic stations recording their data on punched cards, and the idea was to take this data and produce from it a “computer-compiled oceanographic atlas”. The result was a paper:

Click to enlarge

And with statements like “Only a high-speed computer has the capacity and speed to follow the quickly shifting demands and questions of a human mind exploring a large field of numbers” the paper presented visualizations like:

Click to enlarge

These various developments put III in the center of the emerging field of film-meets-computers systems. The company grew, moving its center of operations to Los Angeles, not least to be near the Systems Development Corporation (SDC) which RAND had spun off as its software arm in response to the SAGE project.

But Ed was always having new ideas for III, and defining new directions. Ed had brought Minsky and McCarthy into III as board members and consultants, and for example in 1964 III was proposing to SDC a project to make a new version of LISP (and, yes, with no obvious film-meets-computers applications). The proposal gives some insight into the state of III at the time. It says that “From a one-man operation [in 1962], I.I.I. has grown to the point where our gross volume of business for 1964 is in the neighborhood of $1 million [about $10 million today]”. It explains that III has four divisions: Mathematical and Programming Services, Behavioral Science, Operations, and “New York”. It goes on to list various things III is doing: (1) LISP; (2) Inductive Inference on Sequences; (3) Computer Time-Sharing; (4) Programmable Film Readers; (5) The World Oceanographic Data Display System; and (6) Computer Display Systems.

It’s certainly an eclectic collection, reflecting, as such things often do, the character of the company’s founder. From a modern perspective, one item that catches one’s attention is:

Click to enlarge

One can think of it as an early attempt at AI/machine learning—which 60 years later still hasn’t been solved. (GPT-4 says the next letter should be Q, not O.)

But distractions or not, it was a talented team that assembled at III—with lots of cross-fertilization with MIT. III’s business progressively grew, and perhaps it outgrew Ed—and in 1965 Ed stepped down as CEO. In 1968 he left entirely and (as we’ll discuss below) went to MIT, leaving III in the hands of Al Fenaughty, who, years later (and after nearly 30 years at III), would become the chairman of Yandex.

As someone who’s curious about the ways of company founders, I asked Ed many times about his departure from III. He usually just said: “I had a partner who died”. But it’s only now that I’ve pieced together, partly from my 2014 oral history with Ed, what happened. Ed described it to me as the greatest tragedy of his life.

Shortly after he set up III, Ed persuaded Ben Gurley (designer of the PDP-1) to leave DEC and join him at III. I think Ed had hoped to build computers at III, with Gurley as their designer. But on November 7, 1963, in Concord, MA, just a few miles from where I am as I write this, Ben Gurley was murdered—by a single revolver shot through his dining room window as he was about to sit down for dinner with his wife and 7 children. An engineer from DEC (and Lincoln Labs)—about whom Gurley had recently complained to the police—was arrested, and eventually convicted of the crime (after Ed hired a private detective to help). It later turned out that a few years earlier the same engineer was likely also responsible for shooting (though not killing) another engineer from DEC.

I had always assumed that Ed’s decision to leave III happened just after his “partner had died”. But I now realize that Gurley’s death early in the history of III caused III to go on its path of making things like film readers, rather than the DEC- or IBM-challenging computers I think Ed had hoped for.

Even after Ed left active management of III, he was still its chairman. And in late 1968 something would happen that would change his life forever. Taking tech companies public on the “over-the-counter” market had become a thing, and a broker offered to take III public. And on November 26, 1968, III filed its SEC paperwork:

Click to enlarge

III’s “principal product to date” is described as a “programmable film reader”, but the paperwork notes that as of October 31, 1968, the company has no film readers on order—though there are orders for its new microfilm reader, which it hasn’t delivered yet. It also says that proceeds from the offering will be used to fund its “proposed optical character recognition project”. But for our purposes what’s perhaps more significant is that the paperwork records that Ed owns 57.7% of the company, with the Edward Fredkin Charitable Foundation owning 0.4%.

On January 8, 1969, III went public, and Ed was suddenly, at least on paper, worth more than $10M (or more than $80M today). Two years later (perhaps as soon as a lockup period expired), Ed cashed out, with the SEC notice indicating that Ed would be “repaying personal indebtedness to a bank incurred by him for reasons unrelated to the company or its business” (presumably a loan he’d taken out before he could achieve liquidity):

Click to enlarge

So now Ed—at age 37—was wealthy. And in fact the money he made from III would basically last the rest of his life, even through a long sequence of subsequent business failures.

III’s OCR project was never a great success, but III became a key company in digital-to-film systems (relevant to both movies and printing), and in the early 1970s created some of the very first computer-generated special effects, that eventually made it into movies like Star Wars. III’s stock price hovered around $10 per share for years, and in 1996—after PostScript had pretty much taken the market for prepress printing systems—III was sold to Autologic for $35M in stock, then in 2001 Autologic was sold to Agfa for $42M.

The Island

When III went public in 1969 it was the height of the Cold War (which probably didn’t hurt III’s military sales). And many people—including Ed—thought World War III might be imminent. And so it was that in 1970 Ed decided to buy an island in the Caribbean, close enough to the tropics, he told me subsequently, that, he assumed (incorrectly according to current models), radioactive fallout from a nuclear war wouldn’t reach it.

Apparently Ed was sitting in a dentist’s office when he saw an “Island for Sale” ad in a newspaper. The seller was a shipwreck-scavenging treasure hunter named Bert Kilbride—sometimes called “the last pirate of the Caribbean”—who had started to develop the island (and for several years would manage it for Ed). It’s a fairly small island (about 125 acres, or 0.2 square miles)—in the British Virgin Islands. And its name is Mosquito Island (or sometimes, with some historical justification, Moskito Island). And when Ed bought it, it probably cost something under $1M. (Richard Branson bought the nearby but smaller Necker Island in 1978.)

I visited Ed’s island in January 1982—the first time I met Ed. And, yes, there was a certain “lair of a Bond villain” (think: Dr. No) vibe to the whole thing. Here are pictures I took from a boat leaving the island (notice the just-visible seaplane parked at the island):

Click to enlarge

There was a small resort (and restaurant) on the island, named Drake’s Anchorage (built by the previous owner):

Click to enlarge

And, yes, there were beaches on the island (though I myself have never been much of a beach-goer):

Click to enlarge

And, in keeping with the Bond vibe, there was a seaplane too:

Click to enlarge

There was one house on the island, here pictured from the plane (it so happened that when I visited the island, I was learning to fly small planes myself—so I was interested in the plane):

Click to enlarge

Visiting a nearby island—with its very rundown airport sign—gives some sense of the overall area:

Click to enlarge

Ed claimed it was difficult to run the resort on his island, not least because, he said, “the British Virgin Islands have the lowest average worker productivity in the world”. But he nevertheless, for example, had a functioning restaurant, and here I am there in 1982, along with Charles Bennett, about whom we’ll hear more later:

Click to enlarge

When people talked about Ed, his island was often mentioned, and it projected a general image of overall mystique and extreme wealth. In 1983 a movie called WarGames came out, featuring a reclusive military-oriented computer expert named “Professor Falken”—who had an island. Many people assumed Falken was based on Fredkin (and it now says so all over the internet). However, in writing this piece, I decided to find out what was actually true—so I asked one of the writers of the movie, Walter Parkes. He responded, and, yes, fact is often even stranger than fiction:

Unfortunately I can confirm that Ed was not the inspiration for Stephen Falken. The character was inspired by Steven [sic] Hawking. (Falken = Falcon = Hawking) The movie was first conceived to be about two characters, a young super-genius born into a family incapable of acknowledging his gifts, and a dying scientist in need of a protégé. In the first several drafts Falken was confined to a wheel-chair and was working on understanding the big bang, for which he had created a computer simulation. Little known fact—while writing the character, we had one person in mind to play the role: John Lennon, who was murdered shortly before we finished the script.

(By the way, in a moment of “fact follows fiction”, WarGames featured a computer with lots of flashing lights. I happened to see the movie with Danny Hillis, and as we were walking out of the movie, I said to Danny “Perhaps your computer should have flashing lights too”. And indeed flashing lights became a signature feature of Danny’s Connection Machine computer, as later seen in movies like Jurassic Park.)

Project MAC

After he left III in 1968, Ed’s next stop would be MIT, and specifically Project MAC (the “Multiple Access Computer” Project). But actually Ed had already been involved much earlier with Project MAC. In many ways the project was a follow-on to what Ed had been doing at BBN on timesharing.

In 1963 Ed wrote a long survey article on timesharing:

Click to enlarge

The introduction contains a rather charming window onto the view of computers at the time:

Click to enlarge

And the ads interspersed through the article give a further sense of the time:

Click to enlarge

As illustrations of what can be done with an interactive timeshared computer, there’s a picture from Ed’s vortex ocean simulation—as well as an example of an online “book” about LISP:

Click to enlarge

And, yes, already a kind of “cloud computing” story:

Click to enlarge

There’s also a description of Project MAC—that had just been funded by the Advanced Research Projects Agency (now DARPA). The article said that the “MAC” stood either for “Multiple Access Computer” or “Machine-Aided Cognition”. It included various sections on what might be possible with timesharing:

Click to enlarge

The main text of the article ends with a rousing (?) vision of AI taking over from humans (and, yes, even though this is from 60 years ago it’s not so different from what at least some people might say about the “AI future” today):

Click to enlarge

But there’s a curious piece of backstory to Project MAC—from 1961—that appears as a footnote to Ed’s article:

Click to enlarge

Ed told me versions of this story many times. McCarthy had failed to get tenure at MIT, and was looking for another job. (Yes, in retrospect this seems remarkable given all the things he’d already done by then. But those things were computer science—and MIT didn’t yet have a CS department; McCarthy was in the EE department.) Ed, Minsky and McCarthy were going to an SDC meeting in Los Angeles, and while he was out there McCarthy was going to interview at Caltech (his undergraduate alma mater). They had a free evening, and Ed suggested they meet “someone interesting”. Ed remembered Linus Pauling from his time at Caltech. But Pauling wasn’t in. So Minsky suggested they call Richard Feynman. And he was in, and invited them over to his house.

Feynman apparently showed them things like his nanotech-inspiring tiny motor, etc., but somehow the discussion shifted to AI. And Minsky mentioned work a student of his was doing on the “AI problem” of symbolic integration. Then McCarthy started to explain ways a computer could do algebra. Then, as Ed told it to me in 2014:

Feynman produces this sheaf of papers to show us. It was all algebra. And he says “There’s a problem. I’ve done this calculation, and it’s close to 50 pages. A graduate student has done it too, and Murray Gell-Mann has done it. And the only thing we know for sure is that our three results are mutually inconsistent. And the only conclusion we can arrive at is that a person can’t do this much algebra with the hope of getting it right.” And so the question was could there be some system that could help do a problem like that? So what happened is Marvin [Minsky] and I basically fleshed out the idea of a mathematical thing. And it was agreed that we would do it. Marvin and I decided to divide this task up, that I would do one part, and he would do another. Now, we had one bad idea in there, OK. It’s partly Feynman’s fault, but it’s also Marvin and my fault. He was convinced you could not do [math] by typing it. It had to have some kind of handwriting recognition. So, it was decided I would do the handwriting recognition…

And although I didn’t know this until I was writing this piece, it turns out the original proposal for Project MAC was actually based on the idea of building a system for mathematics, and “Project MAC” was originally the “Project on Mathematics and Computation”. Pretty soon, though, the emphasis of Project MAC would shift to the “infrastructure” of timeshared computing. But there was still a math effort, which in time became the MACSYMA system for computer algebra (written in LISP by students and grandstudents of Minsky).

And here this intersects with my personal story. Because many years later (starting in 1976) I would use that system—along with other early computer algebra systems—to do all sorts of physics calculations. My archives still contain an example of what it was like in 1980 to log in to “Project MAC” over the ARPANET (my username was “swolf” in those days; note the system message, the presence of 15 MITishly-named “lusers” altogether, and yes, mail):

Click to enlarge

But, actually, in late 1979 I had already decided to “do my own thing” and build my own system for doing mathematical computation, and eventually much more. And indeed when I first met Ed in 1982 I had recently finished the first version of SMP, and to commercialize it I had started my first company. In 1986 I started to build Mathematica (and what’s now Wolfram Language)—which was released in 1988. Ed started using Mathematica very soon after it was released, and basically continued to do so for the rest of his life.

But picking up the original Project MAC narrative from 1963: the old group from BBN had dispersed but were still writing together about timesharing (and when they said a “debugging system” they meant essentially what we would now call an operating system):

Click to enlarge

And when Project MAC launched in 1963, its “steering committee” included Minsky, Gurley—and Ed. (John McCarthy had landed at Stanford, where he would remain for the rest of his life. I first met him in 1981, at a time when Stanford was trying to recruit me. There was a lunch with the CS department; people went around the room and introduced themselves. McCarthy unhelpfully—and confusingly—said he was “John Smith”.)

Ed at MIT

In 1968, Ed left III—and Minsky, together with Licklider (who had by then become director of Project MAC), persuaded the MIT EE department to hire Ed as a visiting professor for the year. Ed had been spending most of his time at III in Los Angeles, but III also had a pied-à-terre in the Boston area, and indeed its IPO documents listed its address as 545 Technology Square, Cambridge—the very building in which Project MAC was located.

At MIT, Ed invented and taught a freshman course on “Problem Solving”. He told me many times one of his favorite “problem exercises”. Imagine there’s a person who can cure anyone who’s sick just by touching them. How could one set things up to make the best use of this? I must say I never find such implausible hypotheticals terribly interesting. But Ed was proud of a solution that he’d come up with (I think in discussion with Minsky and McCarthy) that involved systematically shuttling millions of people past the healer.

This probably didn’t come from that particular course, but here are some notes I found in an archive of Ed’s papers at MIT that perhaps suggest some of the flavor of the course (we’ll talk about Ed’s interest in the Soviet Union later):

Click to enlarge

In 1968 MIT—and Project MAC in particular—was at the very center of emerging ideas about computer science and AI. A picture from that time captures Ed (third from left) with a few of the people involved: Claude Shannon, John McCarthy and Joe Weizenbaum (creator of ELIZA, the original chatbot):

Click to enlarge

At the end of the 1968 academic year student reviews from Ed’s course were unexpectedly good, and MIT needed faculty members who could be principal investigators on the government grants that were becoming plentiful for computing—and one of those typical-for-Ed “surprising things” happened: MIT agreed to hire him as a full professor with tenure, despite his lack of academic qualifications. It was a watershed moment for Ed, and I think a piece of validation that he carried with pride for the rest of his life. (For what it’s worth, while Ed was an extreme case, MIT was at that time also hiring at least some other people without the usual PhD qualifications into CS professor positions.)

In 1971 Licklider stepped down from his position as director of Project MAC—and Ed assumed the position. His archives from the time contain lots of administrative material—studies, reports, proposals, budgets, etc.—including many pieces reflecting things like the birth of the ARPANET, the maturing of operating systems and the general enthusiasm about the promise of AI.

One item (conceivably from an earlier time) is Ed’s summary of “Information Processing Terminology” for PDP-1 users, complete with definitions like: “A bit is a binary digit or any thing or state that represents a binary digit. Equivalently, a bit is a set with exactly two members. Note that a bit is not one of the members of such a set”:

Click to enlarge

Ed does not seem to have been very central to the intellectual activities around Project MAC, and the emerging Lab for Computer Science and AI Lab. But his name shows up from time to time. And, for example, in the classic “HAKMEM” collection of 191 math and CS “hacks” from the AI Lab, there are two—both very number oriented—attributed to Ed:

Click to enlarge

Rollo Silver gets mentioned too—notably in connection with “random number generators” involving XORs (and, yes, the code is assembly code—for a PDP-10):

Click to enlarge

Also in HAKMEM is the “munching squares” algorithm—that I was later shown by Bill Gosper:

Click to enlarge

And talking of Gosper (whom I’ve known since 1979, and who almost every week seems to send me mail with a surprising new piece of math he’s found with Mathematica): in 1970 the Game of Life cellular automaton had come on the scene, and Gosper and others at MIT were intensely studying it, with Gosper triumphantly discovering the glider gun in November 1970. Curiously—in view of all his emphasis on cellular automata—Ed doesn’t seem to have been involved.

But he did do other things. In 1972, for example, as a kind of spinoff from his Problem Solving course, he formed a group called “The Army to End the War” (i.e. the Vietnam War), whose idea was that it was time to stop the government fighting an unwinnable war, and this could be achieved by having an organization that would coordinate citizens to threaten a run on banks unless the war was ended. Needless to say, though, this didn’t really fit well with the project Ed ran being funded by the Department of Defense.

Between MIT being what it is, and Ed being who he was, there were often strange things that happened. As Ed tells it, one day he was in Marvin Minsky’s office talking about unrecognized geniuses, and a certain Patrick Gunkel walks in, and identifies himself as such. Ed ended up having a long association with Gunkel, who produced such documents as:

Click to enlarge

(Gunkel’s major goal was to create what he called “ideonomy”, or the “science of ideas”, with divisions like isology, chorology, morology and crinology. I met Gunkel once, in Woods Hole, where he had become something of a local fixture, riding around town with his cat in his bicycle basket.)

But after a few years as director of Project MAC, in 1974 Ed was onto something new: being a visiting scholar at Caltech. After his 1961 encounter, he had gotten to know Richard Feynman—who always enjoyed spending time with “out of the box” people like Ed. And so in 1974 Ed went for a year to Caltech, to be with Feynman.

The Universe as a Cellular Automaton

My own efforts (and successes) with cellular automata may perhaps have had something to do with it. But I think at least in the later part of his life, Ed felt his greatest achievements related to cellular automata and in particular his idea that the universe is a giant cellular automaton. I’m not sure when Ed really first hatched this idea, or indeed started to think about cellular automata. Ed had told me many times that when he’d told John McCarthy “the idea”, McCarthy suggested testing it by looking for “roundoff error” in physics, analogous to roundoff error from finite precision in computers. Ed scoffed at this, accusing McCarthy of imagining that there was literally “an IBM 709 computer in the sky”. And Ed’s implication was that he had gotten further than that, imagining the universe to be made more abstractly from a cellular automaton.

I didn’t know quite when this exchange with McCarthy was supposed to have taken place (and, by the way, some of the emerging experimental implications of our Physics Project are precisely about finding evidence of discrete space through something quite analogous to “roundoff errors” in the equations for spacetime). But Ed’s implication to me was always that he’d started exploring cellular automata sometime before 1960.

In the mid-1990s, researching history for my book A New Kind of Science, (as I’ll discuss below) I had a detailed email exchange and long phone conversation with Ed about this. The result was a statement in my notes about the history of cellular automata:

Click to enlarge

At the time, Ed made it sound very convincing. But in writing this piece, I’ve come to the conclusion it’s almost certainly not correct. And of course that’s disappointing given all the effort I put into the history notes in my book, and the almost complete lack of other errors that have surfaced even after two decades of scrutiny. But in any case, it’s interesting to trace the actual development of Ed’s ideas.

One useful piece of evidence is a 25-page document from 1969 in his archives, entitled “Thinking about New Things”—that seems to outline Ed’s thinking at the time. Ed explains “I am not a Physicist, in fact I know very little about modern physics”—but says he wants to suggest a new way of thinking about physics:

Click to enlarge

Soon he starts talking about the possibility that the universe is “merely a simulation on a giant computer”, and relates a version of what he told me about his interaction with John McCarthy:

Click to enlarge

He talks (in a rather programmer kind of way) about the beginning of the universe:

Click to enlarge

He goes on—again in a charmingly “programmer” way:

Click to enlarge

A bit later, Ed is beginning to get to the concept of cellular automata:

Click to enlarge

And there we have it: Ed gets to (3D) cellular automata, though he calls them “spatial automata”:

Click to enlarge

And now he claims that spatial automata can exhibit “very complex behavior”—although his meaning of that will turn out to be a pale shadow of what I discovered in the early 1980s with things like rule 30:

Click to enlarge

But at this point Ed already seems to think he’s almost there—that he’s almost reproduced physics:

Click to enlarge

A little later he’s discussing doing something very much in my style: enumerating possible rules:

Click to enlarge

And still further on he actually talks about 1D rules. And in some sense it might seem like he’s getting very close to what I did in the early 1980s. But his approach is very different. He’s not doing “science” and “empirically seeing what cellular automata do”. Or even being very interested in cellular automata for their own sake. Instead, he’s trying to engineer cellular automata that can “be the universe”. And so for example he wants to consider only left-right symmetric cellular automata “because the universe is isotropic”. And having also decided he wants cellular automata that are symmetric under interchange of black and white (a property he calls “syntactic symmetry”), he ends up with just 8 rules. He could just have simulated these by running them on a computer. But instead he tries to “prove” by pure thought what the rules will do—and comes up with this table:

Click to enlarge

Had he done simulations he might have made pictures like these (labeled using my rule-numbering scheme):

But as it was he didn’t really come to any particular conclusion, other than what amount to a few simple “theorems” about what “data processing” these cellular automata can do:

Click to enlarge

I must say I find it very odd that—particularly given all the stories about his activities and achievements he told me—Ed never in the four decades I knew him mentioned anything about having thought about 1D cellular automata. Perhaps he didn’t remember, or perhaps—even after everything I wrote about them—he never really knew that I was studying 1D cellular automata.

But in any case, what comes next in the 1969 document is Ed getting back to “pure thought” arguments about how cellular automata might “make physics”:

Click to enlarge

It’s a bit muddled (though, to be fair, this was a document Ed never published), but at the end it’s basically saying that if the universe really is just a cellular automaton then one should be able to replace physical experiments (that would, for example, need particle accelerators) with “digital hardware” that just runs the cellular automaton. The next section is entitled “The Design of a Simulator”, and discusses how such hardware could be constructed, concluding that a 1000×1000×1000 3D grid of cells could be built for $50M (or nearly half a billion dollars today).

After that, there’s one final (perhaps unfinished) section that reads a bit like a caricature of “I’ve-got-a-theory-of-physics-too” mechanical models of physics:

Click to enlarge

But, OK, so what does this all mean? Well, first, I think it makes it rather clear that (despite what he told me) by 1969—let alone 1961—Ed hadn’t actually implemented or run cellular automata in any serious way. It’s also notable that in this 1969 piece Ed isn’t using the term “cellular automaton”. The concept of cellular automata had been invented many times, under many different names. But by 1969 the term “cellular automaton” was pretty firmly established, and in fact 1969 might have represented the very peak up to that point of interest in cellular automata in the world at large. But somehow Ed didn’t know about this—or at least wasn’t choosing to connect with it.

Even at MIT Frederick Hennie in the EE department had actually been studying cellular automata—albeit under the name “iterative arrays”—since the very beginning of the 1960s. In 1968 E. F. Codd from IBM (who laid the foundations for SQL—and who worked with Ed’s friend John Cocke) had published a book entitled Cellular Automata. Alvy Ray Smith—in the same department as John McCarthy at Stanford—was writing his PhD thesis on “cellular automata”. In 1969 Marvin Minsky and Seymour Papert published their Perceptrons book, and were apparently talking a lot about cellular automata. And for example by the fall of 1969 Papert’s student Terry Beyer had written a thesis about the “recognition and transformation of figures by iterative arrays of finite state automata”—under the auspices of Project MAC, presumably right under Ed’s nose. (And, no, the thesis doesn’t mention Ed, though it mentions Minsky.)

Right around that time, though, something happens. Ed had been convinced—probably by Minsky and McCarthy—that any cellular automaton capable of “being the universe” better be computation universal. And now there’s a student named Roger Banks who’s working on seeing what kind of (2D) cellular automaton would be needed to get computation universality. Banks had found examples requiring much fewer than the 29 states von Neumann and Burks had used in the 1950s. But—as he related to me many times—Ed challenged Banks to find a 2-state example (“implementable purely with logic gates”), and Banks soon found it, first describing it in June 1970:

Click to enlarge

Banks had apparently been interacting with the “Life hackers” at MIT, and in November 1970 some of the thunder of his result was stolen when Bill Gosper at MIT discovered the glider gun, which suggested that even the rules of the Game of Life (albeit involving 9 rather than 5 2D neighbors) were likely to be sufficient for computation universality.

But for our efforts to trace history, Banks’s June 1970 report has a number of interesting elements. It relates the history of cellular automata, without any mention of Ed. But then—in its one mention of Ed—it says:

Click to enlarge

The “mod-2 rule” that Ed told me he’d simulated in 1961 has finally made an appearance. In an oral history years later Terry Winograd reported that in 1970 he “went to a lecture of Papert’s in which he described a conjecture about cellular automata [which Winograd] came back with a proof of”.

By January 1971, Banks is finishing his thesis, which is now officially supervised by Ed (even though it’s nominally in the mechanical engineering department):

Click to enlarge

Most of Banks’s work is presented as what amount to “engineering drawings”, but he mentions that he has done some simulations. I don’t know if these included simulations of the mod-2 rule but it seems likely.

So was 1969 or 1970 the first time the mod-2 rule had been heard from? I’m not sure, but I suspect so. But to confuse things there’s a “display hack” known as “munching squares” (described in HAKMEM) that looks in some ways similar, and that was probably already seen in 1962 on the PDP-1. Here are the frames in a small example of munching squares:

Here’s a video of a bigger example:

I expect Ed saw munching squares, perhaps even in 1962. But it’s not the mod-2 rule—or actually a cellular automaton at all. And even though Ed certainly had the capability to simulate cellular automata back at the beginning of the 1960s (and could even have recorded videos of 2D ones with III’s film technology) the evidence we have so far is that he didn’t. And in fact my suspicion is that it was probably only around the time I met Ed in 1982 when it finally happened.

My First Encounter with Ed

In May 1981 there’d been a conference at MIT on the Physics of Computation. I’d been invited, but in the end I couldn’t go—because (in a pattern that has repeated many times in my life) it coincided with the initial release of my SMP software system. Still, in December 1981 I got the following invitation:

Click to enlarge

In January 1982 I was planning to go to England to do a few weeks of intensive SMP development on a computer that a friend’s startup had—and I figured I would go to the Caribbean “on the way”.

It was an interesting group that assembled on January 18, 1982, on Mosquito Island. It was the first time I met my now-longtime friend Greg Chaitin. There were physicists there, like Ken Wilson and David Finkelstein. (Despite the promise of the invitation, Feynman’s health prevented him from coming.) And then there were people who’d worked on reversible computation, like Rolf Landauer and Charles Bennett. There were Tom Toffoli and Norm Margolus, who had their cellular automaton machine with them. And finally there was Ed. At first he seemed a little Gatsby-like, watching and listening, but not saying much. I think it was the next morning that Ed pulled me aside rather conspiratorially and said I should come and see something.

Click to enlarge

There was just one real house (as opposed to cabin) on the island (with enough marble to clinch the Bond-villain-lair vibe). Ed led me to a narrow room in the house—where there was a rather-out-of-place-for-a-tropical-island modern workstation computer. I’d seen workstation computers before; in fact, the company I’d started was at the time (foolishly) thinking of building one. But the computer Ed had was from a company he was CEOing. It was a PERQ 1, made by Three Rivers Computer Corporation, which had been founded by a group from CMU including McCarthy’s former student Raj Reddy. I learned that Three Rivers was a company in trouble, and that Ed had recently jumped in to save it. I also learned that in addition to any other challenges the engineers there might have had, he’d added the requirement that the PERQ be able to successfully operate on a tropical island with almost 100% humidity.

But in any case, Ed wanted to show me something on the screen. And here’s basically what it was:

Ed pressed a button and now this is what happened:

I’d seen plenty of “display hacks” before. Bill Gosper had shown me ones at Xerox PARC back in 1979, and my archives even contain some of the early color laser printer outputs he gave me:

Click to enlarge

I don’t remember the details of what Ed said. And what I saw looked like “display hacks flashing on the screen”. But Ed also mentioned the more science-oriented idea of reversibility. And I’m pretty sure he mentioned the term “cellular automaton”. It wasn’t a long conversation. And I remember that at the end I said I’d like to understand better what he was showing me.

And so it was that Ed handed me a PERQ 8” floppy disk. And now, 41 years later, here it is, sitting— still unread—in my archives:

Click to enlarge

It’s not so easy these days to read something like this—and I’m not even sure it will have “magnetically survived”. But fortunately—along with the floppy—there’s something else Ed gave me that day. Two copies of a 9-page printout, presumably of what’s on the floppy:

Click to enlarge

And what’s there is basically a Pascal program (and the PERQ was a very Pascal-oriented machine; “PERQ” is said to have stood for “Pascal Engine that Runs Quicker”). But what does the program do? The main program is called “CA1”, suggesting that, yes, it was supposed to do something with cellular automata.

There are a few comments:

Click to enlarge

And there’s code for making help text:

Click to enlarge

Apparently you press “b” to “clear the Celluar [sic] Automata boundary”, “n” for “Fredkin’s Pattern” and “p” for “EF1”. And at the end there’s a reference to munching squares. The first pattern above is what you get by pressing “n”; the second by pressing “p”.

Both patterns look pretty messy. But if instead you press “a”, you get something with a lot more structure:

I think Ed showed this to me in passing. But he was more interested in the more complicated patterns, and in the fact that you could get them to reverse what they were doing. And in this animated form, I suspect this just looked to me like another munching squares kind of thing.

But, OK, given that we have the program, can we tell what it actually does? The core of it is a bunch of calls to the function rasterop(). Functions like rasterop() were common in computers with bitmapped displays. Their purpose was to apply a certain Boolean operation to the array of black and white pixels in a region of the screen. Here it’s always rasterop(6, …) which means that the function being applied is Boolean function 6, or Xor (or “sum mod 2”).

And what’s happening is that chunks of the screen are getting Xor’ed together: specifically, chunks that are offset by one pixel in each of the four directions. And this is all happening in two phases, swapping between different halves of the framebuffer. Here are the central parts of the sequence of frames that get generated starting from a single cell:

It helps a lot to see the separate frames explicitly. And, yes, it’s a cellular automaton. In fact, it’s exactly the “reversible mod-2 rule”. Here it is for a few more steps, with its simple “self-reproduction” increasingly evident:

Back in 1982 I think I only saw the PERQ that one time. But in one of the resort cabins on the other side of the island—there was this (as captured in a slightly blurry photograph that I took):

Click to enlarge

It was a “cellular automaton machine” built out of “raw electronics” by Tom Toffoli and Norm Margolus—who were the core of Ed’s “Information Mechanics” group at MIT. It didn’t feel much like science, but more like a video DJ performance. Patterns flashing and dancing on the screen. Constant rewiring to produce new effects. I wanted to slow it all down and “sciencify” it. But Tom and Norm always wanted to show yet another strange thing they’d found.

Looking in my archives today, I find just one other photograph I took of the machine. I think I considered this the most striking pattern I saw the machine produce. And, yes, presumably it’s a 2D cellular automaton—though despite my decades of experience with cellular automata I don’t today immediately recognize it:

Click to enlarge

What did I make of Ed back in 1982? Remember, those were days long before the web, and before one could readily look up people’s backgrounds. So pretty much all I knew was that Ed was connected to MIT, and that he owned the island. And I had the impression that he was some kind of technology magnate (and, yes, the island and the plane helped). But it was all quite mysterious. Ed didn’t engage much in technical conversations. He would make statements that were more like pronouncements—that sounded interesting, but were too vague and general for me to do much more than make up my own interpretations for them. Sometimes I would try to ask for clarification, but the response was usually not an explanation, but instead a tangentially related—though often rather engaging—story.

All these years later, though, one particular exchange stands out in my memory. It was at the end of the conference. We were standing around in the little restaurant on the island, waiting for a boat to arrive. And Ed said out of the blue: “I’ll make a deal with you. You teach me how to write a paper and I’ll teach you how to build a company.” At the time, this struck me as quite odd. After all, writing papers seemed easy to me, and I assumed Ed was doing it if he wanted to. And I’d already successfully started a company the previous year, and didn’t think I particularly needed help with it. (Though, yes, I made plenty of mistakes with that company.) But that one comment from Ed somehow for years cemented my view of him as a business tycoon who didn’t quite “get” science, though had ideas about it and wanted to dabble in it.

Ed and Feynman

Ed would later describe Richard Feynman as his best friend. As we discussed above, they’d first met in 1961, and in 1974 Ed had spent the year at Caltech visiting Feynman, having, as Ed tells it, made a deal (analogous to the one he later proposed to me) that he would teach Feynman about computers, and Feynman would teach him about physics. I myself first got to know Feynman in 1978, and interacted extensively with him not only about physics, but also about symbolic computing—and cellular automata. And in retrospect I have to say I’m quite surprised that he mentioned Ed to me only a few times in passing, and never in detail.

But I think the point was that Feynman and Ed were—more than anything else—personal friends. Feynman tended to find “traditional academics” quite dull, and much preferred to hang out with more “unusual” people—like Ed. Quite often the people Feynman hung out with had quite kooky ideas about things, and I think he was always a little embarrassed by this, even though he often seemed to find it fun to indulge and explore those ideas.

Feynman always liked solving problems, and applying himself to different kinds of areas. But I have to say that even I was a little surprised when in writing this piece I was going through the archives of Ed’s papers at MIT, and found the following letter from Feynman to Ed:

Click to enlarge

Clearly he—like me—viewed Ed as an authority on business. But what on earth was this “cutting machine”, and why was Feynman trying to sell it?

For what it’s worth, the next couple of pages tell the story:

Click to enlarge

Feynman’s next-door neighbor had a company that made swimwear, and this was a machine for cutting the necessary fabric—and Feynman had helped develop it. And much as Feynman had been prepared to help his neighbor with this, he was also prepared to help Ed with some of his ideas about physics. And in the archive of Ed’s papers, there’s a letter from Feynman:

Click to enlarge

I don’t know whether this is the first place the term “Fredkin gate” was ever used. But what’s here is a quintessential example of Feynman diving into some new subject, doing detailed calculations (by hand) and getting a useful answer—in this case about what would become Ed’s best-known invention: reversible logic, and the Fredkin gate.

Feynman had always been interested in “computing”. And indeed when he was recruited to the Manhattan Project it was to run a team of human computers (equipped with mechanical desk calculators). I think Feynman always hoped that physics would “become computational” at least in some sense—and he would for example lament to me that Feynman diagrams were such a bad way to compute things. Feynman always liked the methodology of traditional continuous mathematics, but (as I just noticed) even in 1964 he was saying that “I believe that the theory that space is continuous is wrong, because we get these infinities and other difficulties…”. And elsewhere in his 1964 lectures that became The Character of Physical Law Feynman says:

Click to enlarge

Did Feynman say these things because of his conversations with Ed? I rather doubt it. But as I was writing this piece I learned that Ed thought differently. As he told it:

I never pressed any issue that would sort of give me credit, okay? It’s just my nature. A very weird thing happened toward the end of my time at Caltech. Richard Feynman and I would get into very fierce arguments. . . . I’m trying to convince him of my ideas, that at the bottom is something finite and so on. He suddenly says to me, “You know, I’m sure I had this same idea sometime quite a while ago, but I don’t remember where or how or whether I ever wrote it down.” I said, “I know what you’re talking about. It’s a set of lectures you gave someplace. In those lectures you said perhaps the world is finite.” He just has this little statement in this book. I saw the book on his shelf. I got it out, and he was so happy to see that there. What I didn’t tell him was he gave that lecture years after I’d been haranguing him on this subject. I knew he thought it was his idea, and I left it that way. That was just my nature.

Notwithstanding what he said, I rather suspect he did push the point. And for example when Feynman gave a talk on “Simulating Physics with Computers” at the 1981 MIT Physics of Computation conference that Ed co-organized, he was careful to write that:

Click to enlarge

Ed, by the way, arranged for Feynman to get his first personal computer: a Commodore PET. I don’t think Feynman ended up using it terribly much, though in 1984 he took it with him on a trip to Hawaii where he and his son Carl used it to work out probabilities to try to “crack” the randomness of my rule 30 cellular automaton (needless to say, without success).

Digital Physics & Reversible Logic

Back at MIT in 1975 after his year at Caltech, Ed was no longer the director of Project MAC, but was still on the books as a professor, albeit something of an outcast one. Soon, though, he was teaching a class about his ideas—under the title of “Digital Physics”:

Click to enlarge

Cellular automata weren’t specifically mentioned in the course description—though in the syllabus they were there, with the Game of Life as a key example:

Click to enlarge

Back in the 1960s, cellular automata had been a popular topic in theoretical computer science. But by the mid-1970s the emphasis of the field had switched to things like computational complexity theory—and, as Ed told me many times, his efforts to interest people at MIT in cellular automata failed, with influential CS professor Albert Meyer (whose advisor Patrick Fischer had worked quite extensively on cellular automata) apparently telling Ed that “one can tell someone is out of it if they don’t think cellular automata are dead”. (It’s an amusing irony that around this time, Meyer’s future wife Irene Greif would point John Moussouris—who we’ll meet later—to Ed and his work on cellular automata.)

Ed’s ideas about physics were not well received by the physicists at MIT. And for example when students from Ed’s class asked the well-known MIT physics professor Philip Morrison what he thought of Ed’s approach, he apparently responded that “Of course Fredkin thinks the universe is a computer—he’s a computer person; if instead he were a cheese merchant he’d think it was a big cheese!”

When Ed was at Caltech in 1974 a big focus there—led by Carver Mead—was VLSI design. And this led to increasing interest in the ultimate limits on computation imposed by physics. Ever since von Neumann in the 1950s it had been assumed that every step in a computation would necessarily require dissipation of energy—and this was something Carver Mead took as a given. But if this was true, how could Ed’s cellular automaton for the universe work? Somehow, Ed reasoned, it—and any computation, for that matter—had to be able to run reversibly, without dissipating any energy. And this is what led Ed to his most notable scientific contribution: the idea of reversible logic.

Ordinary logic operations—like And and Or—take two bits of input and give one bit of output. And this means they can’t be reversible: with only one bit in the output there isn’t information to uniquely determine the two bits of input from the output. But if—like Ed—you consider a generalized logic operation that for example has both two inputs and two outputs, then this can be invertible, i.e. reversible.

The concept of an invertible mapping had long existed in mathematics, and under the name “automorphisms of the shift” had even been studied back in the 1950s for the case of what amounted to 1D cellular automata (for applications in cryptography). And in 1973 Charles Bennett had shown that one could make a reversible analog of a Turing machine. But what Ed realized is that it’s possible to make something like a typical computer design—and have it be reversible, by building it out of reversible logic elements.

Looking through the archive of Ed’s papers at MIT, I found what seem to be notes on the beginning of this idea:

Click to enlarge

And I also found this—which I immediately recognized as a sorting network, in which values get sorted through a sequence of binary comparisons:

Click to enlarge

Sorting networks are inevitably reversible. And this particular sorting network I recognized as the largest guaranteed-optimal sorting network that’s known—discovered by Milton Green at SRI (then “Stanford Research Institute”) in 1969. It’s implausible that Ed independently discovered this exact same network, but it’s interesting that he was drawing it (by hand) on a piece of paper.

Ed’s archives also contain a 3-page draft entitled “Conservative Logic”:

Click to enlarge

Ed explains that he is limiting himself to gates that implement permutations

Click to enlarge

and then goes on to construct a “symmetric-majority-parity” gate—which he claims is “computation universal”:

Click to enlarge

It’s not quite a Fredkin gate, but it’s close. And, by the way, it’s worth pointing out that these gates alone aren’t “computation universal” in something like the Turing sense. Rather, the point is that—like with Nand for ordinary logic—any reversible logic operation (i.e. permutation) with any number of inputs can be constructed using just these gates, connected by wires.

Ed didn’t at first publish anything about his reversible logic idea, though he talked about it in his class, and in 1978 there were already students writing term papers about it. But then in 1978, as Ed told it later:

I found this guy Tommaso Toffoli. He had written a paper that showed how you could build a reversible computer by storing everything that an ordinary computer would have to forget. I had figured out how to have a reversible computer that didn’t store anything because all the fundamental activity was reversible. Okay? So I decided to hire him because he was the only person who tried to do it and he didn’t succeed, really, and I had—and I hired him to help me.

Toffoli had done a first PhD in Italy building electronics for cosmic ray detectors, and in 1978 he’d just finished a second PhD, working on 2D cellular automata with Art Burks (who had coined the name “cellular automaton”). Ed brought Toffoli to MIT under a grant to build a cellular automaton machine—leading to the machine I saw on Ed’s island in 1982. But Ed also worked with Toffoli to write a paper about conservative logic—which finally appeared in 1982, and contained both the Fredkin gate, and the Toffoli gate. (Ed later griped to me that Toffoli “really hadn’t done much” for the paper—and that after all the Toffoli gate was just a special case of the Fredkin gate.)

Back in 1980—on the way to this paper—Ed, with Feynman’s encouragement, had had another idea: to imagine implementing reversible logic not just abstractly, but through an explicit physical process, namely collisions between elastic billiard balls. And as we saw above, Feynman quickly got into analyzing this, for example seeing how a Fredkin gate could be implemented just with billiard balls.

But ultimately Ed wanted to implement reversibility not just for things like circuits, but also—imitating the reversibility that he believed was fundamental to physics—for cellular automata. Now the fact is that reversibility for cellular automata had actually been quite well studied since the 1950s. But I don’t think Ed knew that—and so he invented his own way to “get reversibility” in cellular automata.

It came from something Ed had seen on the PDP-1 back in 1961. As Ed tells it, in playing around with the PDP-1 he had come up with a piece of code that surprised him by drawing something close to a circle in pixels on the screen. Minsky had apparently “gone into the debugger” to see how it worked—and in 1972 HAKMEM attributed the algorithm to Minsky (though in the Pascal program I got from Ed in 1982, it appears as a function called efpattern()). Here’s a version of the algorithm:

And, yes, with different divisors d it can give rather different (and sometimes wild) results:

But for our purposes here what’s important is that Ed found out that this algorithm is reversible—and he realized that in some sense the reason is that it’s based on a second-order recurrence. And, once again, the basic ideas here are well known in math (cf. reversibility of the wave equation, which is second order). But Ed had a more computational version: a second-order cellular automaton in which one adds mod 2 the value of a cell two steps back. And I think in 1982 Ed was already talking about this “mod-2 trick”—and perhaps the PERQ program was intended to implement it (though it didn’t).

Ed’s work on reversible logic and “digital physics” in a sense came to a climax with the 1981 Physics of Computation conference at MIT—that brought in quite a Who’s Who of people who’d been interested in related topics (as I mentioned above, I wasn’t there because of a clash with the release of SMP Version 1.0, though I did meet or at least correspond with most of the attendees at one time or another):

Click to enlarge

Originally Ed wanted to call the conference “Physics and Computation”. But Feynman objected, and the conference was renamed. In the end, though, Feynman gave a talk entitled “Simulating Physics with Computers”—which most notably talked about the relation between quantum mechanics and computation, and is often seen as a key impetus for the development of quantum computing. (As a small footnote to history, I worked with Feynman quite a bit on the possibility of both quantum computing and quantum randomness generation, and I think we were both convinced that the process of measurement was ultimately going to get in the way—something that with our Physics Project we are finally now beginning to be able to analyze in much more detail.)

But despite his interactions with Feynman, Ed was never too much into the usual ideas of quantum mechanics, hoping (as he said in the flyer for his course on digital physics) that perhaps quantum mechanics would somehow fall out of a classical cellular-automaton-based universe. But when quantum computing finally became popular in the 1990s, reversible logic was a necessary feature, and the Fredkin gate (also known as CSWAP or “controlled-swap”) became famous. (The Toffoli gate—or CCNOT—is a bit more famous, though.)

In tracing the development of Ed’s ideas, particularly about “digital physics”, there’s another event worthy of mention. In late 1969 Ed learned about an older German tech entrepreneur named Konrad Zuse who’d published an article in 1967 (and a book in 1969) on Rechnender Raum (Calculating Space)—mentioning the term “cellular automata”:

Click to enlarge

Although Zuse was 24 years older than Ed, there were definitely similarities between them. Zuse had been very early to computers, apparently building one during World War II that suffered an air raid (and may yet still lie buried in Berlin). After the war, Zuse started a series of computer companies—and had ideas about many things. He’d been trained as an engineer, and perhaps it was having worked on solving his share of PDEs using finite differences that led him to the idea—a bit like Ed’s—that space might fundamentally be a discrete grid. But unlike Ed, Zuse for the most part seemed to think that—as with finite differences—the values on the grid should be continuous, or at least integers. Ed arranged for Zuse’s book to be translated into English, and for Zuse to visit MIT. I don’t know how much influence Zuse had on Ed, and when Ed talked to me about Zuse it was mostly just to say that people had treated his ideas—like Ed’s—as rather kooky. (I exchanged letters with Zuse in the 1980s and 1990s; he seemed to find my work on cellular automata interesting.)

Ideas & Inventions Galore

It wasn’t just physics that Ed had ideas about. It was lots of other things too. Sometimes the ideas would turn into businesses; more often they’d just stay as ideas. Ed’s archive, for example, contains a document on the “Intermon Idea” that Ed hoped would “provide a permanent solution to the world’s problem of not having a stable medium of exchange”:

Click to enlarge

And, no, Ed wasn’t Satoshi Nakamoto—though he did tell me several times that (although, to his displeasure, it was never acknowledged) he had suggested to Ron Rivest (the “R” of RSA cryptography) the idea of “using factoring as a trapdoor”. And—not content with solving the financial problems of the world, or, for that matter, fundamental physics—Ed also had his “algorithmic plan” to prevent the possibility of World War III.

And then there was the Muse. Marvin Minsky had long been involved with music, and had assembled out of electronic modules a system that generated sequences of musical notes. But in 1970 Ed and Minsky developed what they called the Muse—whose idea was to be a streamlined system that would use integrated circuits to “automatically compose music”:

Click to enlarge

In actuality, the Muse produced sequences of notes determined by a linear feedback shift register—in essence a 1D additive cellular automaton—in which the details of the rule were set on its front panel as “themes”. The results were interesting—if rather R2-D2-like—but weren’t what people usually thought of as “music”. Ed and Minsky started a company named Triadex (note the triangular shape of the Muse), and manufactured a few hundred Muses. But the venture was not a commercial success.

Particularly through interacting with Minsky, Ed was quite involved in “things that should be possible with AI”. The Muse had been about music. But Ed also for example thought about chess—where he wanted to build an array of circuits that could tree out possible moves. Working with Richard Greenblatt (who had developed an earlier chess machine) my longtime friend John Moussouris ended up designing CHEOPS (a “Chess-Oriented Processing System”) while Ed was away at Caltech. (Soon thereafter, curiously enough, Moussouris would go to Oxford and work with Roger Penrose on discrete spacetime—in the form of spin networks. Then in later years he would found two important Silicon Valley microprocessor companies.)

Keeping on the chess theme, Ed would in 1980 (through his Fredkin Foundation) put up the Fredkin Prize for the first computer to beat a world champion at chess. The first “pre-prize” of $5k was awarded in 1981; the second pre-prize of $10k in 1988—and the grand prize of $100k was awarded in 1997 with some fanfare to the IBM Deep Blue team.

Ed also put up a prize for “math AI”, or, more specifically, automated theorem proving. It was administered through the American Math Society and a few “milestone prizes” were given out. But the grand Leibniz Prize “for the proof of a ‘substantial’ theorem in which the computer played a major role” was never claimed, the assets of the Fredkin Foundation withered, and the prize was withdrawn. (I wonder if some of the things done in the 1980s and 1990s by users of Mathematica should have qualified—but Ed and I never made this connection, and it’s too late now.)

Ed the Consultant

Particularly during his time at MIT, Ed did a fair amount of strategy consulting for tech companies—and Ed would tell me many stories about this, particularly related to IBM and DEC (which were in the 1980s the world’s two largest computer companies).

One story (whose accuracy I’ve never been able to determine) related to DEC’s ultimately disastrous decision not to enter the personal computer business. As Ed tells it, a team at DEC did a focus group about PCs—with Ken Olsen (CEO of DEC) watching. There was a young teacher in the group who was particularly enthusiastic. And Olsen seemed to be getting convinced that, yes, PCs were a good idea. As the focus group was concluding, the teacher listed off all sorts of ways PCs could change the world. But then, fatefully, he added right at the end: “And I don’t just mean here on Earth”. Ed claims this was the moment when Olsen decided to kill the PC project at DEC.

Ed told a story from the early 1970s about a giant IBM project called FS (for “Future Systems”):

IBM has this project. They’re going to completely revolutionize everything. The project is to design everything from the smallest computer to the new largest. They’re all to be multiprocessors. The specs were just fantastic. They promised to guarantee their customers 100% uptime. Their plans were, for instance, when you have a new OS, it’s updated. They guarantee 24-hour operation at all times. They plan to be able to update the OS without stopping this process. Things like that, a lot of goals that are very lofty, and so on.

Someone at IBM whom I knew very well, a very senior guy, came to me one day and said, “Look, these guys are in trouble, and maybe MIT could help them.” I organized something. Just under 30 professors of computer science came down to IBM. We got there on Sunday night and starting Monday morning, we got one lecture an hour, eight on Monday, Tuesday, Wednesday, Thursday, and four on Friday, describing the system. It was just spectacular, everything they were trying to do, but it was full of all kinds of idiocy. They were designing things that they’d never used. This whole thing was to be oriented about people looking at displays.

No one at IBM had done anything like that. They think, “Okay, you should have a computer display,” and they came up with certain problems that hadn’t occurred to the rest of us. If you’re looking at the display, how can you tell the difference between what you had put into the computer and what the computer had put in? This worried them. They came up with a hardware fix. When you typed, it always went on the right half of the screen; when the computer did something, it always went on the left half, or I may have it backwards, but that was the hardware.

What happened is I came to realize that they were so over their head in their goal that they were going to annihilate themselves with this thing. It was just going to be the world’s greatest fiasco for it. I started cornering people and saying, “Look, do you realize that you’re never going to make this work?” and so on, so forth. This came to the attention of people at IBM, and it annoyed them. I got a call from someone saying, “Look, you’re driving us nuts. We want to hear you out, so we’re going to conduct a debate.” There’s a guy named Bob [Evans], who was the head of the project. What happened was we’re in the boardroom with IBM, lots of officials there, and he and I have a debate.

I’m debating that they have to kill the project and do something else. He’s debating that they shouldn’t kill the project. I made all my points. He made all his points. Then a guy named Mannie Piore, who was the one who thought of the idea of having a research laboratory, a very senior guy said to me, he said, “Hey, Ed,” he said, “We’ve heard you out.” He says, “This is our company. We can do this product even if you think we shouldn’t.” I said, “Yes, I admit that’s true.” He said, “You presented your case. We’ve heard you out, and we want to do it.” I said, “Okay.” He said, “Can you do us a favor?” I said, “What’s that?’ He said, “Can you stop going around talking to people about why it has to be killed?” I said, “Look, I’ve said my piece. I’ve been heard out.” “Yes. Okay.” “I quit.”

I had only one ally in that room; that was John Cocke. As we were walking out of the room, he came over to me and said, “Don’t worry, Ed.” He said, “It’s going to fall over of its own weight.” I’ll never forget that. Ten days later, it was canceled. A lot of people were very mad at me.

I’m not sure what Ed was like as an operational manager of businesses. But he certainly had no shortage of opinions about how businesses should be run, or at least what their strategies should be. He was always keen on “do-the-big-thing” ideas. I remember him telling me multiple times about a company that did airplane navigation. It had put a certain number of radio navigation beacons into its software. Ed told me he’d asked about others, and the company had said “Well, we only put in the beacons lots of people care about”. Ed said “Just put all of them in”. They didn’t. And eventually they were overtaken by a company that did.

Ed the Businessman

Ed’s great business success—and windfall—was III. But Ed was also involved with a couple dozen other companies—almost all of which failed. There’s a certain charm in the diversity of Ed’s companies. There was Three Rivers Computer Corporation, that made the PERQ computer. There was Triadex, that made the Muse. There was a Boston television station. There was an air taxi service. There was Fredkin Enterprises, importing PCs into the Soviet Union. There was Drake’s Anchorage, the resort on his island. There was Gensym, a maker of AI-oriented process control systems, which was a rare success. And then there was Reliable Water.

Ed’s island—like many tropical islands—had trouble getting fresh water. So Ed decided to invent a solution, coming up with a new, more energy-optimized way to do reverse osmosis—with a dash of AI control. Reliable Water announced its product in May 1987, desalinating water taken from Boston Harbor and serving it to journalists to drink. (Ed told me he was a little surprised how willingly they did so.)

Click to enlarge

Looking at my archives I see I was sufficiently charmed by the picture of Ed posing with his elaborate “intelligent” glass tubing that I kept the article from New Scientist:

Click to enlarge

As Ed told it to me, Reliable Water was just about to sell a major system to an Arab country when his well-pedigreed CEO somehow cheated him, and the deal fell through.

But what about the television station? How did Ed get involved with that? Apparently in 1969 Jerry Wiesner, then president of MIT, encouraged Ed to support a group of Black investors (led by a certain Bertram Lee) who were challenging the broadcasting license of Boston’s channel 7. Years went by, other suitors showed up, and litigation about the license went all the way to the Supreme Court (which described the previous licensee as having shown an “egregious lack of candor” with the FCC). For a while it seemed like channel 7 might just “go dark”. But in early January 1982 (just a couple of weeks before I first met him) Ed took over as president of New England Television Corporation (NETV)—and in May 1982 NETV took over channel 7, leaving Ed with a foot of acquisition documents in his home library, and a television channel to run:

Click to enlarge

There’d been hopes of injecting new ideas, and adding innovative educational and other content. But things didn’t go well and it wasn’t long before Ed stepped down from his role.

A major influence on Ed’s business activities came out of something that happened in his personal life. In 1977 Ed had been married for 20 years and had three almost-grown children. But then he met Joyce. On a flight back from the Caribbean he sat next to a certain Joyce Wheatley who came from a prominent family in the British Virgin Islands and had just graduated with a BS in economics and finance from Bentley College (now Bentley University) in Waltham, MA. As both Ed and Joyce tell it, Ed immediately gave advice like that the best way to overcome a fear of flying was to learn to fly (which much later, Joyce in fact did).

Joyce was starting work at a bank in Boston, but matters with Ed intervened, and in 1980 the two of them were married in the Virgin Islands, with Feynman serving as Ed’s best man (and at the last minute lending Ed a tie for the occasion). In 1981, Ed and Joyce had a son, who they named Richard after Richard Feynman (though now themed as “Rick”)—of whom Ed was very proud.

When Ed died, Joyce and he had been married for 43 years—and Joyce had been Ed’s key business partner all that time. They made many investments together. Sometimes it’d start with a friend or vendor. Sometimes Ed (or Joyce) would meet students or others—who’d be invited over to the house some evening, and leave with a check. Sometimes the investments would be fairly hands-off. Sometimes Ed would get deeply involved, even at times playing CEO (as he did with Three Rivers and NETV).

When the web started to take off, Ed and Joyce created a company called Capital Technologies which did angel investing—and ended up investing in many companies with names like Sourcecraft, SqueePlay, EchoMail, Individual Inc. and Radnet. And—like so many startups of this kind—most failed.

Ed also continued to have all sorts of ideas of his own, some of which turned into patents. And—like so much to do with Ed—they were eclectic. In 1995 (with a couple of other people) there was one based on using evanescent waves (essentially photon tunneling) to more accurately find the distance between the read/write head and the disk in a disk drive or CD-ROM drive. Then in 1999 there was the “Automatic Refueling Station”—using machine vision plus a car database to automate pumping gas into cars:

Click to enlarge

That was followed in 2003 by a patent about securely controlling telephone switching from web clients. In 2006, there was a patent application named simply “Contract System” about an “algorithmic contract system” in which the requirements of buyers and sellers of basically anything would be matched up in a kind of tiling-oriented geometrical way:

Click to enlarge

In 2011 there was “Traffic Negotiation System”, in which cars would have rather-airplane-like displays installed that would get them in effect to “drive in formation” to avoid traffic jams:

Click to enlarge

Ed’s last patent was filed in 2015, and was essentially for a scheme to cache large chunks of the web locally on a user’s computer—a kind of local CDN.

But all these patents represented only a small part of Ed’s “idea output”. And for example Ed told me many other tech ideas he had—a few of which I’ll mention later.

And Ed’s business activities weren’t limited to tech. He did his share of real-estate transactions too. And then there was his island. For years Joyce and Ed continued to operate Drake’s Anchorage, and tried to improve the infrastructure of the island—with Ed, as Joyce tells it, more often to be found helping to fix the generator on the island than partaking of its beaches.

Back in 1978 Ed had acquired a “neighbor” when Richard Branson bought Necker Island, which was a couple of miles further out towards the Atlantic than Moskito Island. Ed told me quite a few stories about Branson, and for years had told me that Branson wanted to buy his island. Ed hadn’t been interested in selling, but eventually agreed to give Branson right of first refusal. Then in 2007 a Czech (or were they a Russian?) showed up and offered to buy the island for cash “to be delivered in a suitcase”. It was all rather sketchy, but Ed and Joyce decided it was finally time to sell, and let Branson exercise his right of first refusal, and buy the island for about $10M.

Ed and His Toys

Ed liked to buy things. Computers. Cars. Planes. Boats. Oh, and extra houses too (Vermont, Martha’s Vineyard, Portola Valley, …)—as well as his island. Ed would typically make decisions quickly. A house he drove by. New tech when it first came out. He was always proud of being an early adopter, and he’d often talk almost conspiratorially about the “secret” features he’d figured out in new tech he’d bought.

But I think Ed’s all-time favorite “toys” were planes—and over the course of his life he owned a long sequence of them. Ed was a serious (and, by all reports, exceptionally good) pilot—with an airplane transport pilot license (plus seaplane and glider licenses). And I always suspected that his cut-and-dried approach to many things reflected his experience in making decisions as a pilot.

Ed at different times had a variety of kinds of planes, usually registered with the vanity tail number N1EF. There were twin-propellor planes. There were high-performance single-propellor planes. There was the seaplane that I’d “met” in the Caribbean. At one time there was a jet—and in typical fashion Ed got himself certified to fly the jet singlehandedly, without a copilot. Ed had all sorts of stories about flying. About running into Tom Watson (CEO of IBM) who was also a pilot. About getting a new type of plane where he thought he was getting #5 off the production line, but it was actually #1—and one day its engine basically melted down, but Ed was still able to land it.

Ed also had gliders, and competed in gliding competitions. Several times he told me a story—as a kind of allegory—about another pilot in a gliding competition. Gliders are usually transported with their wings removed, with the wings attached in order to fly. Apparently there was an extra locking pin used, which the other pilot decided to remove to save weight, because it didn’t seem necessary. But when the glider was flying in the competition its wings fell off. (The pilot had a parachute, but landed embarrassed.) The very pilot-oriented moral as far as Ed was concerned: just because you don’t understand why something is there, don’t assume it’s not necessary.

Ed and the Soviet Union

One of the topics about which Ed often told “you-can’t-make-this-stuff-up” stories was the Soviet Union. Ed’s friend John McCarthy had parents who were active communists, had learned Russian, and regularly took trips to the Soviet Union. And as Ed tells it McCarthy came to Ed one day and said (perhaps as a result of having gotten involved with a Russian woman) “I’m moving to the Soviet Union”, and talked about how he was planning to dramatically renounce his US citizenship. McCarthy began to make arrangements. Ed tried to talk him out of it. And then it was 1968 and the Soviets send their tanks into Czechoslovakia—and McCarthy is incensed, and according to Ed, sends a telegram to a very senior person in the Soviet Union saying “If you invade Czechoslovakia then I’m not coming”. Needless to say, the Soviets ignored him. Ed told me he’d said at the time: “If the Russians were really smart and really understood things, and they had to choose between John McCarthy and Czechoslovakia, they should have chosen John McCarthy.” (McCarthy would later “flip” and become a staunch conservative.)

Perhaps through McCarthy, Ed started visiting the Soviet Union. He didn’t like the tourist arrangements (required to be through the government’s Intourist organization)—and decided to try to do something about it, sending a survey to Americans who’d visited the Soviet Union:

Click to enlarge

A year later, Ed was back in the Soviet Union, attending a somewhat all-star conference (along with McCarthy) on AI—with a rather modern-sounding collection of topics:

Click to enlarge

Here’s a photograph of a bearded Ed in action there—with a very Soviet simultaneous translation booth behind him:

Click to enlarge

Ed used to tell a story about Soviet computers that probably came from that visit. The Soviet Union had made a copy of an IBM mainframe computer—labeling it as a “RYAD” computer. There was a big demo—and the computer didn’t work. The generals in charge asked “Well, did you copy everything?” As it turned out, there was active circuitry in the “IBM” logo—and that needed to be copied too. Or at least that’s what Ed told me.

But Ed’s most significant interaction with the Soviet Union came in the early 1980s. The US had in place its CoCom list that embargoed export of things like personal computers to the Soviet Union. Meanwhile, within the Soviet Union, photocopiers were strictly controlled—to prevent non-state-sanctioned flow of information. But as Ed tells it, he hatched a plan and sold it to the Reagan administration, telling them: “You’re on the wrong track. If we can get personal computers into the Soviet Union, it breaks their lock on the flow of information.” But the problem was he had to convince the Soviets they wanted personal computers.

In 1984 Ed was in Moscow—supposedly tagging along to a physics conference with an MIT physicist named Roman Jackiw. He “dropped in” at the Computation Center of the Academy of Sciences (which, secretly, was a supplier to the KGB of things like speech recognition tech). And there he was told to talk to a certain Evgeny Velikhov, a nuclear physicist who’d just been elected vice president of the Academy of Sciences. Velikhov arranged for Ed to give a talk at the Kremlin to pitch the importance of computers, which apparently he successfully did, after convincing the audience that his motivation was to make the world a safer place by balancing the technical capabilities of East and West.

And as if to back up this point, while he was in the Soviet Union, Ed wrote a 5-page piece from “A Concerned Citizen, Planet Earth” addressed “To whom it may concern” in Moscow and Washington—ending with the suggestion that its plan might be discussed at an upcoming meeting between Andrei Gromyko and Ronald Reagan at the UN:

Click to enlarge

The piece mentions another issue: the fate of prominent, but by then dissident, Soviet physicist Andrei Sakharov, who was in internal exile and reportedly on hunger strike. Ed hatched a kind of PCs-for-Sakharov plan in which the Soviets would get PCs if they freed Sakharov.

Meanwhile, in true arms-dealer-like fashion, he’d established Fredkin Enterprises, S.A. which planned to export PCs to the Soviet Union. He had his student Norm Margolus spend a summer analyzing the CoCom regulations to see what characteristics PCs needed to have to avoid embargo.

In the Reagan Presidential Library there’s now a fairly extensive file entitled “Fredkin Computer Exports to USSR”—which for example contains a memo reporting a call made on August 25, 1984, by then-vice-president George H. W. Bush to Sakharov’s stepdaughter, who was by that time living in Massachusetts (and, yes, Ed was described as a “PhD in computer science” with a “flourishing computer business”):

Click to enlarge

Soon the White House is communicating with the US embassy in Moscow to get a message to Ed:

Click to enlarge

And things are quickly starting to sound as if they were from a Cold War spy drama (there’s no evidence Ed was ever officially involved with the US intelligence services, though):

Click to enlarge

I don’t think Ed ever ended up talking to Sakharov, but on November 6, 1984, Fredkin Enterprises was sent a letter by Velikhov ordering 100 PCs for the Academy of Sciences, and saying they hoped to order 10,000 more. But the US was not as speedy, and in 1985 there was still back and forth about CoCom issues. Ed of course had a plan:

And indeed in the end Ed did succeed in shipping at least some computers to the Soviet Union, adding a hack to support Cyrillic characters. Ed often took his family with him to Moscow, and he told me that his son Rick created quite a stir when at age 6 he was seen there playing a game on a computer. Up to then, computers had always been viewed as expensive tools for adults. But after Rick’s example there were suddenly all sorts of academicians’ kids using computers.

(In the small world that it is, one person Ed got to know in the Academy of Sciences was a certain Arkady Borkovsky—who in 1989 would leave Russia to come work at our company, and who would later co-found Yandex.)

By the way, to fill in a little color of the time, I might relate a story of my own. In 1987 I went to a (rather Soviet) conference in Moscow on “Logic, Methodology and Philosophy of Science.” Like everyone, I was assigned a “guide”. Mine continually tried to pump me for information about the American computer industry. Eventually I just said: “So what do you actually want to know?” He said: “We’ve cloned the Intel 8086 microprocessor, and we want to know if it’s worth cloning the Motorola 68000. Motorola has put a layer of epoxy that makes it hard to reverse engineer.” He assumed that the epoxy was at the request of the US government, to defeat Soviet efforts—and he didn’t believe me when I said I thought it was much more likely there to defeat Intel.

Ed told me another story about his interactions with Soviet computer efforts after Gorbachev came to power:

Before the days of integrated circuits the way IBM and Digital built computers was they put the whole computer together, and then it would sit for six weeks in “system integration” while they made the pieces work together and slowly got the bugs out.

The Russians built computers differently because that seemed logical to them. They’d send all the components down there and then some guy was supposed to plug them together, and they were supposed to work. But they didn’t. With these big computers, they never made any of them work.

The Academy of Sciences had one. And one time I went to see their big computer, so they unlock the doors to this dusty room where the computer is, where it’s not being used because it doesn’t work, and all this information is being kept secret, not from the United States, but from the leadership. When I discovered all this I documented it … and I wrote a 40-page document that explained it.

I was making trips with Rick often and Mike [his older son] very often. On one trip when I arrived, they tell me, “Oh, you have to come to this meeting.”

I don’t speak Russian. I never knew it. I’m seated at this meeting, and there’s a Russian friend of mine [head of the Soviet Space Research Institute] next to me. We’re just sitting there, and things are going on. I still don’t know what that meeting was, but I had this 40-page document. I gave it to my friend. He starts reading. He says, “Oh, this is so interesting.” It got to be about ten o’clock at night and they said, “Everyone come back in the morning. Nine o’clock.”

My friend said, “Can I borrow this [document]? I’ll bring it back in the morning”. I said, “Sure, go ahead.” He comes back next morning. He says to me, “I have good news, and I have bad news.” I said, “What’s the good news?” He says, “Your document has been translated into Russian.” I said, “You left here with a 40-page typewritten document. I don’t believe you.” He said, “Well, my institute recently took on the task of translating scientific American into Russian.

“When I left here, I went to my institute, called in the translators, and they all came in. We divided the document up between them, and it’s all been translated into Russian.”

The document was the analysis of the RYAD situation with the recommendation that the only thing they could do was to cancel it all.

I said, “Okay, what’s the bad news?” He says, “The bad news is it’s classified secret.” When you made a copy or did something, you had to have a government person look at it. They classified it. I said to him, “You can’t classify my documents.” He said, “Of course not. We haven’t. It’s just the Russian one that’s secret.”

Then maybe a week later, he said, “Gorbachev’s read your document.” He canceled it. RYAD. Some people I know were looking to kill me.

In Moscow, there’s a building that’s so unusual. It’s on a highway leading into the city. It’s about five stories high. It’s about a kilometer long, okay? It’s a giant building. I was in it a few years ago, and it’s just a beehive of startups, almost all software startups. That was the RYAD Software Center, okay? 100,000 people got put out of work.

Ed Becomes a Physics Professor

When I first met Ed in 1982, he was in principle a professor at MIT. But he was also CEOing a computer company (Three Rivers), and, though I didn’t know it at the time, had just become president of a television channel. Not to mention a host of other assorted business activities. MIT had a policy that professors could do other things “one day a week”. But Ed was doing other things a lot more than that. Ed used to say he was “tricked” out of his tenured professorship. Because in 1986 he was convinced that with all the other things he had going on, he should become an adjunct professor. But apparently he didn’t realize that tenure doesn’t apply to adjunct professors. And, as Ed told it, the people in the department considered him something of a kook, and without tenure forcing them to keep him, were keen to eject him.

Minsky’s neighbor in Brookline, MA, was a certain Larry Sulak—the very energetic chairman of the physics department at Boston University (and someone I have known since the 1970s). Ed knew Sulak and when Ed was ejected from MIT, Sulak seized the opportunity to bring Ed in as a physics professor at Boston University. Sulak asked me to write a letter about Ed (and, yes, particularly after the research for this piece, there are some things I would change today):

Subject: Re: Ed Fredkin
Date: Aug 24, 1988
From: Stephen Wolfram
To: Larry Sulak

Dear Larry:

In this century, people like Ed Fredkin have been very rare. Ed Fredkin
is a gentleman scientist. He has made several fortunes in business, yet
he chooses to spend much of his time thinking about science.

The main thing he thinks about is what ideas from computing can tell us
about physics. This is an area that I believe has fundamental importance
for physics. There are many issues about the behaviour of complex
physical systems where the best hope for analysis and understanding comes
from computational ideas. There are also many traditional problems
in quantum physics and other fundamental areas that I suspect are most
likely to be solved by thinking about things from a computational point
of view.

Ed Fredkin has had some very good ideas about physics and its relation
to computation. Probably the single most important was his independent discovery
of the possibility of thermodynamically reversible computation.
von Neumann got this wrong — by thinking about things from a computational
point of view, Fredkin got it right.

Fredkin has been convinced for many years that cellular automata —
basically computational models — could describe fundamental physical
processes. As you know, I have worked on using cellular automata to
model various specific physical processes. Fredkin is trying to do something
grander — he wants to show that all of physics can be reproduced by
a cellular automaton. If he is right the discovery would be one whose
importance could be compared to the discovery of quantization.
Of course, what he is trying to show may not be true, but that is a risk
that any new fundamental idea in physics faces.

Ed Fredkin’s style is not typical of scientists. He is more used to
addressing boards of directors than lecture audiences. He learned
the kind of physics that is in the Feynman lectures by spending time
with Dick Feynman rather than reading his books. To some standard
scientists, Fredkin at first seems like a nut. To be sure, some of his
ideas are pretty nutty. But if you listen and think about it, there
is much substance to what Fredkin has to say.

I gather that Fredkin has decided to spend some time around “ordinary
physicists”, to try and work out how his ideas fit in with current
physical thinking. I believe you are very lucky that Fredkin wants
to do this in your department.

Best wishes,
Stephen

And so it was that Ed became a research professor of physics at Boston University (BU). At MIT he’d gotten a DARPA grant that supported Tom Toffoli and Ed’s only “physics PhD student” Norm Margolus in building ever-larger “cellular automaton machines”. And when Ed moved to BU, this effort moved with him, leaving in effect “no trace of Ed” at MIT.

When Ed arrived at BU he found he was assigned to an office with a certain Gerard ‘t Hooft—who happens to be one of the more creative and productive theoretical physicists of the past half-century (and would win a Nobel Prize in 1999 for his efforts). Ed became friends with ‘t Hooft, inviting him and his family to spend time on his island, and later on the boat that Ed bought in the south of France. Feynman died in 1988, and Ed would tell me that he thought he’d “traded” one great physicist for another. (Feynman had suggested Ed try Sidney Coleman, but Coleman wasn’t into it.)

Like Feynman, I think ‘t Hooft felt a little uneasy with Ed’s statements about physics. But in 2016 ‘t Hooft ended up publishing a book entitled The Cellular Automaton Interpretation of Quantum Mechanics. I thought it was a nice recognition of ‘t Hooft’s friendship with Ed. But Ed told me in no uncertain terms that he thought ‘t Hooft hadn’t given him the credit he was due—though in reality I don’t think what ‘t Hooft did was much related to Ed’s actual work and ideas. (And, by the way, it’s not directly related to my efforts either, though conceivably looking at “generational states” in our Physics Project may give something at least somewhat analogous.)

In 1994 Ed’s direct affiliation with BU ended—though he remained on good terms with the department, and after I moved to the Boston area in 2002 I would often see him at an annual dinner the BU physics department put on for “Boston-area physics people”.

In 1998 Ed would summarize himself like this:

Ed Fredkin has worked with a number of companies in the computer field and has held academic positions at a number of universities. He is a computer programmer, a pilot, advisor to businesses and an amatuer [sic] physicist. His main interests concern digital computer like models of basic processes in physics.

For a while, Ed didn’t have a “university affiliation” (except, through Minsky, as a visitor at the MIT Media Lab), but in 2003—through his friend Raj Reddy—he became a professor (now of computer science) at Carnegie Mellon University, for a while spending time at their West Coast outpost, but mostly just making occasional trips in his plane to Pittsburgh.

Forty Years of Interactions with Ed

For a few years after I first met Ed in 1982, I’d see him fairly regularly. In 1983 I invited him to the first “modern” conference on cellular automata, that I co-organized at Los Alamos. I visited his house in Brookline, MA, a few times. I saw him at the Aspen Center for Physics, and at other places around the world. He was always fun and lively—and told great stories about all sorts of things. He gave the impression that he was mostly spending his time doing big things in business, and that science was an avocation for him. Sometimes he would talk about cellular automata—though I now realize that what he said was either very general and philosophical (leaving me to interpret things in my own way), or very specific to particular rules he’d engineered.

It was always a bit uncomfortable when it came to physics. Because the things Ed was saying always seemed to me pretty naive. Quite often I would challenge them—and frustratedly tell Ed that he should learn twentieth-century physics. But Ed would glide over it—and be off telling some other (engaging) story, or some such.

In 1986 I co-organized (with Tom Toffoli and Charles Bennett) a conference called Cellular Automata ’86—at MIT. Ed didn’t come—and I think I had the impression that he’d rather lost interest in cellular automata by that time. I myself went off to start my Center for Complex Systems Research, and then to found Wolfram Research and start the development of Mathematica. Mathematica was released on June 23, 1988—and our records (yes, we’ve kept them!) show that Ed registered his first copy on December 14, 1988. In March 1991 I did a lecture tour about Mathematica 2.0, and saw Ed one last time before diving into work on my book A New Kind of Science—which led me for more than a decade to became an almost complete scientific hermit.

I saw Ed (now 62 years old) when I briefly “came up for air” in connection with the release of Mathematica 3.0 in 1996, and we continued occasionally to exchange pleasant emails:

Date: Sun, 29 Jun 1997 15:49:41 -0400
From: Ed Fredkin
To: Stephen Wolfram

[Reporting the birth of my second child]

For many children its worst when they are teenagers. Some glide through
that period of life without hassle. Rick is doing great (at 15) despite
his unorthodox education. He relishes calling his parents dopes, but
aside from arguments about subjects like how late he should be able to
hang out with his buddies, its clear that he doesn’t think we’re dopes.

I promise to read your book as soon as I get it!

Its nice to hear from you. News here is that I am no longer needed at
Radnet as they now have a great CEO. I got a new airplane in December.
It’s called a Cessna CitationJet. It can carry 7 people at about 440
mph. So far its been a lot of fun. We’ll have to think of an excuse to
go for a ride. We are planning to spend some time at Drake’s Anchorage
in July. Its great for kids so if that interests you, let me or Joyce
know.

I have taken as a challenge to architect a computer (that weighs a few
kilos) that assumes another 100 years of Moore’s Law (10^15 in cost
performance). There are a lot of unsuspected problems lurking in the
details, but everyone of them seems to have easy solutions. I have
given a number of talks (IBM Almaden and Watson labs, Intel, NYU,
etc.). Interest in reversible computing has picked up since heat
dissipation has gotten to be a really hot topic (no pun intended). The
next high end Alpha may dissipate as much as 150 watts. Think of a
light bulb!

I use Mathematica for something almost every week… keep it up!

Best regards,

Ed

Although I didn’t see Ed myself for quite a few years, Ed would always write to ask for betas of new versions of Mathematica, and he would sometimes chat with staff from my company at trade shows. I thought it a bit odd in 1999 when I heard that in such an encounter he said that he was the one who had “introduced me to cellular automata”. And, moreover, that he, Feynman and Murray [Gell-Mann] were the people who’d suggested I write SMP—which was particularly bizarre since, among other things, I hadn’t met (or even heard of) Ed until about 3 years later.

Then, out of the blue on September 13, 2000, Ed calls my assistant, and follows up with an email:

Subject: Invitation
Date: Wed, 13 Sep 2000 23:53:09 -0400
From: Ed Fredkin
To: Stephen Wolfram

Hi,

The primary reason I’m contacting you has to do with a program I’m
organizing at Carnegie Mellon (CMU). I wrote a proposal to the NSF, called
“The Digital Perspective” and got funded. The idea is to invite a number (8
to 10) of guests to come to CMU for a few days, to meet with students and to
give a Distinguished Lecture. The NSF would also like to arrange for the
guests to come to Washington D.C. and give the same lecture there.

By “Digital Perspective” I mean looking at aspects of the world as Digital
Processes. As you know, I am most interested in looking at physics this
way. I have just started getting commitments from potential participants.
Gerard ‘t Hooft has agreed to come and a number of other good physicists are
thinking about it.

Please consider this to be a formal invitation. Of course, CMU will pay
expenses and an honorarium. If the timing works out, it can probably be
arranged for many of the students to have read your book before you come.
You might get some good feedback from bright students who have also gained
familiarity with the thoughts of others who are thinking about the “Digital
Perspective”. The seminar will run throughout the 2000-2001 academic year.

If you can make it to CMU, I expect that it will be fun and interesting;
both for you, for me and for many others.

I responded:

Subject: RE: Invitation
Date: Thu, 14 Sep 2000 06:49:19 -0500
From: Stephen Wolfram
To: Ed Fredkin

Thanks for the invitation, etc.

It sounds like a thing I’d like to do, but I can only consider
*anything* after my book is finished.

If my book is done in time for your program, then, yes, I’d like to
participate (though of course I’d want more details about the actual
plans etc. etc.). But if the book isn’t done, then sadly I just can’t.
If the cutoff time is June 2001, I am not extremely hopeful that the
book will be done … but if it’s fall 2001 the probabilities go up
substantially (though, sadly, they are still not 100%).

And what are you up to these days? Business? Science? Other?

On another topic:
In my book, I’m trying very hard to write accurate history notes about
the things I discuss. And for the notes on the history of cellular
automata I’ve been meaning for ages to ask you some questions…

I’m not sure this is a complete list, but here are a few I’ve been
curious about for a long time that I’d really like to know the answers
to…

I know that history is hard … even if it’s about oneself. I consider
that I have a good memory, but it’s often hard for me to keep straight
what happened when, and why, etc. But anything you can tell me about
these questions … or about other aspects of CA history … I’d be very
grateful for.

1. As far as you know, did you invent the 2D XOR CA rule? (I’m assuming
the answer is “yes”…)

2. In what year did you first simulate this CA? On what computer?
Where?

3. What other CA rules did you study at that time?

4. Do you still have any material from the simulations you did
(printouts, tapes, programs, etc.)?

5. When you learn about the “munching squares” display hack? How did it
relate to your work on the XOR CA?

6. What did you know about the work done by Unger etc. on cellular image
processors? How did this relate to your work?

7. What did you know about von Neumann’s work on cellular automata? How
did it relate to your work?

8. What did you know about Ulam and others’ work at Los Alamos on
simulating cellular automata? How did it relate to your work?

9. Were you aware of work on cryptographic applications of CA-like
systems?

Ed responded:

Subject: RE: Invitation
Date: Fri, 15 Sep 2000 01:25:23 -0400
From: Ed Fredkin
To: Stephen Wolfram

Hi,

Here are some answers and some free association type ramblings.

> And what are you up to these days? Business? Science? Other?

I’m winding down on business (I’m into one last e-business project) and like
you, working on a book. My guess is that mine is nowhere as ambitious as
yours… It’s just to document my ideas about Digital Mechanics (Physics).
In any case, these ideas have made more progress in the last 2 years than in
the previous 40.

I bought a sailboat which is moored in Antibes, France. I spent most of the
summer there and got more science done than in the prior several years.
It’s absolutely the perfect place and circumstance for me to work on my
stuff. Gerard ‘t Hooft (plus wife and daughter) came down and joined us for
a while. You know (I hope) about his interest in CA’s? I’m going back
there for a few weeks on Tuesday.

Here’s a formal proof that you can, at any time, escape all your normal
responsibilities and concentrate exclusively on one really important thing
(hint, hint). The proof is that, at any time, YOU CAN DIE. I don’t mean to
be morbid, but sometimes it makes good sense to consider that proof and
temporarily abandon all but some very important task (or some very exciting
or fun thing).

Ed continued with a long response to my “history questionnaire”:

> 1. As far as you know, did you invent the 2D XOR CA rule? (I’m assuming
> the answer is “yes”…)

Yes, as far as I know I did invent it. Here is what I did. I decided to
look for the simplest possible rule that met certain criteria. I wanted
spatial symmetry and a symmetric rule vis-à-vis the states of the cells.
The thought was to find something so simple that its behavior could be
understood while not so simple as to be totally dull. The first such rule I
tried was the XOR rule. I programmed it first on the PDP-1 (1961, at BBN
and III) where I could see it on the display, and later I wrote a program
for CTSS using a model 33 teletype as a terminal. My motivation was then,
as it is now, to be able to capture more and more properties of physics
within a Digital model. I found an easy proof as to why patterns reappeared
in any number of dimensions. I also found, at the beginning, a formula for
the number of ones as a function of time from a single one as the initial
state. My recollection was that it was something like 2D 2^b(t) where D is
the number of dimensions, t is the time step, and b(t) is the number of bits
that are one in the binary representation of t (the tally function). After
I showed all this to Seymour Papert, he generalized the proof re self
replication from XOR (sum mod 2) to sum mod any prime. (Some time around
1967)

> 2. In what year did you first simulate this CA? On what computer?

Where?

See above.

> 3. What other CA rules did you study at that time?

I found a simple proof that a von Neumann neighborhood CA could exactly
emulate any other (such as the 3×3 neighborhood) and used this as a reason
to look at nothing else. I explored so many different rules that I probably
would have found the game of Life had I not put blinders on. After I came
to MIT (1968), I had 2 things in mind, to find a really simple Universal CA
(I call them UCA’s )and to find Reversible, Universal CA’s (RUCA’s)
As you may know, the search for UCA’s went slowly until I had the idea to
abandon the Turing Machine model and look at modeling digital logic and
wires. Within 15 minutes after this idea occurred to me, I had a 4 state
UCA on my blackboard. At that time the best known was in Codd’s thesis; an
8 state UCA. I showed this to a student of mine, Roger Banks, who had been
struggling for a few years trying to complete an AI PhD thesis. The next
morning both he and I showed up with 3 state UCA’s. He switched his PhD topic
and found a 2 state, von Neumann neighborhood UCA, a thing that Codd
purported to have proved impossible.

While at BBN, after seeing all my 2-D CA’s expanding with simple
kaleidoscope like symmetries, (like the diamond shapes in the XOR rule),
Marvin Minsky challenged me to find a rule (any rule) that showed spherical
propagation. I took the challenge and shortly came up with such a rule.

With respect to reversibility, the first satisfactory RUCA was done by
Norman Margolus. I shortly thereafter found a simple RUCA that didn’t need
the use of the Margolus Neighborhood trick.

> 4. Do you still have any material from the simulations you did
> (printouts, tapes, programs, etc.)?

Yes, Probably, quite a bit

> 5. When you learn about the “munching squares” display hack? How did it
> relate to your work on the XOR CA?

I don’t recall it having any effect. It’s very unlikely that I knew of it
prior to the XOR CA.

> 6. What did you know about the work done by Unger etc. on cellular image
> processors? How did this relate to your work?

I knew of it second hand, but I don’t think it had any effect. Do you know
about Farley and Clark (Wes Clark) and their publication while at MIT’s
Lincoln Labs in the late 50’s?

> 7. What did you know about von Neumann’s work on cellular automata? How
> did it relate to your work?

At the time I did the XOR work I had not read anything about the von Neumann
CA, but I was told about it and I understood the concept very well. Many
years later I read something (by Burkes, I think). I remember knowing that
it was a 29 state system and that it knew left from right in order to extend
and turn its construction arm.

> 8. What did you know about Ulam and others’ work at Los Alamos on
> simulating cellular automata? How did it relate to your work?

All I knew about Ulam and CA is that, like the Hydrogen Bomb, he had key
ideas but probably didn’t get as much credit as he deserved. All my
knowledge re Ulam was anecdotal. As to what he did vs. what von Neumann did
I didn’t really know anything.
I didn’t know anything about anyone else actually simulating CA’s however
I’m pretty sure I assumed that others must have done so. It was so easy and
so obvious. While the use of a computer with a display (such as the Lincoln
Lab TX-0 and TX-2, the Digital PDP-1 and the IBM 709 and 7090 all had or
could have CRT displays, it was easy enough to display simple CA’s with a
printer, even a 10 CPS teletype.

> 9. Were you aware of work on cryptographic applications of CA-like
> systems?

I thought I invented that idea! As soon as I found ways to make RUCA’s it
occurred to me that they could be used for cryptography. As an aside, when
Witt Diffey [Whit Diffie] came up with the idea of public key cryptography,
which needed a trapdoor function, I thought of using the product of 2 large primes.
I had just written the first program, in LISP, to implement Michael Rabin’s first
version of a probabilistic prime test. As soon as I implemented it I
started a search at 10^100 and discovered that 10^100 +35,737 and 10^100
+35,739 were prime. A week later I met Rich Schroeppel in LA (he was
working for my company, III) and knowing a larger prime pair than anyone
else on Earth I told Rich and he was blown away. He was seated at a PDP-10
terminal and all he said was an emphatic “Really!” He then went type, type,
type for a few seconds and turned around and said “You’re right!” which blew
me away! I asked what he did and he said (while knowing nothing of Rabin’s
method) “all I did was look at 3^(n-1) mod n, you know, Fermat’s little
theorem, it usually gives 1 for primes.”

I’m rambling, probably about stuff of no interest to you. Anyway, I stopped
Ron Rivest in the hallway at Tech Square and asked if he had heard of
Diffey’s [Diffie’s] stuff. I don’t remember exactly what he said but I know that when
I told him that Rabin’s new method to find large primes meant that the
product of 2 primes was a good trapdoor function he was surprised and
thought it was a good idea! I never thought any more about it and hadn’t
come up with the idea of using the phi function… Years later, long after
RSA was a big thing Ron reminded me of the event… Don Knuth told me that
he also thought of using the product of 2 primes before RSA, but he couldn’t
have known about Rabin’s method when I did (as Rabin told it to me right
when he thought it up!)

By the way, I have an interesting algorithm for factoring smaller numbers,
such as can be done in less than an hour with Mathematica (normal
FactorInteger or ECM). I’ve written a few terribly unoptimized Mathematica
functions that implement the method. For what its good for, my Mathematica
functions (not compiled or anything) make Mathematica factor in a lot less
real time than Mathematica does with FactorInteger or ECM.

The big news re me and my work is what’s happening right now. Whatever one
thinks about my stuff (Digital Mechanics), it’s vastly improved. However
it’s still very far from a complete theory. Of course, Digital Mechanics is
about CA’s.

If you have any interest in reversibility, I’ve done lots in that area,
ranging from RUCAs, conservative logic, and my transforms. The transforms
are general methods of converting algorithms that calculate the approximate
time evolution of a system (approximate because of round off, truncation and
the finite delta t) which is approximately reversible (by changing delta t
to minus delta t) into an equivalent algorithm that calculates approximately
the same thing going forwards, but which is exactly reversible (being
calculated on a computer with round off and truncation error). I also have
a lot of methods for making RUCA’s with particular properties.

You’ve criticized me in the past for not publishing stuff, but I’m so
ambitious as to what I’m trying to do that I haven’t had the motivation to
publish all the little things I’ve uncovered along the way.

I’m sure I discovered more and better ways to make all kinds of RUCAs before
anyone else with the exception of the rule found by my student, Margolus.

Finally, one last anecdote. You and I were at some meeting long ago (maybe
Santa Fe?) and you brought along an early Sun to demonstrate your collection
of different kinds of 1-D CA’s. After your talk, I asked you why none of
the CA’s you showed were reversible. Your response was “Because all
reversible CA’s are trivial.” That really was a very common belief,
coincident with most people’s intuition. On the spot, I made up a rule,
using your convention for specifying it, of an “interesting” reversible CA.
You typed it in and ran it. Being surprised is one of the best kinds of
experiences we ever have.

As Emerson once quipped, “My apologies for such a long email, I didn’t have
the time to write you a short one.”

I’m having fun; it’s a good thing to do!

Best regards

Ed F

PS If you have any interest in having parts of your book read so that you
can get comments prior to publication, I have an idea that might be useful.

A little later he added:

Subject: error
Date: Fri, 15 Sep 2000 10:12:15 -0400
From: Ed Fredkin
To: Stephen Wolfram

Hi,

Looking at my long email I noticed a boo boo.

Where I wrote, quoting Schroeppel talking about 3^(n-1) mod n, “…it
usually give a 1 for primes…” very true but a bit of an understatement.
Of course, it ALWAYS gives a 1 for primes! What Schroeppel said was that it
usually doesn’t give a 1 for non-primes. It’s incorrect for 91 and 121 and
lots of other small numbers, but seems to work better for large numbers…
but then you probably know much more about such things than I do. Also
looking at your questions, I had the feeling that some might have been
prompted by my circa 1990 Digital Mechanics paper. If so, I guess I
repeated stuff already in the paper and I apologise.

Regards,

Ed F

I responded, asking for various pieces of clarification (and now that I’m writing this piece I would have asked even more, because some key parts of what Ed said I now realize don’t add up):

Subject: Re: your mail
Date: Wed, 20 Sep 2000 21:04:54 -0500
From: Stephen Wolfram
To: Ed Fredkin

>> 1. As far as you know, did you invent the 2D XOR CA rule?
>>

> Yes, as far as I know I did invent it. Here is what I did. ….
>

Very interesting.

1a. Did you ever look at 1D CAs? If not, why not?

1b. Did you think about analogies between XOR rules and linear feedback
shift registers?

1c. Did you think about analogies between XOR rules and Pascal’s
triangle?

By the way, the result about the number of binomial coefficients mod a
prime has been independently discovered a remarkable number of times
(including by me). The earliest references I know are Edouard Lucas
(1877) and James Glaisher (1899).

….

>> 3. What other CA rules did you study at that time?

> … I explored so many different rules that I probably
would have found the game of Life had I not put blinders on.

By the way, I happened to have a long phone conversation recently with
John Conway about the history of the Game of Life. I still haven’t
quite got to the bottom of exactly what Conway was doing and why (I
think he wants some of the history lost, which is a pity, because it is
interesting and reflects much better on him than he seems to
believe…) But what is clear is that Conway (and his various helpers)
had much more serious motivations from recursive function theory etc.
than is ever usually mentioned. It was just not a “find an amusing
game” etc. piece of work.

> Marvin Minsky challenged me to find a rule (any rule) that showed spherical
propagation. I took the challenge and shortly came up with such a rule.

I don’t believe I’ve ever seen your rule of this kind. I showed such a
rule to Marvin in 1984 and he said “that’s very interesting; we were
looking for these but hadn’t found any”. So I’m confused about
this….

>> 6. What did you know about the work done by Unger etc. on cellular image
processors? How did this relate to your work?

> I knew of it second hand, but I don’t think it had any effect.

Wasn’t BBN quite involved with cellular image processing? And I believe
you worked on aerial photography analysis. Did you use cellular
automata for image processing?

>> 9. Were you aware of work on cryptographic applications of CA-like
systems?

> I thought I invented that idea!

There was a lot of work done on 1D CAs by some distinguished
mathematicians consulting for the NSA in the late 1950s. I think much
of it is still classified. But over the years I’ve talked to many of
the people involved (Gustav Hedlund, Andrew Gleason, John Milnor, some
NSA folk, etc. etc.), and read their unclassified papers. They figured
out some interesting stuff. They thought of it as related to nonlinear
feedback shift registers.

> As soon as I found ways to make RUCA’s it
occurred to me that they could be used for cryptography.

How?

There’s a 1D CA (rule 30) that I studied in 1984 that has been
extensively used as a randomness generator (e.g. Random[Integer] in
Mathematica uses it), and that has been used a bit as a cryptosystem.

I tried to make a good public key system out of CAs in the mid-1980s
(mostly in collaboration with John Milnor), but did not come up with
anything satisfactory. …

> I also have a lot of methods for making RUCA’s with particular properties.

I am definitely somewhat interested in these things. They don’t happen
to be central to my grand scheme. But they are obviously worthwhile …
AND WORTH (you) WRITING DOWN!!

I’m sure I discovered more and better ways to make all kinds of RUCAs before
anyone else with the exception of the rule found by my student, Margolus.

Interesting. You probably know that the general problem of telling
whether an arbitrary 2D CA is reversible is undecidable (the question
can be mapped to the tiling problem).

So I’m taking it that you have some good methods for generating 2D
reversible CAs. That’s obviously interesting.

> Finally, one last anecdote. … I asked you why none of
the CA’s you showed were reversible. Your response was “Because all
reversible CA’s are trivial.” …

This anecdote can’t be quite right. I have known since 1982 that there
are nontrivial things that can happen in CAs that are made reversible by
your mod 2 trick. What is true (and may have been what I was saying)
is that none of the 2-color nearest neighbor CAs that are reversible are
non-trivial. With more colors or more neighbors, that changes. I’m
guessing that what you showed me was a 4-color nearest neighbor CA that
is reversible … and that is of course quite easy to get by recoding a
2-color one that has your mod 2 trick.

By the way, I heard third hand a while back that you had “introduced me
to CAs”. For what it’s worth, that isn’t correct. My first “CA
experience” was actually in 1973 (when I was 13) when I tried to program
molecular dynamics on a very small computer, and ended up with something
equivalent to the square CA fluid model. My next CA experience was in
summer 1981. I was trying to make models of “self organizing” systems
(now I hate that term), particularly self-gravitating gases. I ended up
simplifying the models until I got 1D CAs. That fall I spent a month at
the Institute for Advanced Study, and spent a lot of time studying von
Neumann’s work, etc., and analysing all sorts of features of 1D CAs. I
came for a day to give a talk at MIT, and was having dinner with some
LCS people (Rich Zippel was one of them), and they told me about your
work. Later that fall I talked with Feynman a certain amount about what
I was doing with CAs, and he again mentioned you. (I think he had been
to your Physics of Computation meeting, which was perhaps in June 1981,
but I didn’t discuss the CA aspects of the meeting with him.) Then in
[January 1982] I came to the meeting you had on your island, and Tom Toffoli
showed me his 2D CA machine (at the time he gave me the impression of
95% hackery, 5% science), and you showed me the 2D XOR CA on a PERQ
computer.

Ed didn’t respond to this, but three days later we talked on the phone. I sent some (unvarnished) notes from the call to a research assistant of mine:

Subject: Fredkin conversation
Date: Sat, 23 Sep 2000 03:05:43 -0500
From: Stephen Wolfram

I had a long conversation with Ed this evening.

About his work in science, my work in science, etc.

A few things mentioned:

– He feels bitter that his paper on reversible logic, coauthored with
Tom Toffoli, was actually all his (Ed’s) work

– He is pleased that I will discuss history even when people haven’t
published things (of course he has published little)

– He says he has written about 150 pages about his views of physics; he
is planning to prepare something, perhaps for publication, in about a
year

– He says he missed not being able to bounce ideas off Dick Feynman …
even though Feynman often ended up screaming at him (Ed) about how dumb
his ideas were

– He said that his main problem was that he has been trying to get
people to steal his ideas for years, but nobody was interested

– He said that now “for some reason” he is becoming more concerned about
matters of credit

– He is a serious fan of Mathematica, the Mathematica Book, etc.

– He made an effort again (he’s been trying for 20 years) to get me to
coauthor a paper with him. He recognizes that he can’t write a credible
scientific paper, but he’s “sure he has some ideas I haven’t thought
of”. I told him that unfortunately I haven’t written a paper for 15
years.

– I told him that particularly when I’m in the Boston area, I’ll look
forward to chatting with him about physics etc.

– He said he’s tried to interact some with Gerhardt ‘t Hooft, but that
‘t Hooft keeps on rushing off in traditional physics directions that Ed
(and I, by the way) think are stupid

– He wanted to know if I really believed that all of physics etc. was
ultimately discrete; he expressed the opinion that he and I may be the
only people in the world who actually believe that right now

– He told a bizarre story about how Don Knuth gave a talk at MIT
recently on computers and religion, and how 1/4 of it was stuff that Don
had heard about from Ed. Apparently Guy Steele asked a question about
how Don’s stuff related to Ed’s, and Don said something meaningless.

I talked to him a little more about the CA history stuff. He mentioned
that around 1961 a certain Henry Stommel (sp?) told him that CA-like
models had been used in studying sand dunes in the 1930s. I have a
feeling this may be another cat gut search, but perhaps we can follow
up. (You could email Ed at the appropriate time.)

I asked Ed if he had ever looked at cryptography (as in NSA style stuff)
with CAs. He said no. But that in the late 1960’s he had had a student
who had studied ways to make counters out of JK flip flops … and that
that person’s work had made something that Ed thought could be used for
cryptography. This was followed up by a certain Vera Pless
subsequently.

I didn’t hear anything more from Ed for a while, though a public records search indicates that, yes, he had successfully “worked the system” to get $100k from the NSF for “The Digital Perspective Project”. And on May 1, 2001, I received a rather formal email from Ed (for some reason Americans born before about 1955 seem to reflexively call me “Steve”):

Subject: Workshop on the Digital Perspective 24-26 July, Washington DC
Date: Tue, 1 May 2001 21:20:45 -0400
From: Ed Fredkin
To: Steve Wolfram

We are sending this email to invite you to an NSF-sponsored workshop on the
Digital Perspective in Physics planned for July 24th through the 26th,
Tuesday, Wednesday and Thursday. It will be held in the NSF building,
Arlington Virginia. Gerard ‘t Hooft has already agreed to present a paper
and we hope that you will also be willing to contribute. We intend to
combine the papers presented at the workshop into a monograph that will be
published later this year. Two earlier workshops on related subjects were
held at Moskito Island and this was a central theme at a meeting held at
MIT’s Endicott house in 1982. Participants at previous meetings included
Charles Bennett, Richard Feynman, Ed Fredkin, Leo Kadanoff, Rolf Landauer,
Norman Margolus, Tomasso Toffoli, John Wheeler, Ken Wilson, Stephen Wolfram
and others.

When I didn’t immediately respond, Ed called my assistant, saying that he was “calling regarding a meeting he spoke with [me] about on the phone”. I responded by email later the same day:

Subject: I gather you called…
Date: Tue, 15 May 2001 15:18:57 -0500
From: Stephen Wolfram
To: Ed Fredkin

Sorry for not getting back to you sooner….

I myself am right now trying to work at absolutely full capacity to finish my
book/project. I haven’t done any travelling at all for a long time, and won’t
until my book is done.

And I also don’t yet have anything public to say about my work on physics.

Hopefully by the end of the year my book will be done and I will have quite a
bit to say.

However, it occurs to me that one or two of my assistants might be very good
people to come to your workshop.

Who all is coming?

One person you should definitely invite is someone who has been an assistant of
mine, and now works part time for me, and part time on his own projects. His
name is David Hillman, and he’s been interested in discrete models of spacetime
for a long time. (He got his PhD working on some kind of generalization of
cellular automata intended as a spacetime model.)

I have two physics assistants, and one math one, who might be relevant for your
workshop.

Just let me know in more detail who might be coming, and I’ll try to figure out
the correct person/people to suggest.

Of course I’d love to come myself if I were a free man. But not until the book
is done.

In haste,

— Stephen

Ed responded pleasantly enough:

Subject: RE: I gather you called…
Date: Tue, 15 May 2001 17:08:39 -0400
From: Ed Fredkin
To: Stephen Wolfram

Hi,

Sorry you can’t make it.

About half of those coming are veterans of some Moskito Island workshop.
Newcomers include Gerry Sussman, Tom Knight, Gerard ‘t Hooft, John Negele,
John Conway, Raj Reddy, Jack Wisdom, Seth Lloyd, David di Vincenzo, plus a
number of students, etc. A couple of those mentioned are still struggling
with scheduling issues.

But, in any case, I would be pleased to have David Hillman come to the
workshop. Send me his email address and I will send him an invitation.

Best regards and good luck on the book!

Ed

I responded and suggested an additional person from our team for his workshop. Nearly a month passed with no word from Ed, so I pinged him asking what was going on. No response. It was a very busy time for me, and this wasn’t something I wanted to be chasing (I saw myself as doing Ed a favor by suggesting sending people to his workshop) … so I sent a slightly exasperated email:

Subject: your conference, again
Date: Fri, 15 Jun 2001 06:08:39 -0500
From: Stephen Wolfram
To: Ed Fredkin

Look … I’m now in a bit of an embarrassing situation: following your
initial response, I told David Hillman and David Reiss about your conference
… assuming you’d want to invite them … and they both became quite
interested in it. But they never heard from anyone about it. So now of
course they’re wondering what’s going on. And so am I. What should I tell
them? I’m now embarrassed about having suggested this…

This seems peculiarly un-you-like. I was thinking you must have been away
or something. But isn’t the conference coming up very soon?

I hope everything’s OK…

Still no response from Ed. A week later I called him, and we talked for two hours. It wasn’t clear why he hadn’t already reached out to the people I’d suggested, but he quickly said he would. And then Ed launched into telling me about the “astounding” cellular automaton models he said he’d just created that “had charge, energy, momentum, angular momentum, etc.”. He talked about things like the idea of what he called an “infoton” that would be an “information particle” that would “make Feynman diagrams reversible”. I explained why that didn’t make any real sense given how Feynman diagrams actually work. It was the same kind of conversation I had many times with Ed. I kept trying to explain what was known in physics, and he kept on coming back with things that, yes, I think I understood, but that seemed close to typical crackpot fare to me. But Ed seemed convinced he had discovered something great (though exactly what I couldn’t divine). And eventually—having obviously not convinced me of what he was doing “on its merits”—he just came out and said “It must be related to stuff you’re doing, one way or another”.

I explained that I really didn’t think that was very likely, not least because I emphatically wasn’t trying to use cellular automata as models of fundamental physics. And with that, Ed launched into a long speech about giving credit, particularly to him. I explained that I was trying hard to write correct history, and reiterated some of the questions I’d asked him before. He didn’t really tell me more, but instead regaled me with stories (that I’d mostly heard many times before from him) about how he’d been the first to figure out this and that—apparently oblivious to historical research I tried to tell him. But eventually we both had to go—and the conversation ended pleasantly enough, with him confirming the email addresses for the two people for his workshop.

As the workshop approached, the people from my team had made arrangements to go to Washington, DC—but still didn’t know where exactly the workshop was. With days to go, one of them simply called Ed to ask. But Ed told them that actually they couldn’t come, because “Raj Reddy says there is no room for you”. Really? No extra chair to be found? Ed was the organizer, wasn’t he? Why was he laying this on someone else? It seemed to me that Ed was playing some kind of game. But at that moment I was too busy trying to finish my book to think about it. (Now that I’m writing this piece, however, I realize that Ed was perhaps following an “algorithm” he’d established years earlier when he was proud to have organized a meeting to push forward his ideas about timesharing—by inviting just people who supported his ideas, and not inviting ones who didn’t. I don’t know if the meeting actually happened, or what went on there. I don’t think the writeup promised in the invitation and in the NSF contract ever materialized.)

In January 2002 A New Kind of Science was off to the printer, and review copies were starting to be sent out. In late March a seasoned journalist named Steven Levy (who had written about my work on cellular automata in the mid-1980s) was talking to someone from my team and reported that Ed had told him that “Minsky had told [Ed] to publish his stuff on the web to stake out priority” before my book came out. (And it’s a pity Ed didn’t do that, because it might have made it clear to him and everyone else how different what he was saying was from what I was saying.) But in any case Levy said that Ed seemed to be saying the same things as he’d said 15 years ago—and Levy knew that regardless of anything else I’d done incredibly much more since then.

After his conversation with Levy, Ed sent me mail:

Subject: The Book
Date: Fri, 22 Mar 2002 16:38:29 -0800 (PST)
From: Ed Fredkin
To: Stephen Wolfram

Congratulations on finishing!!!

I ordered the book from someplace, so long ago I can’t
remember from who. I’m wondering if, when its
possible, I could get a copy in advance of whenever my
ordered copy is going to appear. I just don’t want to
be the last on the block to see it. Of course I’d be
happy to pay if you can tell me how to do it.

Thanks,

Ed Fredkin

The book was going to be published on May 14; on May 4 I signed a copy for Ed:

Click to enlarge

The book mentioned Ed a total of 7 times. (The person with the most mentions overall was Alan Turing, at 19; Minsky had 13; Feynman 10.)

Ed never told me he’d received the book. And I’m not sure he ever seriously looked at it. But somehow he was convinced that since he knew it talked a lot about cellular automata, and had a section about physics, it must be about his big idea—that the universe is a cellular automaton. As one witty friend pointed out to me in connection with writing this piece, my book says only one thing about the universe being a cellular automaton: that it isn’t! But in any case, Ed apparently seemed to feel that I was stealing credit from him for his big idea—and, as I now realize, started an urgent campaign to right the perceived wrong, basically by telling people that somehow (despite all my efforts to describe the history) I wasn’t giving anyone enough credit and that “he was there first”. The New York Times rather diplomatically quoted Ed as saying “For me this is a great event. Wolfram is the first significant person to believe in this stuff. I’ve been very lonely”. It followed up by saying that “Mr. Fredkin, who said he was a longtime friend, said Dr. Wolfram had ‘an egregious lack of humility’”. (In some contexts, I suppose that might be a compliment.)

In writing this piece I asked Steven Levy what Ed had actually said in the interview he did. His first summary in reviewing his notes was “He says he considers you a friend and then goes on endlessly about what an egomaniac you are”. But then he sent me his actual notes, and they’re somewhat revealing. Ed doesn’t claim he introduced me to cellular automata, perhaps because he realizes that Levy knows from the 1980s that that isn’t true. But then Ed tells the story about showing me reversible cellular automata, which I’d explained to Ed wasn’t true. Ed goes on to say that “Everyone who’s in science wants credit, driven probably by wanting to become famous. [Wolfram] has a larger than normal dose”. Ed says that when he had said that cellular automata underlie physics, I’d said that was crazy. (Yup, that’s true.) But then Ed said “Now he denies this”. Huh? Ed went on: “He’s a prisoner of some kind of overactive ego. I believe he might not know. Wolfram deserves loads & loads of credit, but he has this personality flaw”. And so on.

A month later Ed writes to me:

From: Ed Fredkin
To: Steve Wolfram
Sent: Friday, June 14, 2002 2:48 PM
Subject: ANKOS critics

Steve,

Sometime soon I’d like to get together and talk.

I’ve read a lot of your book.

Take a look at the draft of a little paper of mine (attached). I’d
appreciate comments.

Ed F

The following is my response someone else’s response [Gerry Sussman] to a review of ANKOS.

My comments are only with regard to Wolfram’s ideas on modeling physics.
I don’t happen to like his network model but we are in agreement that
some kind of discrete process might underlie QM.

Not everything Wolfram says is wrong.

The ideas that some kinds of discrete space-time processes (such as
CA’s) might underlie physics or other processes in nature is the BABY.
Everything else in ANKOS (or missing from ANKOS) is the BATH WATER.

Ed’s attached paper was basically yet another restatement of cellular automata as models of fundamental physics.

A few weeks later there was a strange (if in some ways charming) incident when a reporter for the San Francisco Chronicle decided to investigate what seemed to be a science feud between Ed and me. After a nod to medieval metaphysicians, the article (under the title “Cosmic Computer”) opens with “Nowadays, with a daring that might have dazzled St. Augustine and St. Thomas Aquinas, two titans of the computer world argue that everything in the universe is a kind of computer.” After analogizing me to Britney Spears, the article goes on to say “The excitement has also brought tension to the long-standing friendship between Wolfram and Fredkin, who are now wrestling with one of the bigger bummers of any scientist’s life: a dispute over originality.” The article reports: “Last week, the two men had a long, heartfelt phone conversation with each other, in which they tried to resolve their strong disagreement over priority. The conversation was amicable, but they failed to reach agreement.”

And so things remained until March 2003 when Ed sent the following:

Subject: Re: NKS 2003 Conference & Minicourse
Date: Thu, 20 Mar 2003 17:24:59 -0500
From: Edward Fredkin
To: Stephen Wolfram

Dear Stephen,

I guess I’m on a Wolfram mailing list for potential attendees for your
Boston conference. I hope you don’t mind a little plain speaking. I
consider that I am a friend of yours and therefor I take the risk of
telling the emperor about his new clothes. Of course, few others
would do so as a friend. Please don’t be offended as the plain talking
that follows is my attempt at trying to be constructive.

Your work is acquiring a reputation amongst the scientific community
that is much less than it deserves. I find myself often in the
position of defending you, your work and your accomplishments against
the negative views that many hold, even though they have little
understanding of the significance of what you have done. They are
turned off by your egregious behavior; it distracts much of the
scientific community rendering it barely possible for them to take you
seriously . You have invented and discovered quite a few things, but
so have others. You told me you would try to give credit in ANKOS
where credit was due; I believed you and I believe that you tried your
best but nevertheless you failed miserably. I guess you simply didn’t
know how. Consider this conference: Must this conference be a one man
show or might it actually be better for the ideas in ANKOS and better
for SW and his overall scientific reputation if it were a real
conference where others might address the same questions? Please don’t
kid yourself into thinking that no one else has anything original,
novel, important or interesting to say.

Of course, this so-called “…first ever conference” devoted to the “…
ideas and implications…” of concepts found in ANKOS might be nothing
more than a marketing tool for Mathematica and for sales of the ANKOS
book. If so, you ought to call a spade “spade”.

You’ve done enough things (and hopefully will continue to do so) to
ensure your reputation as a pioneer in various areas. This flood of
self puffery simply detracts, in the minds of many whose opinions you
ought to value, from the positive reputation you deserve.

I’m not one of those whose opinion of your work is in any way affected
by your unfortunate behavior. I see and understand exactly what you’ve
done and I know and understand what your work is based on. I am human,
so I find it interesting when you now and then claim to have discovered
an idea or fact that I personally explained to you when it was
perfectly clear at the time that to you, the idea was absolutely novel.
My model of you is that your overpowering motivation results in your
mind playing tricks on you. I really believe that you actually forget;
that you actually re-remember the past differently than it happened.

But I am the eternal optimist. I believe that even Stephen Wolfram
might someday come around and join the collegial scientific community
where you receive credit and give credit; both nearly effortlessly.
The world actually might voluntarily heap honors on you as opposed to
SW having to orchestrate “conferences” for the glorification of SW and
all the ideas claimed by SW. No one knows better than me how slow and
torturous this process can be for new and novel big concepts, but
patience and modestly [sic] still seems like the better path.

Please try to not be offended. I actually mean well. If you ever have
an actual, real conference, invite me to be a speaker; I’ll come. If I
organize another conference you can rest assured that you will be
invited again (as you were for the NSF Workshop) and I hope you will
come to talk about your ideas and maybe — maybe even stay to hear what
others have to say on the subject. It’s not too healthy to the
scientific mind to be the only real speaker at conferences you organize
and hype for yourself.

Among the very few who really are able to appreciate what you’ve done,
I am one of your greatest supporters. But I am not your average person
with more or less normal reactions. When you reach for extra glory
and credit by stealing one of my ideas, my reaction is: “I admire your
good taste”.

Best regards

Ed F

On Thursday, March 20, 2003, at 02:19 PM, Stephen Wolfram wrote:

> In June of this year we’re going to be holding the first-ever
> conference devoted to the ideas and implications of A NEW KIND OF
> SCIENCE. I think it’ll be an exciting and unique event. And if you’re
> interested in any facet of NKS or its implications, you should plan
> to come!
>
> I’ll be giving a series of in-depth lectures to explain the core
> ideas of NKS. There’ll be more specialized sessions exploring
> implications and applications in areas such as computer science,
> biology, social science, physics, mathematics, philosophy, and future
> of technology. And there’ll also be workshops and case studies about
> such issues as modelling, computer experimentation, defining NKS
> problems, NKS-based education–as well as a gallery of NKS-based
> art pieces.
>
> I’d expected that it’d be a few years before it would make sense to
> start having NKS conferences. But things have gone faster than I
> expected, and the enthusiasm and energy we’ve seen in the ten months
> since the book was published has made it clear that it’s time to have
> the first NKS conference.
>
> In planning NKS 2003, we want to cater to as broad a range of
> attendees as possible. There’ll be many professional scientists
> coming, as well as technologists and other researchers from a very
> wide range of fields. There’ll also be a large number of educators
> and students, as well as all sorts of individuals with general
> interests in the ideas and implications of NKS.
>
> We’ll be holding NKS 2003 near Boston over the weekend of June 27-29,
> 2003. There’s more information and registration details at
> http://www.wolframscience.com/nks2003
>
> It’s going to be an extremely stimulating weekend–and a unique
> opportunity to meet a broad cross-section of people interested in new
> ideas.
>
> I hope you’ll be able to be part of this pioneering event!
>
>
> — Stephen Wolfram

I responded:

Subject: Re: NKS 2003 Conference & Minicourse
Date: Sat, 22 Mar 2003 22:19:33 -0500
From: Stephen Wolfram
To: Edward Fredkin

Ed —

I must say that I am reluctant to respond to a note like the one below, but
it seems a pity to let things end this way.

I can tell you’re very angry … but beyond that I really can’t tell too
much.

I’d always thought we had a fine, largely social, relationship. We talked
about many kinds of things. It was fun. Occasionally we talked about
science. In the early 1980s I learned a few things about cellular automata
from you. None were extremely influential to me, but they were fine things
that you should be proud of having figured out—and in fact I took some
trouble to mention them in the notes to NKS.

You also told me some of your thinking about fundamental physics. I was (I
hope) polite, and tried to be helpful. But I always found what you were
saying quite vague and slippery—and when it became definite it usually
seemed very naive. I think it’s a great pity that you’ve never taken the
time to learn the technical details of physics as it’s currently practiced.
There’s a lot known. And if you understood it, I think you’d be able to
tell quite quickly which of your ideas are totally naive, and which might
actually be interesting.

I think it’s also a pity that—so far as I can tell—you’ve never really
taken the time to understand what I’ve done. It’s in the end pretty
nontrivial stuff. It’s not just saying something like “the universe is a
cellular automaton” or “I have a philosophy that the universe is like a
computer”. It’s a big and rich intellectual structure, built on a lot of
solid results and detailed, careful, analysis. That among other things
happens to give a bunch of ideas about how physics might actually
work—that have (so far as I know) almost nothing to do with things you’ve
been talking about.

I do agree with your belief that the universe is ultimately discrete. But
of course many people have for a long time said that they thought the
universe might at some level be discrete. Some of those people (like
Wheeler, Penrose, Finkelstein, etc.) are sophisticated physicists, and what
they’ve said has lots of real content—it’s not just vague essay-type
stuff. Now, I don’t happen to think what they’ve specifically proposed is
correct. But you would be completely wrong to think (as you seem to) that
somehow the idea that the universe might be discrete originated with you.

I really encourage you to read NKS in detail, including the notes at the
back. I think there’s a lot more there than you imagine. And I think if
you really understood it, you would be completely embarrassed to write a
note like the one below.

You’ve never struck me as being someone who is terribly interested in other
peoples’ ideas. And that’s of course fine. But you shouldn’t assume you
know their ideas just on the basis of a few buzzphrases or some such. In
some areas of business, that approach often works. Because, as we both
know, the ideas typically aren’t that deep. But it won’t work with a
character like me doing science. There’s too much nontrivial content. You
have to actually dig in to understand it. And from the things you say you
obviously haven’t.

For twenty years I thought we had a fine personal relationship. I thought
it was a little odd that you seemed to go around telling people that you had
introduced me to cellular automata. We talked about this a few times, and
you admitted this wasn’t a true story. But while I thought it was a little
unreasonable for you to keep on saying something you knew wasn’t true, I
didn’t pay much attention. It never really got in the way of our
relationship.

And then there was the incident of your NSF-funded conference. You invited
me. I said I couldn’t come. And suggested two alternates. You said fine.
But then you never contacted these people. Which was rather embarrassing
for me. And then, when David Reiss contacted you, you told him the
conference “was full”.

Later, when we talked about it, you admitted that that was a lie—and then
blamed the lie on Raj Reddy.

Frankly, I was flabbergasted by all this. That’s not the kind of
interaction someone like me expects to have with a seasoned high-level
operative like yourself. Yes, that’s the kind of thing some sleazy young
businessperson might do. But not a mature businessperson who has run
companies and things.

I still have no idea what you were thinking of. But it thoroughly shook my
confidence in you as someone I could interact straightforwardly with.

And then, of course, there’s the question of what you’ve said to journalists
etc. about NKS. In detail, I don’t have much idea. But something fishy
was surely going on. I haven’t gone and studied all the quotes from you.
But certainly my impression was that you were trying to claim that really
lots of key things in NKS were things you had done or said first.

You know that I tried to research the history carefully. And unless I
missed something quite huge, your contributions to NKS were extremely minor,
and are certainly accurately represented in the history notes. Now of
course if you don’t actually understand what I’ve done in NKS, that may be
hard to see. But I can’t really help you with that.

OK, where do we go from here?

We talked at some length when that reporter from a San Francisco paper was
trying to write a story about you and NKS. I thought we had a decent
conversation. But then, so far as I could tell, you went right ahead and
told the reporter—again—exactly a bunch of things we’d agreed in our
conversation weren’t true.

It was the same pattern as with telling people that you’d introduced me to
cellular automata. And it resonated in a bad way with the lie you told
about your conference.

I would have expected vastly better from you. I must say that I was
personally most disappointed. And I concluded with much regret that I must
have seriously misjudged you all these years.

I would like nothing more than to be able to mend our relationship, and go
back to the kind of pleasant social interactions we have always had.

How can that be achieved? Perhaps it’s impossible. But one step is that
you might actually try to understand what I’ve done in NKS. That would
surely help.

— Stephen

Subject: Delayed reply
Date: Thu, 3 Apr 2003 23:34:17 -0500
From: Edward Fredkin
To: Stephen Wolfram

Stephen,

I have been traveling and more recently have had my time gobbled up by
a most urgent matter.

I appreciate your quick reply to my email and I will get back to you
sometime soon. Rather than trying to respond to everything you brought up,
I will be limited to dealing with a couple of issues at a time.

What I can tell you is that I am not angry, and was not angry or upset.
I have always been a non-emotional observer with regard to whatever it is
that comes my way. That’s just my nature. It has come in handy at times,
such as when someone’s stupid mistake caused the single engine jet fighter
I was flying have an engine fire on take off. This required shutting down the
engine and taking other drastic actions very quickly; no time to get mad.

The gist of my comments to you was not related to the work you
documented in NKS, but rather to the style and methodology you are using
while trying to get people to understand and appreciate what it is you
have done. I certainly agree with the fact that it is extraordinarily
difficult to get the scientific establishment to pay attention, listen,
understand and appreciate what you’ve done. Nevertheless, I think there
might be a better approach to that problem than the one you are following.

So, as soon as I can get a little breathing room, I’ll respond to some
of your comments. I do value our friendship and whatever I do in this regard
will be an attempt at honest and unemotional communication with the goal of
some better mutual understanding.

By the way, I have taken the time to read and understand what you’ve
done in NKS. I’m pretty sure that I am better able than most to appreciate
the effort, persistence and creativity that went into that work.

You have made some comments about me and my own work, and I wonder what
you actually know about it beyond our conversations and the things you
referenced in NKS.

As soon as I can get some time I’ll continue with some further thoughts.

Best regards,

Ed

Ed didn’t send the promised followup. But a couple of months later New Scientist sent our media email address a note titled: “cover feature on Fredkin, Wolfram right to reply”, which asked for “comment on the suggestion that you first became familiar with cellular automata first at Fredkin’s lab in the 1970s and that examples in A New Kind of Science came out of work done in the lab”. I told Ed he should correct that—and he responded to me:

Subject: Re: cover feature on Fredkin, Wolfram right to reply.
Date: Thu, 29 May 2003 13:55:02 -0400
From: Ed Fredkin
To: Stephen Wolfram

Stephen,

I carefully and clearly told the author of the NS article that to my
knowledge it is not true “… that Wolfram first became familiar with
cellular automata at Fredkin’s lab in the ’70’s…” and further that
you already knew about CA’s.

My guess is that magazines see value in controversy and they would like
to attribute statements to each of us that helps them titillate their
readers. I tried in every way I could to correct any wrong impressions
the author had. But what they end up doing is beyond my control.

As to cracking the fundamental theory of physics, I did read and did
understand what you wrote about in NKS, however my interests lie in
models that are regular and based on a simple underlying Cartesian
lattice. The models I have been working on for the past few years are
called “Salt” as they are CA’s similar to an NaCl crystal. You can
read about it at www.digitalphilosophy.org.

My approach to being consistent with QM, SR and GR is related to the
fact that CA models of physics can exactly conserve such quantities as
momentum, energy, charge etc. By means of a variant of Noether’s
Theorem, the physics of such CA’s can exhibited the all the symmetries
we currently attribute to physics, but doing so asymptotically at
scales above the lattice.

Thus, in my concept of a theory of physics, translation symmetry,
rotation symmetry etc. would all be violated as we currently understand
is true for time symmetry, parity symmetry and charge symmetry .

No one suggests that you should agree with all my ideas, however your
comment in your prior email to me is unnecessarily condescending:

> “I think it’s a great pity that you’ve never taken
> the time to learn the technical details of physics as it’s
> currently practiced. There’s a lot known. And if you understood
> it, I think you’d be able to tell quite quickly which of
> your ideas are totally naive, and which might actually be interesting.”

What is certain is that there’s no “great pity” necessary. I actually
do know a lot about the technical details of physics. In any case,
thirty years ago Feynman thought that I needed to learn more about
certain aspects of QM. He was specific in what he felt was everything
more that I needed to know (in order to make progress with my CA
ideas). He offered to work with me, which was accomplished during the
course of the year I spent at Caltech (1974-1975). I studied, learned
more about QM and passed the final exam that Feynman gave me. While
we argued a lot, Feynman never accused me of having naive ideas.

As to NKS 2003, it doesn’t make a lot of sense for me to come to be a
member of the audience. If you would like me to participate in some
meaningful way, let me know.

Best regards,

Ed F

And after that exchange, Ed and I basically went back to being as we had been before—having pleasant interactions, without any particular scientific engagement. And in a sense for many years I kept out of Ed’s scientific way—not seriously working on physics again until 2019.

Since 2002 I’d been living in the Boston area, so Ed and I ran into each other more often. And although Ed’s behavior over A New Kind of Science had disappointed and upset me, it gave me a better understanding of Ed as a human being, and a vulnerable one at that.

The Later Ed

It was always a little hard to tell just what was going on with Ed. In July 2003, for example, he wrote to me:

Subject: Gunkel
Date: Thu, 24 Jul 2003 19:07:56 -0400
From: Ed Fredkin
To: Stephen Wolfram

Stephen,

First I must apologize for this long letter. Pat Gunkel sent me an
email telling of your visit. It prompted me (who hardly ever writes
anything) to type up my thoughts for whatever they’re worth.

You might be surprised at the number of wise and intelligent people who
really appreciate Pat and his works. Yet after more than 30 years of
fitful, diverse yet nearly continuous support, Pat has come to a
situation that, to him, looks like the end of the line.

I, unfortunately, am no longer in a position to personally provide the
kind of modest support that Pat needs to continue his church-mouse kind
of existence.

There is no doubt that Pat can be a difficult person to help, but I
notice that he has mellowed with age. Of course, Wolfram Research is
not a charitable institution. But I believe that Pat’s ideas on
ideonomy are really important and that those ideas may form the basis
of interesting future applications. The point of all this is that if
what Pat is doing seems interesting to you, some arrangement with
Wolfram Research might make sense.

(True to form, Gunkel followed up with a very forthright note, including a scathing critique he’d written of A New Kind of Science—as well as of Ed’s theories. That wouldn’t have deterred me, but I couldn’t see anything Gunkel could actually do for us, so I never pursued this.)

But did Ed’s note imply that Ed was running out of money? I’d always assumed some kind of vast business empire lurking in the background, but now I wasn’t sure.

I saw Ed only a few times in the next couple of years—at events like a Festschrift for Sulak and a bat mitzvah for one of Feynman’s granddaughters. But as usual, he was eager to tell stories, some of which I hadn’t heard before—mostly about things far in the past. He said that in the early 1960s John Cocke had stolen the idea of RISC architecture from his murdered friend Ben Gurley, though it had taken him two decades to get it taken seriously. He said that around the same time he’d been pulled in by the Air Force to help with analysis of blast waves from nuclear tests (and that story came with descriptions of B-52s doing loop-the-loop maneuvers when they dropped atomic bombs). He said that he’d once demoed the Muse music system (which, he emphasized, he, not Minsky, had invented) to an astonished audience in the Soviet Union. He said that he’d advised Richard Branson on his transatlantic balloon trip, telling him his butane burners weren’t correctly mounted—and in fact they fell off. And so on.

In 2005 Ed told me he’d been working with a programmer in California named Dan Miller (who’d developed audio compression software [and been at the NKS 2003 conference that Ed had been so upset about]) on the new 3D cellular automaton he’d invented that he called the “SALT architecture” because its pattern of updates were like the Na and Cl in a salt crystal.

But then in 2008 Ed told me he’d sold his island—presumably relieving whatever financial issues he’d had before—and suddenly Ed started to show up much more. He told me (as he did quite a few times) that he was working on a book (which never materialized). He told me he was teaching a course at Carnegie Mellon on the “Physics of Theoretical Computation”—which was apparently actually a very-much-as-before “engineering-style” effort to explore building features of physics from a cellular automaton, now with his SALT architecture. He invited me to a dinner at his house in honor of ‘t Hooft, photographed here with Ed, me and Sulak:

Click to enlarge

That fall, Ed came to the Midwest NKS Conference in Indiana, here photographed in a discussion with Greg Chaitin, me and others:

Click to enlarge

I would interact with Ed quite regularly after that—most often with him telling me about his use of Mathematica and soon Wolfram|Alpha. In 2012 Ed—now aged 78—sent me a nice “I have an idea” email (I made the requested introduction, though I’m not sure if this ever went anywhere):

Subject: Alpha and Problem Solving
Date: Fri, 19 Oct 2012 20:12:01 +0000
From: Edward Fredkin
To: Steve Wolfram

Steve,

The first thing I taught at MIT was a course in general problem solving (in 1968).
I’m now developing a new course on General Problem Solving which I expect to
offer first at Harvard’s HILR program. Part of the motivation came from watching
Joyce struggle with a Harvard course on Chemistry, where a lot of the homework
involved units conversions. I noticed that Alpha promptly solved many of Joyce’s
homework problems including some involving chemical reactions. (The course
was really for students planning to take the MCAT Exam in order to get into Medical
School). One clue that you might give to the Alpha developers, is to work toward
getting Alpha to have more of the capabilities necessary to pass different standard
tests that involve various kinds of quantitative analysis. (Of course, you might
have already done so.)

You might recall that I discussed the issue of units conversion with you long ago
(before Mathematica), and you described the idea you then had that turned into
Convert in Mathematica.

In any case, Alpha is fantastic, and getting better all the time. My plan is that every
one of my students must use Alpha for every problem that involves numbers, along
with some that don’t involve numbers. My motto is John McCarthy’s dictum:
“Those who refuse to do arithmetic are doomed to talk nonsense!” However, with
Alpha, the problem solver doesn’t have to do the arithmetic or the units
conversions; Alpha can do it!

It would be helpful if I could get a little bit of cooperation from someone in the Alpha
group. Basically, I will want to talk to an Alpha expert from time to time to make sure
I’m taking advantage of the best that Alpha can do along with resources already
developed for introducing Alpha to new users. My initial students will be drawn from
a group of retirees who, while clearly above average in intelligence, may have few
recently used skills in mathematics. I also expect that almost all of my initial
students will be first time Alpha users. Again, I might profit from discussion
with someone who has thought about how to introduce Alpha to beginners.

Let me know what you think or, if you like, we could get together to talk about it.

Best regards and Congratulations!

Ed

In 2014, when I recorded some oral history with Ed—now age 80—he was again brimming with ideas. The one he was most excited about had to do with weather prediction. It started from the observation that most smartphones have pressure sensors in them. Ed’s idea was to use these—and more—to create a sensor net that would continuously collect billions of pressure measurements, to be fed as input to weather forecast codes. Channeling his lifelong interest in reversible computing he imagined that the codes could be made reversible, and that running backwards from an incorrect prediction could tell one where more data had to be collected. Then Ed imagined doing this by having tiny balloons all over the place—with nothing that would cause trouble if a plane ran into it. He had a whole plan for partners he wanted to get (and, yes, he wanted us to be part of it too). And in typical Ed fashion, it was all laced with stories:

You know, I had this personal experience with weather. I was flying a glider along at 16,000 feet, and I encountered sink. You know, sink is wind blowing down. And the speed of the sink was 10,000 feet a minute. I was at 16,000 feet. And two minutes later, I was on the ground landing. Not on purpose. You know my attitude was—if I don’t see a big grading on the ground—[the wind] can’t keep going this way all the way down, so I won’t be killed. Actually, in that same storm, one of the pilots was killed.

The weather people just aren’t into the vertical movement of air. They do everything in layers. But this went through a lot of layers all at once in an organized fashion. So the point is that to talk about thousands or even millions of sensors makes no sense. You’re not going to do good weather until you get billions of sensors. That’s my opinion.

We talked about whether sensitive dependence on initial conditions destroys all predictability in fluid dynamics. I have theoretical and computational reasons to think it doesn’t. But Ed had a story:

There’s a mountain in California I happen to know, and I have a picture of a cloud street that starts on that mountain because it has a very peculiar geometry, and then runs for 2,000 miles.

So this particular mountain has an area of its rock that faces towards the east and it’s big. And what happens is when the Sun is shining on that and the wet wind is coming from the Pacific and so on, you get this big cumulus cloud that flows back this way, and then you get another one and it pulses. You get one after another. And these are very stable things and they travel a very long way. So my point is that amidst all the randomness there’s a lot of order that can be found and understood. There are regions that have funny properties. They’re much more temperature stable. There’s like islands of stability. And things like that get ignored by everything people are doing today, you know what I mean?

I would send things I’d written to Ed. I didn’t really think he’d read them. But I thought he might at least enjoy their concepts. And often he would respond with ideas of his own. I sent him an announcement about our Tweet-a-Program project (now reconfigured because of Twitter changes) with the one-line comment (reflecting his “best programmer” self-characterization): “A new frontier of programming prowess?” He responded, in typical Ed fashion, with an idea—that’s actually a little reminiscent of modern AI image generation:

Subject: Re: Tweet-a-Program
Date: Fri, 19 Sep 2014 21:25:47 +0000
From: Edward Fredkin
To: Stephen Wolfram

Hi,

I like it! As usual, it gave me ideas that might be outside of your
current concept.

We should talk sometime, so that I can explain something closely related
to [Tweet-a-Program] but decidedly different and perhaps even more fun.
Strangely, it has to do with Haiku.

What I have figured out is that there could be a new kind of Haiku, where
the text is interpreted by Mathematica to generate an image.

The trick will be having the image reflect something of the Haiku meaning,
even if only abstractly. I don’t know how to do this so that it does the perfect
thing every time, but I have thought of something that could be fun, and a
person could become skilled at creating Mathematica Haikus that seem to reflect
some aspects of the feeling of the words in an image with some increasing
probability of doing it well, as a result of practice.

Ed

Late in 2014 Ed sent me another piece of mail saying he was starting a project to produce a “new cellular automaton system”—and he wanted to use our technology to do it. He also sent me a paper he’d written about his SALT cellular automaton:

Click to enlarge

Finally—and without my help—Ed seemed to have mastered the art of academic papers. This one was on the arXiv preprint server. Others—with titles like “An Introduction to Digital Philosophy”—had appeared in academic journals. (Ones with titles like “A New Cosmogony” and “Finite Nature ” were more privately circulated.) But what most struck me about this particular paper was that—for the first time—it seemed to have actual images of cellular automaton behavior. Ever since those few minutes with the PERQ computer on Ed’s island in 1982 I hadn’t seen Ed ever show anything like that. And now Ed was again chasing that old question Minsky had asked, of making a circle with a cellular automaton.

At the time, I didn’t have a chance to see what Ed had actually done, and whether he’d finally solved it. But in writing this piece, I decided I’d better try to find out. The actual rule—that Ed and Dan Miller called “BusyBoxes”—is quite complicated, involving knight’s-move neighborhoods, etc. Their claim was that starting with a string of cells in a particular configuration, the average of their positions would trace out what in the limit of a long string would be a circle:

At first it looks like a kind of magic trick (and no, nothing is bouncing off any “walls”; the direction changes are just a consequence of the initial pattern of cells). But if you keep all the locations that get visited, things start to seem less mysterious—because what you realize is that the “basket” that gets “woven” is actually just a cube, viewed from a corner:

Where does the apparent circle come from? The details are a bit complicated—and I’ve put them in an appendix below. But suffice it say to that Ed’s old nemesis—calculus—comes in very handy. And in fact it lets one show that although one gets almost a circle, it’s not quite a circle; even with an infinite string, its radius is still wiggling by about 0.5% as one goes around the “circle”:

And—as we’ll see below—remarkably enough one can get a closed-form result for the amount of wiggliness (here computed as the ratio of maximum to minimum radius):

In earlier years, Ed might have tried to say that generating a circle (which this doesn’t) was tantamount to showing that a cellular automaton could reproduce physics. But by now I think he realized that it was really much more complicated than that. And he wasn’t mentioning physics much to me anymore. But—perhaps not least because many of his longtime interlocutors had by then died—he was interacting with me more than before. And perhaps he was even beginning to think that I might have a bit more to contribute than he’d assumed.

In December 2015 I sent Ed a piece I’d written to celebrate the bicentenary of Ada Lovelace, and he responded:

Date: Fri, 11 Dec 2015 15:58:14 +0000
From: Edward Fredkin
To: Stephen Wolfram

Stephen,

I was truly blown away by your essay re Ada Lovelace! You’ve got a lot
more to give the world than I had imagined, and I, more than anyone else,
appreciate what you might still be capable of accomplishing.

It’s too bad that some persons at MIT, for far too long, hung onto one
dimensional views focussed on what Macsyma might have been. My own
impressions have always been different, I recognized your potential long
ago and consequently invited you to one of my Mosqito Island conferences
some 3.5 decades ago.

In any case, much of what Mathematica makes possible is very important and
valuable to me. As you know I was an early user and continue to be a
user.

Many of my interests have run along many paths opened up by activities you
have instigated at Wolfram. Wolfram Alpha and its connections to Siri,
are examples.

Your new book “An Elementary Introduction to the Wolfram Language” (I
don’t yet have a hard copy) fits in with a project I had in mind for my
grandchild Robert, who at age 6 already seems to be extraordinarily
talented mathematically.

To cut to the chase, I want to make a proposal: Although I’m too old to
be a regular employee, I’d nevertheless like to have an association with
Wolfram, where I might be able to contribute ideas, and solve problems
(I’m still quite good at that).

I won’t need much from you other than your opening the door to my
involvement at Wolfram. What I have in mind would be an arrangement
where I could work for Wolfram, with some kind of arrangement other than
full time employment.

I’ve attached something I wrote recently.

Ed

Gosh! That was an unexpected development. Flattering, I suppose. But my main reaction was a kind of sadness. Yes, after all these years, Ed had finally read something I’d written. But somehow his response sounded like he was surrendering. This wasn’t the “I-want-to-do-everything-for-myself” Ed I had known all this time. This was an Ed who somehow felt he needed us to support him. And while our company has been able to absorb a great many “unusual” people—with terrific success—Ed seemed like he was pretty far outside our envelope.

At the time, I didn’t look at the attachment Ed sent with his email. But opening it now adds to my sense of sadness. It was a 13-page document about a system Ed imagined that would help people with “various forms of cognitive disabilities”, including a section on “Dementia and Alzheimer’s”:

Click to enlarge

It wasn’t until 2017 that Ed explicitly mentioned to me that his short-term memory was failing—though in talking to him it had been increasingly obvious for several years. He said he’d joined a group of people who were writing their memoirs. I told him I’d look forward to seeing his, though I’m not sure he ever made much progress on them.

Ed continued to send me ideas and proposals. There was a very Ed-like “global idea” about creating a system “GM” (presumably for “General Mathematician”) that would effectively “learn all of mathematics” by automatically reading math books, etc. (yes, definite overtones of what’s happening with LLM-meets-Wolfram-Language):

Click to enlarge

Later there were several pieces of mail about a new idea for factoring integers. In the first of them (from 2016), Ed told me that when the NeXT computer first came out (in 1989) he’d used Mathematica on it to simulate a reversible hardware multiplier. And being reminded of this by a historical piece I’d written, he said it had “started me thinking, again, about that problem and I had a new insight that appears to so greatly reduce the complexity of a reversible multiplier so as to possibly make it better at factoring large integers than current algorithms.” He wrote me about this several more times, suggesting various kinds of collaborations. Finally, in 2018 he told me how the method worked, saying it involved doing reversible arithmetic using balanced ternary. (Strangely enough, years earlier Ed had told me about Soviet computers that also used balanced ternary.)

I think that was the last technical conversation I had with Ed. A couple of years later I sent him the book about our Physics Project with the inscription:

Click to enlarge

And I would see him at least once a year at the Boston-area physics get-together organized by Boston University. He would always tell me stories. Often the same stories, and sometimes stories about me. And indeed as I was writing this piece I actually found a video Ed made in November 2020 that has such a story, albeit by this point seriously muddled (and, no, I’ve basically never “run” a cellular automaton by hand in my life!):

I used to organize meetings in the Caribbean and I did this because I had an island in the Caribbean … I invited Wolfram to come down. Wolfram had done pioneering work in cellular automata. … He was a great guy, you know, and I wanted him to get on the bandwagon … He shows up at the meeting and he had done all his work by hand as had everyone else in cellular automata. He didn’t think of using a computer. [!] I had a display processor that I modified to be able to run a cellular automaton with the stuff that it used to put text up on the screen. And so I’m showing him a cellular automata running at 60 frames a second continuously like a movie. This was 10,000 times faster than doing it by hand which is what he’d always done. He never thought of using a computer to do cellular automata and he turns around and walks out and and he left the island and went back to someplace else. So [later] I went to his meeting at Los Alamos and I ran into him again and he was now doing computer work. And I said to him “How come in all your work you don’t have a reversible [rule]”, and he says to me “Oh, reversible ones are all trivial”. And I went up and this is the most telling thing about his intellect: he’s a very smart guy [and when I] showed him how he could change his rule slightly and make it reversible his eyes just about popped out of his head and he knew I was correct.

I may have introduced him to this field but what he has done is he is far better than I at getting other people involved. I’ve never bothered and I don’t have the talent that he has for that. What he did was he came up with similar ideas and initially he didn’t give me the credit I thought I deserved. But it became apparent to me that he did this independently and he’s better at writing things and better at hiring bright people who can do things than I ever was.

And right after that, Ed ends the video with:

As I look back on my career I’ve had a fantastic life and I’m not unhappy about any aspect of it because, you know, I’ve accomplished everything I might have done and in spite of various handicaps—like not being a writer—I still have done a lot and the world, uh, understands me, I think, and appreciates what I’ve done.

When I saw Ed in 2022 he wasn’t able to say much. But, though it was a struggle, he was keen to make one point to me, that seemed to matter a lot to him: “You’ve managed to get people to follow you”, he said “I was never able to do that”. I saw Ed one last time this May. Joyce explained that Ed had “bumped his head”, and, in a very Ed-like way, she was avoiding a repeat by getting him to wear a bike helmet. She wanted someone to snap a picture of me and her with Ed:

Click to enlarge

Six weeks later, Ed died, at the age of 88.

I went to see Joyce and Rick a few weeks later, among other things to check facts for this piece. I’d heard from Ed that his ancestors had provided wood for the imperial palace in St. Petersburg. But I’d also heard from someone else that Ed had said he was descended from Mongolian royalty. And as I was about to leave, I thought I might as well ask. “Oh yes”, they said. “And Ed’s father even wrote a historical novel about it”. And they showed me two books (both from the mid-1980s):

Click to enlarge

I’m not sure who Sarah, Queen of Mongolia was, but the book blurb claims that Ed’s father was her great-great-great-grandson—and goes on to speak of the “strong family inheritance of a mind that analyzes not only the injustice of human oppression but offers realistic and beneficial solutions”.

Summing Up Ed

“Can that really be true?” I often asked myself when hearing yet another of Ed’s implausible stories. And of course it didn’t help that stories he told—even to me—about me weren’t true. But the remarkable thing in writing this piece is that I’ve been able to verify that a lot of Ed’s stories—implausible though they may have sounded—were in fact true. Yes, they were often embellished, and parts that didn’t reflect so well on Ed were omitted. But together they defined a remarkable tapestry of a life.

It was in many ways a very independent life. Ed had friends and family members to whom he stayed close throughout his life. But mostly it was “Ed for himself, against the world”. He didn’t want to learn anything from anyone else; he wanted to figure out everything for himself. He wanted to invent his own ideas; he wasn’t too interested in other people’s. In a rather Air-Force-pilot kind of way (“eject or not?”) he liked to be decisive—and he liked to be incisive too, always figuring out a clear, simple thing to say. Sometimes that came across as naive. And sometimes it was in fact naive. But mostly Ed didn’t seem to mind much; he would just go on to another idea.

Ed was a great storyteller, and an engaging speaker. For some reason he developed the theory that he couldn’t write—but there’s ample evidence, going back even to his teenage years, that this wasn’t true. If there was a problem, it was with content, not writing. And the issue with the content was that it tended to just be too Ed-specific—too insular—and not connected enough for other people to be able to understand or appreciate it.

I don’t know what Ed was like as a manager; I rather suspect he may have suffered from trying to be a bit too clever, with too many ideas and too much gamification. In the end, he felt he’d failed as a leader, and perhaps that was inevitable given how independent he always wanted to be. Despite his stints as an academic administrator and as a CEO, Ed was in the end fundamentally a lone warrior (and problem solver), not a general.

And what about all those ideas? Most never developed very far. Some were pretty wild. But many had at least a kernel of visionary insight. The details of the universe as a cellular automaton didn’t make sense. But the idea that the universe is somehow computational is surely correct. And spread over the course of more than six decades, Ed spun out nuggets of ideas that would later appear—usually much more developed—in a remarkable range of areas.

Ed projected a kind of personal serenity—yet he was in many ways deeply competitive. Most of the time, though, he was able to define the arena of his competitiveness so idiosyncratically that there really weren’t other contenders. And I think in the end Ed felt pretty good about all the things he’d managed to do in his life. It was fitting that he owned an actual island. Because somehow an island was a metaphor for Ed’s life: separate, independent and unique.

Thanks

I’ve had help with information for this piece from many people, including Joyce Fredkin, Rick Fredkin, Simson Garfinkel, Andrea Gerlach, Bill Gosper, Howard Gutowitz, Steven Levy, Norm Margolus, Margaret Minsky, Dave Moon, John Moussouris, Mark Nahabedian, Walter Parkes, David Reiss, Brian Silverman, George Sulak, Larry Sulak and Matthew Szudzik. (Tom Toffoli agreed to talk, but didn’t show up.) I thank the Department of Distinctive Collections at the MIT Library for access to the Fredkin papers archive there. Thanks also to Brad Klee and Nik Murzin for technical help.

Appendix: Analyzing the Not-Quite-Circle

Here’s what the SALT cellular automaton does for two sizes of initial “string”:

For an initial string of length n (with n > 2), the overall period is 54n – 43, and the envelope “woven” going through all configurations is:

The “circle” is obtained by averaging the positions of all cells present at a given time step. The “circle” is always planar, but its effective radius varies with direction (i.e. as the system steps through each cycle):

Ed and Dan Miller looked at the standard deviation of the effective radius as a function of n, computing it up to n = 20, and getting the following results:

Click to enlarge

It looked as if the standard deviation was just going to go smoothly to zero—so that for an infinite string one would get a perfect circle. But that turns out not to be true, as one can see by extending the computation to slightly larger values of n:

And actually there’s a minimum at n = 43, with standard deviation 0.0012 (and fractional size discrepancy 0.0048)—and it doesn’t look like even for n ∞ one will get a perfect circle.

But how can one work out the n ∞ case? It’s actually a nice application for calculus.

First, notice that the “basket” consists of a series of layers of a cube viewed from one of its corners, or in other words a sequence of shapes like this:

Here’s how these are formed as one sweeps through the cube:

One can think of the string in the cellular automaton as spanning these “layers”, and successively moving around all of them as the cellular automaton evolves. In the continuum limit, there’s effectively a parameter t that defines where on each “layer curve” one is at a particular time. Conveniently enough, the length of all the layer curves is the same (for a unit cube it is 3 ≈ 4.2). With successive layers parametrized by a variable s (running from 0 to 1) the corners of the layer curves (all normalized to have length 1) are given by:

Now we need to find the actual x, y positions of string elements (AKA infinitesimal cells) as a function of s and t. Since the edges of the layer polygons are always straight, in each of a series of “piecewise regions” in s and t (with breakpoints defined by the corners of the polygons), we get expressions for x and y that are linear in s and t:

One subtlety is that the string in essence turns as time progresses, so that it effectively samples a different t value for different layers s. To correct for this, we have to find for which t we get x = 0 for a given s. It’s convenient to put the center of all our layer curves at {0, 0}, and we can do this now by subtracting . Then the (first) value of t at which x = 0 is given simply by:

The parametric surface we now get as a function of t is (with discrete lines indicating particular values of s):

Now we can slice the parametric surface not in discrete s values but instead in discrete t values—thus getting what’s basically a sequence of effective strings at discrete times:

The centroids of the strings are indicated in green, and these are then points on our potential circle. Using what we did above, the radius of this “circle” as a function of t can then be found by integrating over s. The result is algebraically complicated, but has a closed form:

Integrating this over t we get the “average radius”, normalized to “circumference 1” from the fact that t varies from 0 to 1 going “around the circle”:

(This means that the “effective π” for this circle is about 3.437.)

Now we can plot the “wiggle” of the radius as a function of “angle” (i.e. t):

It looks a bit like a sine curve, but it’s not one. And, for example, it isn’t even symmetrical. Its maxima (which occur at odd multiples of 30°) are

while its minima (at even multiples of 30°) are

and dividing by the average radius these are about 1.00734 and 0.992175.

The ratio of maximum to minimum (effectively “wiggle amplitude”) is:

Meanwhile, the standard deviation can be obtained as an integral over t, and the final result is

which is about 2.4 times larger than what we get at n = 100. We can see the approach to the asymptotic value by computing integrals over t for progressively larger numbers of discrete values of s (which, we should emphasize, is similar to values of n, but not quite the same, particularly for small n):

Will AIs Take All Our Jobs and End Human History—or Not? Well, It’s Complicated…

16 mars 2023 à 02:41

The Shock of ChatGPT

Just a few months ago writing an original essay seemed like something only a human could do. But then ChatGPT burst onto the scene. And suddenly we realized that an AI could write a passable human-like essay. So now it’s natural to wonder: How far will this go? What will AIs be able to do? And how will we humans fit in?

My goal here is to explore some of the science, technology—and philosophy—of what we can expect from AIs. I should say at the outset that this is a subject fraught with both intellectual and practical difficulty. And all I’ll be able to do here is give a snapshot of my current thinking—which will inevitably be incomplete—not least because, as I’ll discuss, trying to predict how history in an area like this will unfold is something that runs straight into an issue of basic science: the phenomenon of computational irreducibility.

But let’s start off by talking about that particularly dramatic example of AI that’s just arrived on the scene: ChatGPT. So what is ChatGPT? Ultimately, it’s a computational system for generating text that’s been set up to follow the patterns defined by human-written text from billions of webpages, millions of books, etc. Give it a textual prompt and it’ll continue in a way that’s somehow typical of what it’s seen us humans write.

The results (which ultimately rely on all sorts of specific engineering) are remarkably “human like”. And what makes this work is that whenever ChatGPT has to “extrapolate” beyond anything it’s explicitly seen from us humans it does so in ways that seem similar to what we as humans might do.

Inside ChatGPT is something that’s actually computationally probably quite similar to a brain—with millions of simple elements (“neurons”) forming a “neural net” with billions of connections that have been “tweaked” through a progressive process of training until they successfully reproduce the patterns of human-written text seen on all those webpages, etc. Even without training the neural net would still produce some kind of text. But the key point is that it won’t be text that we humans consider meaningful. To get such text we need to build on all that “human context” defined by the webpages and other materials we humans have written. The “raw computational system” will just do “raw computation”; to get something aligned with us humans requires leveraging the detailed human history captured by all those pages on the web, etc.

But so what do we get in the end? Well, it’s text that basically reads like it was written by a human. In the past we might have thought that human language was somehow a uniquely human thing to produce. But now we’ve got an AI doing it. So what’s left for us humans? Well, somewhere things have got to get started: in the case of text, there’s got to be a prompt specified that tells the AI “what direction to go in”. And this is the kind of thing we’ll see over and over again. Given a defined “goal”, an AI can automatically work towards achieving it. But it ultimately takes something beyond the raw computational system of the AI to define what us humans would consider a meaningful goal. And that’s where we humans come in.

What does this mean at a practical, everyday level? Typically we use ChatGPT by telling it—using text—what we basically want. And then it’ll fill in a whole essay’s worth of text talking about it. We can think of this interaction as corresponding to a kind of “linguistic user interface” (that we might dub a “LUI”). In a graphical user interface (GUI) there’s core content that’s being rendered (and input) through some potentially elaborate graphical presentation. In the LUI provided by ChatGPT there’s instead core content that’s being rendered (and input) through a textual (“linguistic”) presentation.

You might jot down a few “bullet points”. And in their raw form someone else would probably have a hard time understanding them. But through the LUI provided by ChatGPT those bullet points can be turned into an “essay” that can be generally understood—because it’s based on the “shared context” defined by everything from the billions of webpages, etc. on which ChatGPT has been trained.

There’s something about this that might seem rather unnerving. In the past, if you saw a custom-written essay you’d reasonably be able to conclude that a certain irreducible human effort was spent in producing it. But with ChatGPT this is no longer true. Turning things into essays is now “free” and automated. “Essayification” is no longer evidence of human effort.

Of course, it’s hardly the first time there’s been a development like this. Back when I was a kid, for example, seeing that a document had been typeset was basically evidence that someone had gone to the considerable effort of printing it on printing press. But then came desktop publishing, and it became basically free to make any document be elaborately typeset.

And in a longer view, this kind of thing is basically a constant trend in history: what once took human effort eventually becomes automated and “free to do” through technology. There’s a direct analog of this in the realm of ideas: that with time higher and higher levels of abstraction are developed, that subsume what were formerly laborious details and specifics.

Will this end? Will we eventually have automated everything? Discovered everything? Invented everything? At some level, we now know that the answer is a resounding no. Because one of the consequences of the phenomenon of computational irreducibility is that there’ll always be more computations to do—that can’t in the end be reduced by any finite amount of automation, discovery or invention.

Ultimately, though, this will be a more subtle story. Because while there may always be more computations to do, it could still be that we as humans don’t care about them. And that somehow everything we care about can successfully be automated—say by AIs—leaving “nothing more for us to do”.

Untangling this issue will be at the heart of questions about how we fit into the AI future. And in what follows we’ll see over and over again that what might at first essentially seem like practical matters of technology quickly get enmeshed with deep questions of science and philosophy.

Intuition from the Computational Universe

I’ve already mentioned computational irreducibility a couple of times. And it turns out that this is part of a circle of rather deep—and at first surprising—ideas that I believe are crucial to thinking about the AI future.

Most of our existing intuition about “machinery” and “automation” comes from a kind of “clockwork” view of engineering—in which we specifically build systems component by component to achieve objectives we want. And it’s the same with most software: we write it line by line to specifically do—step by step—whatever it is we want. And we expect that if we want our machinery—or software—to do complex things then the underlying structure of the machinery or software must somehow be correspondingly complex.

So when I started exploring the whole computational universe of possible programs in the early 1980s it was a big surprise to discover that things work quite differently there. And indeed even tiny programs—that effectively just apply very simple rules repeatedly—can generate great complexity. In our usual practice of engineering we haven’t seen this, because we’ve always specifically picked programs (or other structures) where we can readily foresee how they’ll behave, so that we can explicitly set them up to do what we want. But out in the computational universe it’s very common to see programs that just “intrinsically generate” great complexity, without us ever having to explicitly “put it in”.

And having discovered this, we realize that there’s actually a big example that’s been around forever: the natural world. And indeed it increasingly seems as if the “secret” that nature uses to make the complexity it so often shows is exactly to operate according to the rules of simple programs. (For about three centuries it seemed as if mathematical equations were the ultimate way to describe the natural world—but in the past few decades, and particularly poignantly with our recent Physics Project, it’s become clear that simple programs are in general a more powerful approach.)

How does all this relate to technology? Well, technology is about taking what’s out there in the world, and harnessing it for human purposes. And there’s a fundamental tradeoff here. There may be some system out in nature that does amazingly complex things. But the question is whether we can “slice off” certain particular things that we humans happen to find useful. A donkey has all sorts of complex things going on inside. But at some point it was discovered that we can use it “technologically” to do the rather simple thing of pulling a cart.

And when it comes to programs out in the computational universe it’s extremely common to see ones that do amazingly complex things. But the question is whether we can find some aspect of those things that’s useful to us. Maybe the program is good at making pseudorandomness. Or distributedly determining consensus. Or maybe it’s just doing its complex thing, and we don’t yet know any “human purpose” that this achieves.

One of the notable features of a system like ChatGPT is that it isn’t constructed in an “understand-every-step” traditional engineering way. Instead one basically just starts from a “raw computational system” (in the case of ChatGPT, a neural net), then progressively tweaks it until its behavior aligns with the “human-relevant” examples one has. And this alignment is what makes the system “technologically useful”—to us humans.

Underneath, though, it’s still a computational system, with all the potential “wildness” that implies. And free from the “technological objective” of “human-relevant alignment” the system might do all sorts of sophisticated things. But they might not be things that (at least at this time in history) we care about. Even though some putative alien (or our future selves) might.

OK, but let’s come back to the “raw computation” side of things. There’s something very different about computation from all other kinds of “mechanisms” we’ve seen before. We might have a cart that can move forward. And we might have a stapler that can put staples in things. But carts and staplers do very different things; there’s no equivalence between them. But for computational systems (at least ones that don’t just always behave in obviously simple ways) there’s my Principle of Computational Equivalence—which implies that all these systems are in a sense equivalent in the kinds of computations they can do.

This equivalence has many consequences. One of them is that one can expect to make something equally computationally sophisticated out of all sorts of different kinds of things—whether brain tissue or electronics, or some system in nature. And this is effectively where computational irreducibility comes from.

One might think that given, say, some computational system based on a simple program it would always be possible for us—with our sophisticated brains, mathematics, computers, etc.—to “jump ahead” and figure out what the system will do before it’s gone through all the steps to do it. But the Principle of Computational Equivalence implies that this won’t in general be possible—because the system itself can be as computationally sophisticated as our brains, mathematics, computers, etc. are. So this means that the system will be computationally irreducible: the only way to find out what it does is effectively just to go through the same whole computational process that it does.

There’s a prevailing impression that science will always eventually be able do better than this: that it’ll be able to make “predictions” that allow us to work out what will happen without having to trace through each step. And indeed over the past three centuries there’s been lots of success in doing this, mainly by using mathematical equations. But ultimately it turns out that this has only been possible because science has ended up concentrating on particular systems where these methods work (and then these systems have been used for engineering). But the reality is that many systems show computational irreducibility. And in the phenomenon of computational irreducibility science is in effect “deriving its own limitedness”.

Contrary to traditional intuition, try as we might, in many systems we’ll never be able find “formulas” (or other “shortcuts”) that describe what’s going to happen in the systems—because the systems are simply computationally irreducible. And, yes, this represents a limitation on science, and on knowledge in general. But while at first this might seem like a bad thing, there’s also something fundamentally satisfying about it. Because if everything were computationally reducible, we could always “jump ahead” and find out what will happen in the end, say in our lives. But computational irreducibility implies that in general we can’t do that—so that in some sense “something irreducible is being achieved” by the passage of time.

There are a great many consequences of computational irreducibility. Some—that I have particularly explored recently—are in the domain of basic science (for example, establishing core laws of physics as we perceive them from the interplay of computational irreducibility and our computational limitations as observers). But computational irreducibility is also central in thinking about the AI future—and in fact I increasingly feel that it adds the single most important intellectual element needed to make sense of many of the most important questions about the potential roles of AIs and humans in the future.

For example, from our traditional experience with engineering we’re used to the idea that to find out why something happened in a particular way we can just “look inside” a machine or program and “see what it did”. But when there’s computational irreducibility, that won’t work. Yes, we could “look inside” and see, say, a few steps. But computational irreducibility implies that to find out what happened, we’d have to trace through all the steps. We can’t expect to find a “simple human narrative” that “says why something happened”.

But having said this, one feature of computational irreducibility is that within any computationally irreducible systems there must always be (ultimately, infinitely many) “pockets of computational reducibility” to be found. So for example, even though we can’t say in general what will happen, we’ll always be able to identify specific features that we can predict. (“The leftmost cell will always be black”, etc.) And as we’ll discuss later we can potentially think of technological (as well as scientific) progress as being intimately tied to the discovery of these “pockets of reducibility”. And in effect the existence of infinitely many such pockets is the reason that “there’ll always be inventions and discoveries to be made”.

Another consequence of computational irreducibility has to do with trying to ensure things about the behavior of a system. Let’s say one wants to set up an AI so it’ll “never do anything bad”. One might imagine that one could just come up with particular rules that ensure this. But as soon as the behavior of the system (or its environment) is computationally irreducible one will never be able to guarantee what will happen in the system. Yes, there may be particular computationally reducible features one can be sure about. But in general computational irreducibility implies that there’ll always be a “possibility of surprise” or the potential for “unintended consequences”. And the only way to systematically avoid this is to make the system not computationally irreducible—which means it can’t make use of the full power of computation.

“AIs Will Never Be Able to Do That”

We humans like to feel special, and feel as if there’s something “fundamentally unique” about us. Five centuries ago we thought we lived at the center of the universe. Now we just tend to think that there’s something about our intellectual capabilities that’s fundamentally unique and beyond anything else. But the progress of AI—and things like ChatGPT—keep on giving us more and more evidence that that’s not the case. And indeed my Principle of Computational Equivalence says something even more extreme: that at a fundamental computational level there’s just nothing fundamentally special about us at all—and that in fact we’re computationally just equivalent to lots of systems in nature, and even to simple programs.

This broad equivalence is important in being able to make very general scientific statements (like the existence of computational irreducibility). But it also highlights how significant our specifics—our particular history, biology, etc.—are. It’s very much like with ChatGPT. We can have a generic (untrained) neural net with the same structure as ChatGPT, that can do certain “raw computation”. But what makes ChatGPT interesting—at least to us—is that it’s been trained with the “human specifics” described on billions of webpages, etc. In other words, for both us and ChatGPT there’s nothing computationally “generally special”. But there is something “specifically special”—and it’s the particular history we’ve had, particular knowledge our civilization has accumulated, etc.

There’s a curious analogy here to our physical place in the universe. There’s a certain uniformity to the universe, which means there’s nothing “generally special” about our physical location. But at least to us there’s still something “specifically special” about it, because it’s only here that we have our particular planet, etc. At a deeper level, ideas based on our Physics Project have led to the concept of the ruliad: the unique object that is the entangled limit of all possible computational processes. And we can then view our whole experience as “observers of the universe” as consisting of sampling the ruliad at a particular place.

It’s a bit abstract (and a long story, which I won’t go into in any detail here), but we can think of different possible observers as being both at different places in physical space, and at different places in rulial space—giving them different “points of view” about what happens in the universe. Human minds are in effect concentrated in a particular region of physical space (mostly on this planet) and a particular region of rulial space. And in rulial space different human minds—with their different experiences and thus different ways of thinking about the universe—are in slightly different places. Animal minds might be fairly close in rulial space. But other computational systems (like, say, the weather, which is sometimes said to “have a mind of its own”) are further away—as putative aliens might also be.

So what about AIs? It depends what we mean by “AIs”. If we’re talking about computational systems that are set up to do “human-like things” then that means they’ll be close to us in rulial space. But insofar as “an AI” is an arbitrary computational system it can be anywhere in rulial space, and it can do anything that’s computationally possible—which is far broader than what we humans can do, or even think about. (As we’ll talk about later, as our intellectual paradigms—and ways of observing things—expand, the region of rulial space in which we humans operate will correspondingly expand.)

But, OK, just how “general” are the computations that we humans (and the AIs that follow us) are doing? We don’t know enough about the brain to be sure. But if we look at artificial neural net systems—like ChatGPT—we can potentially get some sense. And in fact the computations really don’t seem to be that “general”. In most neural net systems data that’s given as input just “ripples once through the system” to produce output. It’s not like in a computational system like a Turing machine where there can be arbitrary “recirculation of data”. And indeed without such “arbitrary recirculation” the computation is necessarily quite “shallow” and can’t ultimately show computational irreducibility.

It’s a bit of a technical point, but one can ask whether ChatGPT, with its “re-feeding of text produced so far” can in fact achieve arbitrary (“universal”) computation. And I suspect that in some formal sense it can (or at least a sufficiently expanded analog of it can)—though by producing an extremely verbose piece of text that for example in effect lists successive (self-delimiting) states of a Turing machine tape, and in which finding “the answer” to a computation will take a bit of effort. But—as I’ve discussed elsewhere—in practice ChatGPT is presumably almost exclusively doing “quite shallow” computation.

It’s an interesting feature of the history of practical computing that what one might consider “deep pure computations” (say in mathematics or science) were done for decades before “shallow human-like computations” became feasible. And the basic reason for this is that for “human-like computations” (like recognizing images or generating text) one needs to capture lots of “human context”, which requires having lots of “human-generated data” and the computational resources to store and process it.

And, by the way, brains also seem to specialize in fundamentally shallow computations. And to do the kind of deeper computations that allow one to take advantage of more of what’s out there in the computational universe, one has to turn to computers. As we’ve discussed, there’s plenty out in the computational universe that we humans don’t (yet) care about: we just consider it “raw computation”, that doesn’t seem to be “achieving human purposes”. But as a practical matter it’s important to make a bridge between the things we humans do care about and think about, and what’s possible in the computational universe. And in a sense that’s at the core of the project I’ve put so much effort into in the Wolfram Language of creating a full-scale computational language that describes in computational terms the things we think about, and experience in the world.

OK, people have been saying for years: “It’s nice that computers can do A and B, but only humans can do X”. What X is supposed to be has changed—and narrowed—over the years. And ChatGPT provides us with a major unexpected new example of something more that computers can do.

So what’s left? People might say: “Computers can never show creativity or originality”. But—perhaps disappointingly—that’s surprisingly easy to get, and indeed just a bit of randomness “seeding” a computation can often do a pretty good job, as we saw years ago with our WolframTones music-generation system, and as we see today with ChatGPT’s writing. People might also say: “Computers can never show emotions”. But before we had a good way to generate human language we wouldn’t really have been able to tell. And now it already works pretty well to ask ChatGPT to write “happily”, “sadly”, etc. (In their raw form emotions in both humans and other animals are presumably associated with rather simple “global variables” like neurotransmitter concentrations.)

In the past people might have said: “Computers can never show judgement”. But by now there are endless examples of machine learning systems that do well at reproducing human judgement in lots of domains. People might also say: “Computers don’t show common sense”. And by this they typically mean that in a particular situation a computer might locally give an answer, but there’s a global reason why that answer doesn’t make sense, that the computer “doesn’t notice”, but a person would.

So how does ChatGPT do on this? Not too badly. In plenty of cases it correctly recognizes that “that’s not what I’ve typically read”. But, yes, it makes mistakes. Some of them have to do with it not being able to do—purely with its neural net—even slightly “deeper”computations. (And, yes, that’s something that can often be fixed by it calling Wolfram|Alpha as a tool.) But in other cases the problem seems to be that it can’t quite connect different domains well enough.

It’s perfectly capable of doing simple (“SAT-style”) analogies. But when it comes to larger-scale ones it doesn’t manage them. My guess, though, is that it won’t take much scaling up before it starts to be able to make what seem like very impressive analogies (that most of us humans would never even be able to make)—at which point it’ll probably successfully show broader “common sense”.

But so what’s left that humans can do, and AIs can’t? There’s—almost by definition—one fundamental thing: define what we would consider goals for what to do. We’ll talk more about this later. But for now we can note that any computational system, once “set in motion”, will just follow its rules and do what it does. But what “direction should it be pointed in”? That’s something that has to come from “outside the system”.

So how does it work for us humans? Well, our goals are in effect defined by the whole web of history—both from biological evolution and from our cultural development—in which we are embedded. But ultimately the only way to truly participate in that web of history is to be part of it.

Of course, we can imagine technologically emulating every “relevant” aspect of a brain—and indeed things like the success of ChatGPT may suggest that that’s easier to do than we might have thought. But that won’t be enough. To participate in the “human web of history” (as we’ll discuss later) we’ll have to emulate other aspects of “being human”—like moving around, being mortal, etc. And, yes, if we make an “artificial human” we can expect it (by definition) to show all the features of us humans.

But while we’re still talking about AIs as—for example—“running on computers” or “being purely digital” then, at least as far as we’re concerned, they’ll have to “get their goals from outside”. One day (as we’ll discuss) there will no doubt be some kind of “civilization of AIs”—which will form its own web of history. But at this point there’s no reason to think that we’ll still be able to describe what’s going on in terms of goals that we recognize. In effect the AIs will at that point have left our domain of rulial space. And—as we’ll discuss—they’ll be operating more like the kind of systems we see in nature, where we can tell there’s computation going on, but we can’t describe it, except rather anthropomorphically, in terms of human goals and purposes.

Will There Be Anything Left for the Humans to Do?

It’s been an issue that’s been raised—with varying degrees of urgency—for centuries: with the advance of automation (and now AI), will there eventually be nothing left for humans to do? Back in the early days of our species, there was lots of hard work of hunting and gathering to do, just to survive. But at least in the developed parts of the world, that kind of work is now at best a distant historical memory.

And yet at each stage in history—at least so far—there always seem to be other kinds of work that keep people busy. But there’s a pattern that increasingly seems to repeat. Technology in some way or another enables some new occupation. And eventually that occupation becomes widespread, and lots of people do it. But then there’s a technological advance, and the occupation gets automated—and people aren’t needed to do it anymore. But now there’s a new level of technology, that enables new occupations. And the cycle continues.

A century ago the increasingly widespread use of telephones meant that more and more people worked as switchboard operators. But then telephone switching was automated—and those switchboard operators weren’t needed anymore. But with automated switching there could be huge development of telecommunications infrastructure, opening up all sorts of new types of jobs, that in aggregate employ vastly more people than were ever switchboard operators.

Something somewhat similar happened with accounting clerks. Before there were computers, one needed to have people laboriously tallying up numbers. But with computers, that was all automated away. But with that automation came the ability to do more complex financial computations—which allowed for more complex financial transactions, more complex regulations, etc., which in turn led to all sorts of new types of jobs.

And across a whole range of industries, it’s been the same kind of story. Automation obsoletes some jobs, but enables others. There’s quite often a gap in time, and a change in the skills that are needed. But at least so far there always seems to have been a broad frontier of jobs that have been made possible—but haven’t yet been automated.

Will this at some point end? Will there come a time when everything we humans want (or at least need) is delivered automatically? Well, of course, that depends on what we want, and whether, for example, that evolves with what technology has made possible. But could we just decide that “enough is enough”; let’s stop here, and just let everything be automated?

I don’t think so. And the reason is ultimately because of computational irreducibility. We try to get the world to be “just so”, say set up so we’re “predictably comfortable”. Well, the problem is that there’s inevitably computational irreducibility in the way things develop—not just in nature, but in things like societal dynamics too. And that means that things won’t stay “just so”. There’ll always be something unpredictable that happens; something that the automation doesn’t cover.

At first we humans might just say “we don’t care about that”. But in time computational irreducibility will affect everything. So if there’s anything at all we care about (including, for example, not going extinct), we’ll eventually have to do something—and go beyond whatever automation was already set up.

It’s easy to find practical examples. We might think that when computers and people are all connected in a seamless automated network, there’d be nothing more to do. But what about the “unintended consequence” of computer security issues? What might have seemed like a case where “technology finished things” quickly creates a new kind of job for people to do. And at some level, computational irreducibility implies that things like this must always happen. There must always be a “frontier”. At least if there’s anything at all we want to preserve (like not going extinct).

But let’s come back to the situation here and now with AI. ChatGPT just automated all sorts of text-related tasks. It used to take lots of effort—and people—to write customized reports, letters, etc. But (at least so long as one’s dealing with situations where one doesn’t need 100% “correctness”) ChatGPT just automated a lot of that, so people aren’t needed for it anymore. But what will this mean? Well, it means that there’ll be a lot more customized reports, letters, etc. that can be produced. And that will lead to new kinds of jobs—managing, analyzing, validating etc. all that mass-customized text. Not to mention the need for prompt engineers (a job category that just didn’t exist until a few months ago), and what amount to AI wranglers, AI psychologists, etc.

But let’s talk about today’s “frontier” of jobs that haven’t been “automated away”. There’s one category that in many ways seems surprising to still be “with us”: jobs that involve lots of mechanical manipulation, like construction, fulfillment, food preparation, etc. But there’s a missing piece of technology here: there isn’t yet good general-purpose robotics (as there is general-purpose computing), and we humans still have the edge in dexterity, mechanical adaptability, etc. But I’m quite sure that in time—and perhaps quite suddenly—the necessary technology will be developed (and, yes, I have ideas about how to do it). And this will mean that most of today’s “mechanical manipulation” jobs will be “automated away”—and won’t need people to do them.

But then, just as in our other examples, this will mean that mechanical manipulation will become much easier and cheaper to do, and more of it will be done. Houses might routinely be built and dismantled. Products might routinely be picked up from wherever they’ve ended up, and redistributed. Vastly more ornate “food constructions” might become the norm. And each of these things—and many more—will open up new jobs.

But will every job that exists in the world today “on the frontier” eventually be automated? What about jobs where it seems like a large part of the value is just “having a human be there”? Jobs like flying a plane where one wants the “commitment” of the pilot being there in the plane. Caregiver jobs where one wants the “connection” of a human being there. Sales or education jobs where one wants “human persuasion” or “human encouragement”. Today one might think “only a human can make one feel that way”. But that’s typically based on the way the job is done now. And maybe there’ll be different ways found that allow the essence of the task to be automated, almost inevitably opening up new tasks to be done.

For example, something that in the past needed “human persuasion” might be “automated” by something like gamification—but then more of it can be done, with new needs for design, analytics, management, etc.

We’ve been talking about “jobs”. And that term immediately brings to mind wages, economics, etc. And, yes, plenty of what people do (at least in the world as it is today) is driven by issues of economics. But plenty is also not. There are things we “just want to do”—as a “social matter”, for “entertainment”, for “personal satisfaction”, etc.

Why do we want to do these things? Some of it seems intrinsic to our biological nature. Some of it seems determined by the “cultural environment” in which we find ourselves. Why might one walk on a treadmill? In today’s world one might explain that it’s good for health, lifespan, etc. But a few centuries ago, without modern scientific understanding, and with a different view of the significance of life and death, that explanation really wouldn’t work.

What drives such changes in our view of what we “want to do”, or “should do”? Some seems to be driven by the pure “dynamics of society”, presumably with its own computational irreducibility. But some has to do with our ways of interacting with the world—both the increasing automation delivered by the advance of technology, and the increasing abstraction delivered by the advance of knowledge.

And there seem to be similar “cycles” seen here as in the kinds of things we consider to be “occupations” or “jobs”. For a while something is hard to do, and serves as a good “pastime”. But then it gets “too easy” (“everybody now knows how to win at game X”, etc.), and something at a “higher level” takes its place.

About our “base” biologically driven motivations it doesn’t seem like anything has really changed in the course of human history. But there are certainly technological developments that could have an effect in the future. Effective human immortality, for example, would change many aspects of our motivation structure. As would things like the ability to implant memories or, for that matter, implant motivations.

For now, there’s a certain element of what we want to do that’s “anchored” by our biological nature. But at some point we’ll surely be able to emulate with a computer at least the essence of what our brains are doing (and indeed the success of things like ChatGPT makes it seems like the moment when that will happen is closer at hand than we might have thought). And at that point we’ll have the possibility of what amount to “disembodied human souls”.

To us today it’s very hard to imagine what the “motivations” of such a “disembodied soul” might be. Looked at “from the outside” we might “see the soul” doing things that “don’t make much sense” to us. But it’s like asking what someone from a thousand years ago would think about many of our activities today. These activities make sense to us today because we’re embedded in our whole “current framework”. But without that framework they don’t make sense. And so it will be for the “disembodied soul”. To us, what it does may not make sense. But to it, with its “current framework”, it will.

Could we “learn how to make sense of it”? There’s likely to be a certain barrier of computational irreducibility: in effect the only way to “understand the soul of the future” is to retrace its steps to get to where it is. So from our vantage point today, we’re separated by a certain “irreducible distance”, in effect in rulial space.

But could there be some science of the future that will at least tell us general things about how such “souls” behave? Even when there’s computational irreducibility we know that there will always be pockets of computational reducibility—and thus features of behavior that are predictable. But will those features be “interesting”, say from our vantage point today? Maybe some of them will be. Maybe they’ll show us some kind of metapsychology of souls. But inevitably they can only go so far. Because in order for those souls to even experience the passage of time there has to be computational irreducibility. If too much of what happens is too predictable, it’s as if “nothing is happening”—or at least nothing “meaningful”.

And, yes, this is all tied up with questions about “free will”. Even when there’s a disembodied soul that’s operating according to some completely deterministic underlying program, computational irreducibility means its behavior can still “seem free”—because nothing can “outrun it” and say what it’s going to be. And the “inner experience” of the disembodied soul can be significant: it’s “intrinsically defining its future”, not just “having its future defined for it”.

One might have assumed that once everything is just “visibly operating” as “mere computation” it would necessarily be “soulless” and “meaningless”. But computational irreducibility is what breaks out of this, and what allows there to be something irreducible and “meaningful” achieved. And it’s the same phenomenon whether one’s talking about our life now in the physical universe, or a future “disembodied” computational existence. Or in other words, even if absolutely everything—even our very existence—has been “automated by computation”, that doesn’t mean we can’t have a perfectly good “inner experience” of meaningful existence.

Generalized Economics and the Concept of Progress

If we look at human history—or, for that matter, the history of life on Earth—there’s a certain pervasive sense that there’s some kind of “progress” happening. But what fundamentally is this “progress”? One can view it as the process of things being done at a progressively “higher level”, so that in effect “more of what’s important” can happen with a given effort. This idea of “going to a higher level” takes many forms—but they’re all fundamentally about eliding details below, and being able to operate purely in terms of the “things one cares about”.

In technology, this shows up as automation, in which what used to take lots of detailed steps gets packaged into something that can be done “with the push of a button”. In science—and the intellectual realm in general—it shows up as abstraction, where what used to involve lots of specific details gets packaged into something that can be talked about “purely collectively”. And in biology it shows up as some structure (ribosome, cell, wing, etc.) that can be treated as a “modular unit”.

That it’s possible to “do things at a higher level” is a reflection of being able to find “pockets of computational reducibility”. And—as we mentioned above—the fact that (given underlying computational irreducibility) there are necessarily an infinite number of such pockets means that “progress can always go on forever”.

When it comes to human affairs we tend to value such progress highly, because (at least for now) we live finite lives, and insofar as we “want more to happen”, “progress” makes that possible. It’s certainly not self-evident that having more happen is “good”; one might just “want a quiet life”. But there is one constraint that in a sense originates from the deep foundations of biology.

If something doesn’t exist, then nothing can ever “happen to it”. So in biology, if one’s going to have anything “happen” with organisms, they’d better not be extinct. But the physical environment in which biological organisms exist is finite, with many resources that are finite. And given organisms with finite lives, there’s an inevitability to the process of biological evolution, and to the “competition” for resources between organisms.

Will there eventually be an “ultimate winning organism”? Well, no, there can’t be—because of computational irreducibility. There’ll in a sense always be more to explore in the computational universe—more “raw computational material for possible organisms”. And given any “fitness criterion” (like—in a Turing machine analog—“living longer before halting”) there’ll always be a way to “do better” with it.

One might still wonder, however, whether perhaps biological evolution—with its underlying process of random genetic mutation—could “get stuck” and never be able to discover some “way to do better”. And indeed simple models of evolution might give one the intuition that this would happen. But actual evolution seems more like deep learning with a large neural net—where one’s effectively operating in an extremely high-dimensional space where there’s typically always a “way to get there from here”, at least given enough time.

But, OK, so from our history of biological evolution there’s a certain built-in sense of “competition for scarce resources”. And this sense of competition has (so far) also carried over to human affairs. And indeed it’s the basic driver for most of the processes of economics.

But what if resources aren’t “scarce” anymore? What if progress—in the form of automation, or AI—makes it easy to “get anything one wants”? We might imagine robots building everything, AIs figuring everything out, etc. But there are still things that are inevitably scarce. There’s only so much real estate. Only one thing can be “the first ___”. And, in the end, if we have finite lives, we only have so much time.

Still, the more efficient—or high level—the things we do (or have) are, the more we’ll be able to get done in the time we have. And it seems as if what we perceive as “economic value” is intimately connected with “making things higher level”. A finished phone is “worth more” than its raw materials. An organization is “worth more” than its separate parts. But what if we could have “infinite automation”? Then in a sense there’d be “infinite economic value everywhere”, and one might imagine there’d be “no competition left”.

But once again computational irreducibility stands in the way. Because it tells us there’ll never be “infinite automation”, just as there’ll never be an ultimate winning biological organism. There’ll always be “more to explore” in the computational universe, and different paths to follow.

What will this look like in practice? Presumably it’ll lead to all sorts of diversity. So that, for example, a chart of “what the components of an economy are” will become more and more fragmented; it won’t just be “the single winning economic activity is ___”.

There is one potential wrinkle in this picture of unending progress. What if nobody cares? What if the innovations and discoveries just don’t matter, say to us humans? And, yes, there is of course plenty in the world that at any given time in history we don’t care about. That piece of silicon we’ve been able to pick out? It’s just part of a rock. Well, until we start making microprocessors out of it.

But as we’ve discussed, as soon as we’re “operating at some level of abstraction” computational irreducibility makes it inevitable that we’ll eventually be exposed to things that “require going beyond that level”.

But then—critically—there will be choices. There will be different paths to explore (or “mine”) in the computational universe—in the end infinitely many of them. And whatever the computational resources of AIs etc. might be, they’ll never be able to explore all of them. So something—or someone—will have to make a choice of which ones to take.

Given a particular set of things one cares about at a particular point, one might successfully be able to automate all of them. But computational irreducibility implies there will always be a “frontier”, where choices have to be made. And there’s no “right answer”; no “theoretically derivable” conclusion. Instead, if we humans are involved, this is where we get to define what’s going to happen.

How will we do that? Well, ultimately it’ll be based on our history—biological, cultural, etc. We’ll get to use all that irreducible computation that went into getting us to where we are to define what to do next. In a sense it’ll be something that goes “through us”, and that uses what we are. It’s the place where—even when there’s automation all around—there’s still always something us humans can “meaningfully” do.

How Can We Tell the AIs What to Do?

Let’s say we want an AI (or any computational system) to do a particular thing. We might think we could just set up its rules (or “program it”) to do that thing. And indeed for certain kinds of tasks that works just fine. But the deeper the use we make of computation, the more we’re going to run into computational irreducibility, and the less we’ll be able to know how to set up particular rules to achieve what we want.

And then, of course, there’s the question of defining what “we want” in the first place. Yes, we could have specific rules that say what particular pattern of bits should occur at a particular point in a computation. But that probably won’t have much to do with the kind of overall “human-level” objective that we typically care about. And indeed for any objective we can even reasonably define, we’d better be able to coherently “form a thought” about it. Or, in effect, we’d better have some “human-level narrative” to describe it.

But how can we represent such a narrative? Well, we have natural language—probably the single most important innovation in the history of our species. And what natural language fundamentally does is to allow us to talk about things at a “human level”. It’s made of words that we can think of as representing “human-level packets of meaning”. And so, for example, the word “chair” represents the human-level concept of a chair. It’s not referring to some particular arrangement of atoms. Instead, it’s referring to any arrangement of atoms that we can usefully conflate into the single human-level concept of a chair, and from which we can deduce things like the fact that we can expect to sit on it, etc.

So, OK, when we’re “talking to an AI” can we expect to just say what we want using natural language? We can definitely get a certain distance—and indeed ChatGPT helps us get further than ever before. But as we try to make things more precise we run into trouble, and the language we need rapidly becomes increasingly ornate, as in the “legalese” of complex legal documents. So what can we do? If we’re going to keep things at the level of “human thoughts” we can’t “reach down” into all the computational details. But yet we want a precise definition of how what we might say can be implemented in terms of those computational details.

Well, there’s a way to deal with this, and it’s one that I’ve personally devoted many decades to: it’s the idea of computational language. When we think about programming languages, they’re things that operate solely at the level of computational details, defining in more or less the native terms of a computer what the computer should do. But the point of a true computational language (and, yes, in the world today the Wolfram Language is the sole example) is to do something different: to define a precise way of talking in computational terms about things in the world (whether concretely countries or minerals, or abstractly computational or mathematical structures).

Out in the computational universe, there’s immense diversity in the “raw computation” that can happen. But there’s only a thin sliver of it that we humans (at least currently) care about and think about. And we can view computational language as defining a bridge between the things we think about and what’s computationally possible. The functions in our computational language (7000 or so of them in the Wolfram Language) are in effect like words in a human language—but now they have a precise grounding in the “bedrock” of explicit computation. And the point is to design the computational language so it’s convenient for us humans to think and express ourselves in (like a vastly expanded analog of mathematical notation), but so it can also be precisely implemented in practice on a computer.

Given a piece of natural language it’s often possible to give a precise, computational interpretation of it—in computational language. And indeed this is exactly what happens in Wolfram|Alpha. Give a piece of natural language and the Wolfram|Alpha NLU system will try to find an interpretation of it as computational language. And from this interpretation, it’s then up to the Wolfram Language to do the computation that’s specified, and give back the results—and potentially synthesize natural language to express them.

As a practical matter, this setup is useful not only for humans, but also for AIs—like ChatGPT. Given a system that produces natural language, the Wolfram|Alpha NLU system can “catch” natural language it is “thrown”, and interpret it as computational language that precisely specifies a potentially irreducible computation to do.

With both natural language and computational language one’s basically “directly saying what one wants”. But an alternative approach—more aligned with machine learning—is just to give examples, and (implicitly or explicitly) say “follow these”. Inevitably there has to be some underlying model for how to do that following—typically in practice just defined by “what a neural net with a certain architecture will do”. But will the result be “right”? Well, the result will be whatever the neural net gives. But typically we’ll tend to consider it “right” if it’s somehow consistent with what we humans would have concluded. And in practice this often seems to happen, presumably because the actual architecture of our brains is somehow similar enough to the architecture of the neural nets we’re using.

But what if we want to “know for sure” what’s going to happen—or, for example, that some particular “mistake” can never be made? Well then we’re presumably thrust back into computational irreducibility, with the result that there’s no way to know, for example, whether a particular set of training examples can lead to a system that’s capable of doing (or not doing) some particular thing.

OK, but let’s say we’re setting up some AI system, and we want to make sure it “doesn’t do anything bad”. There are several levels of issues here. The first is to decide what we mean by “anything bad”. And, as we’ll discuss below, that in itself is very hard. But even if we could abstractly figure this out, how should we actually express it? We could give examples—but then the AI will inevitably have to “extrapolate” from them, in ways we can’t predict. Or we could describe what we want in computational language. It might be difficult to cover “every case” (as it is in present-day human laws, or complex contracts). But at least we as humans can read what we’re specifying. Though even in this case, there’s an issue of computational irreducibility: that given the specification it won’t be possible to work out all its consequences.

What does all this mean? In essence it’s just a reflection of the fact that as soon as there’s “serious computation” (i.e. irreducible computation) involved, one isn’t going to be immediately able to say what will happen. (And in a sense that’s inevitable, because if one could say, it would mean the computation wasn’t in fact irreducible.) So, yes, we can try to “tell AIs what to do”. But it’ll be like many systems in nature (or, for that matter, people): you can set them on a path, but you can’t know for sure what will happen; you just have to wait and see.

A World Run by AIs

In the world today, there are already plenty of things that are being done by AIs. And, as we’ve discussed, there’ll surely be more in the future. But who’s “in charge”? Are we telling the AIs what to do, or are they telling us? Today it’s at best a mixture: AIs suggest content for us (for example from the web), and in general make all sorts of recommendations about what we should do. And no doubt in the future those recommendations will be even more extensive and tightly coupled to us: we’ll be recording everything we do, processing it with AI, and continually annotating with recommendations—say through augmented reality—everything we see. And in some sense things might even go beyond “recommendations”. If we have direct neural interfaces, then we might be making our brains just “decide” they want to do things, so that in some sense we become pure “puppets of the AI”.

And beyond “personal recommendations” there’s also the question of AIs running the systems we use, or in fact running the whole infrastructure of our civilization. Today we ultimately expect people to make large-scale decisions for our world—often operating in systems of rules defined by laws, and perhaps aided by computation, and even what one might call AI. But there may well come a time when it seems as if AIs could just “do a better job than humans”, say at running a central bank or waging a war.

One might ask how one would ever know if the AI would “do a better job”. Well, one could try tests, and run examples. But once again one’s faced with computational irreducibility. Yes, the particular tests one tries might work fine. But one can’t ultimately predict everything that could happen. What will the AI do if there’s suddenly a never-before-seen seismic event? We basically won’t know until it happens.

But can we be sure the AI won’t do anything “crazy”? Could we—with some definition of “crazy”—effectively “prove a theorem” that the AI can never do that? For any realistically nontrivial definition of crazy we’ll again run into computational irreducibility—and this won’t be possible.

Of course, if we’ve put a person (or even a group of people) “in charge” there’s also no way to “prove” that they won’t do anything “crazy”—and history shows that people in charge quite often have done things that, at least in retrospect, we consider “crazy”. But even though at some level there’s no more certainty about what people will do than about what AIs might do, we still get a certain comfort when people are in charge if we think that “we’re in it together”, and that if something goes wrong those people will also “feel the effects”.

But still, it seems inevitable that lots of decisions and actions in the world will be taken directly by AIs. Perhaps it’ll be because this will be cheaper. Perhaps the results (based on tests) will be better. Or perhaps, for example, things will just have to be done too quickly and in numbers too large for us humans to be in the loop.

But, OK, if a lot of what happens in our world is happening through AIs, and the AIs are effectively doing irreducible computations, what will this be like? We’ll be in a situation where things are “just happening” and we don’t quite know why. But in a sense we’ve very much been in this situation before. Because it’s what happens all the time in our interaction with nature.

Processes in nature—like, for example, the weather—can be thought of as corresponding to computations. And much of the time there’ll be irreducibility in those computations. So we won’t be able to readily predict them. Yes, we can do natural science to figure out some aspects of what’s going to happen. But it’ll inevitably be limited.

And so we can expect it to be with the “AI infrastructure” of the world. Things are happening in it—as they are in the weather—that we can’t readily predict. We’ll be able to say some things—though perhaps in ways that are closer to psychology or social science than to traditional exact science. But there’ll be surprises—like maybe some strange AI analog of a hurricane or an ice age. And in the end all we’ll really be able to do is to try to build up our human civilization so that such things “don’t fundamentally matter” to it.

In a sense the picture we have is that in time there’ll be a whole “civilization of AIs” operating—like nature—in ways that we can’t readily understand. And like with nature, we’ll coexist with it.

But at least at first we might think there’s an important difference between nature and AIs. Because we imagine that we don’t “pick our natural laws”—yet insofar as we’re the ones building the AIs we imagine we can “pick their laws”. But both parts of this aren’t quite right. Because in fact one of the implications of our Physics Project is precisely that the laws of nature that we perceive are the way they are because we are observers who are the way we are. And on the AI side, computational irreducibility implies that we can’t expect to be able to determine the final behavior of the AIs just from knowing the underlying laws we gave them.

But what will the “emergent laws” of the AIs be? Well, just like in physics, it’ll depend on how we “sample” the behavior of the AIs. If we look down at the level of individual bits, it’ll be like looking at molecular dynamics (or the behavior of atoms of space). But typically we won’t do this. And just like in physics, we’ll operate as computationally bounded observers—measuring only certain aggregated features of an underlying computationally irreducible process. But what will the “overall laws of AIs” be like? Maybe they’ll show close analogies to physics. Or maybe they’ll seem more like psychological theories (superegos for AIs?). But we can expect them in many ways to be like large-scale laws of nature of the kind we know.

Still, there’s one more difference between at least our interaction with nature and with AIs. Because we have in effect been “co-evolving” with nature for billions of years—yet AIs are “new on the scene”. And through our co-evolution with nature we’ve developed all sorts of structural, sensory and cognitive features that allow us to “interact successfully” with nature. But with AIs we don’t have these. So what does this mean?

Well, our ways of interacting with nature can be thought of as leveraging pockets of computational reducibility that exist in natural processes—to make things seem at least somewhat predictable to us. But without having found such pockets for AIs, we’re likely to be faced with much more “raw computational irreducibility”—and thus much more unpredictability. It’s been a conceit of modern times that—particularly with the help of science—we’ve been able to make more and more of our world predictable to us, though in practice a large part of what’s led to this is the way we’ve built and controlled the environment in which we live, and the things we choose to do.

But for the new “AI world”, we’re effectively starting from scratch. And to make things predictable in that world may be partly a matter of some new science, but perhaps more importantly a matter of choosing how we set up our “way of life” around the AIs there. (And, yes, if there’s lots of unpredictability we may be back to more ancient points of view about the importance of fate—or we may view AIs as a bit like the Olympians of Greek mythology, duking it out among themselves and sometimes having an effect on mortals.)

Governance in an AI World

Let’s say the world is effectively being run by AIs, but let’s assume that we humans have at least some control over what they do. Then what principles should we have them follow? And what, for example, should their “ethics” be?

Well, the first thing to say is that there’s no ultimate, theoretical “right answer” to this. There are many ethical and other principles that AIs could follow. And it’s basically just a choice which ones should be followed.

When we talk about “principles” and “ethics” we tend to think more in terms of constraints on behavior than in terms of rules for generating behavior. And that means we’re dealing with something more like mathematical axioms, where we ask things like what theorems are true according to those axioms, and what are not. And that means there can be issues like whether the axioms are consistent—and whether they’re complete, in the sense that they can “determine the ethics of anything”. But now, once again, we’re face to face with computational irreducibility, here in the form of Gödel’s theorem and its generalizations.

And what this means is that it’s in general undecidable whether any given set of principles is inconsistent, or incomplete. One might “ask an ethical question”, and find that there’s a “proof chain” of unbounded length to determine what the answer to that question is within one’s specified ethical system, or whether there is even a consistent answer.

One might imagine that somehow one could add axioms to “patch up” whatever issues there are. But Gödel’s theorem basically says that it’ll never work. It’s the same story as so often with computational irreducibility: there’ll always be “new situations” that can arise, that in this case can’t be captured by a finite set of axioms.

OK, but let’s imagine we’re picking a collection of principles for AIs. What criteria could we use to do it? One might be that these principles won’t inexorably lead to a simple state—like one where the AIs are extinct, or have to keep looping doing the same thing forever. And there may be cases where one can readily see that some set of principles will lead to such outcomes. But most of the time, computational irreducibility (here in the form of things like the halting problem) will once again get in the way, and one won’t be able to tell what will happen, or successfully pick “viable principles” this way.

So this means that there are going to be a wide range of principles that we could in theory pick. But presumably what we’ll want is to pick ones that make AIs give us humans some sort of “good time”, whatever that might mean.

And a minimal idea might be to get AIs just to observe what we humans do, and then somehow imitate this. But most people wouldn’t consider this the right thing. They’d point out all the “bad” things people do. And they’d perhaps say “let’s have the AIs follow not what we actually do, but what we aspire to do”.

But where should we get these aspirations from? Different people, and different cultures, can have very different aspirations—with very different resulting principles. So whose should we pick? And, yes, there are pitifully few—if any—principles that we truly find in common everywhere. (Though, for example, the major religions all tend to share things like respect for human life, the Golden Rule, etc.)

But do we in fact have to pick one set of principles? Maybe some AIs can have some principles, and some can have others. Maybe it should be like different countries, or different online communities: different principles for different groups or in different places.

Right now that doesn’t seem plausible, because technological and commercial forces have tended to make it seem as if powerful AIs always have to be centralized. But I expect that this is just a feature of the present time, and not something intrinsic to any “human-like” AI.

So could everyone (and maybe every organization) have “their own AI” with its own principles? For some purposes this might work OK. But there are many situations where AIs (or people) can’t really act independently, and where there have to be “collective decisions” made.

Why is this? In some cases it’s because everyone is in the same physical environment. In other cases it’s because if there’s to be social cohesion—of the kind needed to support even something like a language that’s useful for communication—then there has to be certain conceptual alignment.

It’s worth pointing out, though, that at some level having a “collective conclusion” is effectively just a way of introducing certain computational reducibility to make it “easier to see what to do”. And potentially it can be avoided if one has enough computation capability. For example, one might assume that there has to be a collective conclusion about which side of the road cars should drive on. But that wouldn’t be true if every car had the computation capability to just compute a trajectory that would for example optimally weave around other cars using both sides of the road.

But if we humans are going to be in the loop, we presumably need a certain amount of computational reducibility to make our world sufficiently comprehensible to us that we can operate in it. So that means there’ll be collective—“societal”—decisions to make. We might want to just tell the AIs to “make everything as good as it can be for us”. But inevitably there will be tradeoffs. Making a collective decision one way might be really good for 99% of people, but really bad for 1%; making it the other way might be pretty good for 60%, but pretty bad for 40%. So what should the AI do?

And, of course, this is a classic problem of political philosophy, and there’s no “right answer”. And in reality the setup won’t be as clean as this. It may be fairly easy to work out some immediate effects of different courses of action. But inevitably one will eventually run into computational irreducibility—and “unintended consequences”—and so one won’t be able to say with certainty what the ultimate effects (good or bad) will be.

But, OK, so how should one actually make collective decisions? There’s no perfect answer, but in the world today, democracy in one form or another is usually viewed as the best option. So how might AI affect democracy—and perhaps improve on it? Let’s assume first that “humans are still in charge”, so that it’s ultimately their preferences that matter. (And let’s also assume that humans are more or less in their “current form”: unique and unreplicable discrete entities that believe they have independent minds.)

The basic setup for current democracy is computationally quite simple: discrete votes (or perhaps rankings) are given (sometimes with weights of various kinds), and then numerical totals are used to determine the winner (or winners). And with past technology this was pretty much all that could be done. But now there are some new elements. Imagine not casting discrete votes, but instead using computational language to write a computational essay to describe one’s preferences. Or imagine having a conversation with a linguistically enabled AI that can draw out and debate one’s preferences, and eventually summarize them in some kind of feature vector. Then imagine feeding computational essays or feature vectors from all “voters” to some AI that “works out the best thing to do”.

Well, there are still the same political philosophy issues. It’s not like 60% of people voted for A and 40% for B, so one chose A. It’s much more nuanced. But one still won’t be able to make everyone happy all the time, and one has to have some base principles to know what to do about that.

And there’s a higher-order problem in having an AI “rebalance” collective decisions all the time based on everything it knows about people’s detailed preferences (and perhaps their actions too): for many purposes—like us being able to “keep track of what’s going on”—it’s important to maintain consistency over time. But, yes, one could deal with this by having the AI somehow also weigh consistency in figuring out what to do.

But while there are no doubt ways in which AI can “tune up” democracy, AI doesn’t seem—in and of itself—to deliver any fundamentally new solution for making collective decisions, and for governance in general.

And indeed, in the end things always seem to come down to needing some fundamental set of principles about how one wants things to be. Yes, AIs can be the ones to implement these principles. But there are many possibilities for what the principles could be. And—at least if we humans are “in charge”—we’re the ones who are going to have to come up with them.

Or, in other words, we need to come up with some kind of “AI constitution”. Presumably this constitution should basically be written in precise computational language (and, yes, we’re trying to make it possible for the Wolfram Language to be used), but inevitably (as yet another consequence of computational irreducibility) there’ll be “fuzzy” definitions and distinctions, that will rely on things like examples, “interpolated” by systems like neural nets. Maybe when such a constitution is created, there’ll be multiple “renderings” of it, which can all be applied whenever the constitution is used, with some mechanism for picking the “overall conclusion”. (And, yes, there’s potentially a certain “observer-dependent” multicomputational character to this.)

But whatever its detailed mechanisms, what should the AI constitution say? Different people and groups of people will definitely come to different conclusions about it. And presumably—just as there are different countries, etc. today with different systems of laws—there’ll be different groups that want to adopt different AI constitutions. (And, yes, the same issues about collective decision making apply again when those AI constitutions have to interact.)

But given an AI constitution, one has a base on which AIs can make decisions. And on top of this one imagines a huge network of computational contracts that are autonomously executed, essentially to “run the world”.

And this is perhaps one of those classic “what could possibly go wrong?” moments. An AI constitution has been agreed on, and now everything is being run efficiently and autonomously by AIs that are following it. Well, once again, computational irreducibility rears its head. Because however carefully the AI constitution is drafted, computational irreducibility implies that one won’t be able to foresee all its consequences: “unexpected” things will always happen—and some of them will undoubtedly be things “one doesn’t like”.

In human legal systems there’s always a mechanism for adding “patches”—filling in laws or precedents that cover new situations that have come up. But if everything is being autonomously run by AIs there’s no room for that. Yes, we as humans might characterize “bad things that happen” as “bugs” that could be fixed by adding a patch. But the AI is just supposed to be operating—essentially axiomatically—according to its constitution, so it has no way to “see that it’s a bug”.

Similar to what we discussed above, there’s an interesting analogy here with human law versus natural law. Human law is something we define and can modify. Natural law is something the universe just provides us (notwithstanding the issues about observers discussed above). And by “setting an AI constitution and letting it run” we’re basically forcing ourselves into a situation where the “civilization of the AIs” is some “independent stratum” in the world, that we essentially have to take as it is, and adapt to.

Of course, one might wonder if the AI constitution could “automatically evolve”, say based on what’s actually seen to happen in the world. But one quickly returns to the exact same issues of computational irreducibility, where one can’t predict whether the evolution will be “right”, etc.

So far, we’ve assumed that in some sense “humans are in charge”. But at some level that’s an issue for the AI constitution to define. It’ll have to define whether AIs have “independent rights”—just like humans (and, in many legal systems, some other entities too). Closely related to the question of independent rights for AIs is whether an AI can be considered autonomously “responsible for its actions”—or whether such responsibility must always ultimately rest with the (presumably human) creator or “programmer” of the AI.

Once again, computational irreducibility has something to say. Because it implies that the behavior of the AI can go “irreducibly beyond” what its programmer defined. And in the end (as we discussed above) this is the same basic mechanism that allows us humans to effectively have “free will” even when we’re ultimately operating according to deterministic underlying natural laws. So if we’re going to claim that we humans have free will, and can be “responsible for our actions” (as opposed to having our actions always “dictated by underlying laws”) then we’d better claim the same for AIs.

So just as a human builds up something irreducible and irreplaceable in the course of their life, so can an AI. As a practical matter, though, AIs can presumably be backed up, copied, etc.—which isn’t (yet) possible for humans. So somehow their individual instances don’t seem as valuable, even if the “last copy” might still be valuable. As humans, we might want to say “those AIs are something inferior; they shouldn’t have rights”. But things are going to get more entangled. Imagine a bot that no longer has an identifiable owner but that’s successfully befriending people (say on social media), and paying for its underlying operation from donations, ads, etc. Can we reasonably delete that bot? We might argue that “the bot can feel no pain”—but that’s not true of its human friends. But what if the bot starts doing “bad” things? Well, then we’ll need some form of “bot justice”—and pretty soon we’ll find ourselves building a whole human-like legal structure for the AIs.

So Will It End Badly?

OK, so AIs will learn what they can from us humans, then they’ll fundamentally just be running as autonomous computational systems—much like nature runs as an autonomous computational system—sometimes “interacting with us”. What will they “do to us”? Well, what does nature “do to us”? In a kind of animistic way, we might attribute intentions to nature, but ultimately it’s just “following its rules” and doing what it does. And so it will be with AIs. Yes, we might think we can set things up to determine what the AIs will do. But in the end—insofar as the AIs are really making use of what’s possible in the computational universe—there’ll inevitably be computational irreducibility, and we won’t be able to foresee what will happen, or what consequences it will have.

So will the dynamics of AIs in fact have “bad” effects—like, for example, wiping us out? Well, it’s perfectly possible nature could wipe us out too. But one has the feeling that—extraterrestrial “accidents” aside—the natural world around us is at some level enough in some kind of “equilibrium” that nothing too dramatic will happen. But AIs are something new. So maybe they’ll be different.

And one possibility might be that AIs could “improve themselves” to produce a single “apex intelligence” that would in a sense dominate everything else. But here we can see computational irreducibility as coming to the rescue. Because it implies that there can never be a “best at everything” computational system. It’s a core result of the emerging field of metabiology: that whatever “achievement” you specify, there’ll always be a computational system somewhere out there in the computational universe that will exceed it. (A simple example is that there’s always a Turing machine that can be found that will exceed any upper bound you specify on the time it takes to halt.)

So what this means is that there’ll inevitably be a whole “ecosystem” of AIs—with no single winner. Of course, while that might be an inevitable final outcome, it might not be what happens in the shorter term. And indeed the current tendency to centralize AI systems has a certain danger of AI behavior becoming “unstabilized” relative to what it would be with a whole ecosystem of “AIs in equilibrium”.

And in this situation there’s another potential concern as well. We humans are the product of a long struggle for life played out over the course of the history of biological evolution. And insofar as AIs inherit our attributes we might expect them to inherit a certain “drive to win”—perhaps also against us. And perhaps this is where the AI constitution becomes important: to define a “contract” that supersedes what AIs might “naturally” inherit from effectively observing our behavior. Eventually we can expect the AIs to “independently reach equilibrium”. But in the meantime, the AI constitution can help break their connection with our “competitive” history of biological evolution.

Preparing for an AI World

We’ve talked quite a bit about the ultimate future course of AIs, and their relation to us humans. But what about the short term? How today can we prepare for the growing capabilities and uses of AIs?

As has been true throughout history, people who use tools tend to do better than those who don’t. Yes, you can go on doing by direct human effort what has now been successfully automated, but except in rare cases you’ll increasingly be left behind. And what’s now emerging is an extremely powerful combination of tools: neural-net-style AI for “immediate human-like tasks”, along with computational language for deeper access to the computational universe and computational knowledge.

So what should people do with this? The highest leverage will come from figuring out new possibilities—things that weren’t possible before but have now “come into range” as a result of new capabilities. And as we discussed above, this is a place where we humans are inevitably central contributors—because we’re the ones who must define what we consider has value for us.

So what does this mean for education? What’s worth learning now that so much has been automated? I think the fundamental answer is how to think as broadly and deeply as possible—calling on as much knowledge and as many paradigms as possible, and particularly making use of the computational paradigm, and ways of thinking about things that directly connect with what computation can help with.

In the course of human history a lot of knowledge has been accumulated. But as ways of thinking have advanced, it’s become unnecessary to learn directly that knowledge in all its detail: instead one can learn things at a higher level, abstracting out many of the specific details. But in the past few decades something fundamentally new has come on the scene: computers and the things they enable.

For the first time in history, it’s become realistic to truly automate intellectual tasks. The leverage this provides is completely unprecedented. And we’re only just starting to come to terms with what it means for what and how we should learn. But with all this new power there’s a tendency to think something must be lost. Surely it must still be worth learning all those intricate details—that people in the past worked so hard to figure out—of how to do some mathematical calculation, even though Mathematica has been able to do it automatically for more than a third of a century?

And, yes, at the right time it can be interesting to learn those details. But in the effort to understand and best make use of the intellectual achievements of our civilization, it makes much more sense to leverage the automation we have, and treat those calculations just as “building blocks” that can be put together in “finished form” to do whatever it is we want to do.

One might think this kind of leveraging of automation would just be important for “practical purposes”, and for applying knowledge in the real world. But actually—as I have personally found repeatedly to great benefit over the decades—it’s also crucial at a conceptual level. Because it’s only through automation that one can get enough examples and experience that one’s able to develop the intuition needed to reach a higher level of understanding.

Confronted with the rapidly growing amount of knowledge in the world there’s been a tremendous tendency to assume that people must inevitably become more and more specialized. But with increasing success in the automation of intellectual tasks—and what we might broadly call AI—it becomes clear there’s an alternative: to make more and more use of this automation, so people can operate at a higher level, “integrating” rather than specializing.

And in a sense this is the way to make the best use of our human capabilities: to let us concentrate on setting the “strategy” of what we want to do—delegating the details of how to do it to automated systems that can do it better than us. But, by the way, the very fact that there’s an AI that knows how to do something will no doubt make it easier for humans to learn how to do it too. Because—although we don’t yet have the complete story—it seems inevitable that with modern techniques AIs will be able to successfully “learn how people learn”, and effectively present things an AI “knows” in just the right way for any given person to absorb.

So what should people actually learn? Learn how to use tools to do things. But also learn what things are out there to do—and learn facts to anchor how you think about those things. A lot of education today is about answering questions. But for the future—with AI in the picture—what’s likely to be more important is to learn how to ask questions, and how to figure out what questions are worth asking. Or, in effect, how to lay out an “intellectual strategy” for what to do.

And to be successful at this, what’s going to be important is breadth of knowledge—and clarity of thinking. And when it comes to clarity of thinking, there’s again something new in modern times: the concept of computational thinking. In the past we’ve had things like logic, and mathematics, as ways to structure thinking. But now we have something new: computation.

Does that mean everyone should “learn to program” in some traditional programming language? No. Traditional programming languages are about telling computers what to do in their terms. And, yes, lots of humans do this today. But it’s something that’s fundamentally ripe for direct automation (as examples with ChatGPT already show). And what’s important for the long term is something different. It’s to use the computational paradigm as a structured way to think not about the operation of computers, but about both things in the world and abstract things.

And crucial to this is having a computational language: a language for expressing things using the computational paradigm. It’s perfectly possible to express simple “everyday things” in plain, unstructured natural language. But to build any kind of serious “conceptual tower” one needs something more structured. And that’s what computational language is about.

One can see a rough historical analog in the development of mathematics and mathematical thinking. Up until about half a millennium ago, mathematics basically had to be expressed in natural language. But then came mathematical notation—and from it a more streamlined approach to mathematical thinking, that eventually made possible all the various mathematical sciences. And it’s now the same kind of thing with computational language and the computational paradigm. Except that it’s a much broader story, in which for basically every field or occupation “X” there’s a “computational X” that’s emerging.

In a sense the point of computational language (and all my efforts in the development of the Wolfram Language) is to be able to let people get “as automatically as possible” to computational X—and to let people express themselves using the full power of the computational paradigm.

Something like ChatGPT provides “human-like AI” in effect by piecing together existing human material (like billions of words of human-written text). But computational language lets one tap directly into computation—and gives the ability to do fundamentally new things, that immediately leverage our human capabilities for defining intellectual strategy.

And, yes, while traditional programming is likely to be largely obsoleted by AI, computational language is something that provides a permanent bridge between human thinking and the computational universe: a channel in which the automation is already done in the very design (and implementation) of the language—leaving in a sense an interface directly suitable for humans to learn, and to use as a basis to extend their thinking.

But, OK, what about the future of discovery? Will AIs take over from us humans in, for example, “doing science”? I, for one, have used computation (and many things one might think of as AI) as a tool for scientific discovery for nearly half a century. And, yes, many of my discoveries have in effect been “made by computer”. But science is ultimately about connecting things to human understanding. And so far it’s taken a human to knit what the computer finds into the whole web of human intellectual history.

One can certainly imagine, though, that an AI—even one rather like ChatGPT—could be quite successful in taking a “raw computational discovery” and “explaining” how it might relate to existing human knowledge. One could also imagine that the AI would be successful at identifying what aspects of some system in the world could be picked out to describe in some formal way. But—as is typical for the process of modeling in general—a key step is to decide “what one cares about”, and in effect in what direction to go in extending one’s science. And this—like so much else—is inevitably tied into the specifics of the goals we humans set ourselves.

In the emerging AI world there are plenty of specific skills that won’t make sense for (most) humans to learn—just as today the advance of automation has obsoleted many skills from the past. But—as we’ve discussed—we can expect there to “be a place” for humans. And what’s most important for us humans to learn is in effect how to pick “where next to go”—and where, out of all the infinite possibilities in the computational universe, we should take human civilization.

Afterword: Looking at Some Actual Data

OK, so we’ve talked quite a bit about what might happen in the future. But what about actual data from the past? For example, what’s been the actual history of the evolution of jobs? Conveniently, in the US, the Census Bureau has records of people’s occupations going back to 1850. Of course, many job titles have changed since then. Switchmen (on railroads), chainmen (in surveying) and sextons (in churches) aren’t really things anymore. And telemarketers, aircraft pilots and web developers weren’t things in 1850. But with a bit of effort, it’s possible to more or less match things up—at least if one aggregates into large enough categories.

So here are pie charts of different job categories at 50-year intervals:

And, yes, in 1850 the US was firmly an agricultural economy, with just over half of all jobs being in agriculture. But as agriculture got more efficient—with the introduction of machinery, irrigation, better seeds, fertilizers, etc.—the fraction dropped dramatically, to just a few percent today.

After agriculture, the next biggest category back in 1850 was construction (along with other real-estate-related jobs, mainly maintenance). And this is a category that for a century and a half hasn’t changed much in size (at least so far), presumably because, even though there’s been greater automation, this has just allowed buildings to be more complex.

Looking at the pie charts above, we can see a clear trend towards greater diversification in jobs (and indeed the same thing is seen in the development of other economies around the world). It’s an old theory in economics that increasing specialization is related to economic growth, but from our point of view here, we might say that the very possibility of a more complex economy, with more niches and jobs, is a reflection of the inevitable presence of computational irreducibility, and the complex web of pockets of computational reducibility that it implies.

Beyond the overall distribution of job categories, we can also look at trends in individual categories over time—with each one in a sense providing a certain window onto history:

One can definitely see cases where the number of jobs decreases as a result of automation. And this happens not only in areas like agriculture and mining, but also for example in finance (fewer clerks and bank tellers), as well as in sales and retail (online shopping). Sometimes—as in the case of manufacturing—there’s a decrease of jobs partly because of automation, and partly because the jobs move out of the US (mainly to countries with lower labor costs).

There are cases—like military jobs—where there are clear “exogenous” effects. And then there are cases like transportation+logistics where there’s a steady increase for more than half a century as technology spreads and infrastructure gets built up—but then things “saturate”, presumably at least partly as a result of increased automation. It’s a somewhat similar story with what I’ve called “technical operations”—with more “tending to technology” needed as technology becomes more widespread.

Another clear trend is an increase in job categories associated with the world becoming an “organizationally more complicated place”. Thus we see increases in management, as well as administration, government, finance and sales (which all have recent decreases as a result of computerization). And there’s also a (somewhat recent) increase in legal.

Other areas with increases include healthcare, engineering, science and education—where “more is known and there’s more to do” (as well as there being increased organizational complexity). And then there’s entertainment, and food+hospitality, with increases that one might attribute to people leading (and wanting) “more complex lives”. And, of course, there’s information technology which takes off from nothing in the mid-1950s (and which had to be rather awkwardly grafted into the data we’re using here).

So what can we conclude? The data seems quite well aligned with what we discussed in more general terms above. Well-developed areas get automated and need to employ fewer people. But technology also opens up new areas, which employ additional people. And—as we might expect from computational irreducibility—things generally get progressively more complicated, with additional knowledge and organizational structure opening up more “frontiers” where people are needed. But even though there are sometimes “sudden inventions”, it still always seems to take decades (or effectively a generation) for there to be any dramatic change in the number of jobs. (The few sharp changes visible in the plots seem mostly to be associated with specific economic events, and—often related—changes in government policies.)

But in addition to the different jobs that get done, there’s also the question of how individual people spend their time each day. And—while it certainly doesn’t live up to my own (rather extreme) level of personal analytics—there’s a certain amount of data on this that’s been collected over the years (by getting time diaries from randomly sampled people) in the American Heritage Time Use Study. So here, for example, are plots based on this survey for how the amount of time spent on different broad activities has varied over the decades (the main line shows the mean—in hours—for each activity; the shaded areas indicate successive deciles):

And, yes, people are spending more time on “media & computing”, some mixture of watching TV, playing videogames, etc. Housework, at least for women, takes less time, presumably mostly as a result of automation (appliances, etc.). (“Leisure” is basically “hanging out” as well as hobbies and social, cultural, sporting events, etc.; “Civic” includes volunteer, religious, etc. activities.)

If one looks specifically at people who are doing paid work

one notices several things. First, the average number of hours worked hasn’t changed much in half a century, though the distribution has broadened somewhat. For people doing paid work, media & computing hasn’t increased significantly, at least since the 1980s. One category in which there is systematic increase (though the total time still isn’t very large) is exercise.

What about people who—for one reason or another—aren’t doing paid work? Here are corresponding results in this case:

Not so much increase in exercise (though the total times are larger to begin with), but now a significant increase in media & computing, with the average recently reaching nearly 6 hours per day for men—perhaps as a reflection of “more of life going online”.

But looking at all these results on time use, I think the main conclusion that over the past half century, the ways people (at least in the US) spend their time have remained rather stable—even as we’ve gone from a world with almost no computers to a world in which there are more computers than people.

A 50-Year Quest: My Personal Journey with the Second Law of Thermodynamics

2 février 2023 à 22:19

When I Was 12 Years Old…

I’ve been trying to understand the Second Law now for a bit more than 50 years.

It all started when I was 12 years old. Building on an earlier interest in space and spacecraft, I’d gotten very interested in physics, and was trying to read everything I could about it. There were several shelves of physics books at the local bookstore. But what I coveted most was the largest physics book collection there: a series of five plushly illustrated college textbooks. And as a kind of graduation gift when I finished (British) elementary school in June 1972 I arranged to get those books. And here they are, still on my bookshelf today, just a little faded, more than half a century later:

Click to enlarge

For a while the first book in the series was my favorite. Then the third. The second. The fourth. The fifth one at first seemed quite mysterious—and somehow more abstract in its goals than the others:

Click to enlarge

What story was the filmstrip on its cover telling? For a couple of months I didn’t look seriously at the book. And I spent much of the summer of 1972 writing my own (unseen by anyone else for 30+ years) Concise Directory of Physics

Click to enlarge

that included a rather stiff page about energy, mentioning entropy—along with the heat death of the universe.

But one afternoon late that summer I decided I should really find out what that mysterious fifth book was all about. Memory being what it is I remember that—very unusually for me—I took the book to read sitting on the grass under some trees. And, yes, my archives almost let me check my recollection: in the distance, there’s the spot, except in 1967 the trees are significantly smaller, and in 1985 they’re bigger:

Click to enlarge

Of course, by 1972 I was a little bigger than in 1967—and here I am a little later, complete with a book called Planets and Life on the ground, along with a tube of (British) Smarties, and, yes, a pocket protector (but, hey, those were actual ink pens):

Click to enlarge

But back to the mysterious green book. It wasn’t like anything I’d seen before. It was full of pictures like the one on the cover. And it seemed to be saying that—just by looking at those pictures and thinking—one could figure out fundamental things about physics. The other books I’d read had all basically said “physics works like this”. But here was a book saying “you can figure out how physics has to work”. Back then I definitely hadn’t internalized it, but I think what was so exciting that day was that I got a first taste of the idea that one didn’t have to be told how the world works; one could just figure it out:

Click to enlarge

I didn’t yet understand quite a bit of the math in the book. But it didn’t seem so relevant to the core phenomenon the book was apparently talking about: the tendency of things to become more random. I remember wondering how this related to stars being organized into galaxies. Why might that be different? The book didn’t seem to say, though I thought maybe somewhere it was buried in the math.

But soon the summer was over, and I was at a new school, mostly away from my books, and doing things like diligently learning more Latin and Greek. But whenever I could I was learning more about physics—and particularly about the hot area of the time: particle physics. The pions. The kaons. The lambda hyperon. They all became my personal friends. During the school vacations I would excitedly bicycle the few miles to the nearby university library to check out the latest journals and the latest news about particle physics.

The school I was at (Eton) had five centuries of history, and I think at first I assumed no particular bridge to the future. But it wasn’t long before I started hearing mentions that somewhere at the school there was a computer. I’d seen a computer in real life only once—when I was 10 years old, and from a distance. But now, tucked away at the edge of the school, above a bicycle repair shed, there was an island of modernity, a “computer room” with a glass partition separating off a loudly humming desk-sized piece of electronics that I could actually touch and use: an Elliott 903C computer with 8 kilowords of 18-bit ferrite core memory (acquired by the school in 1970 for £12,000, or about about $300k today):

Click to enlarge

At first it was such an unfamiliar novelty that I was satisfied writing little programs to do things like compute primes, print curious patterns on the teleprinter, and play tunes with the built-in diagnostic tone generator. But it wasn’t long before I set my sights on the goal of using the computer to reproduce that interesting picture on the book cover.

I programmed in assembler, with my programs on paper tape. The computer had just 16 machine instructions, which included arithmetic ones, but only for integers. So how was I going to simulate colliding “molecules” with that? Somewhat sheepishly, I decided to put everything on a grid, with everything represented by discrete elements. There was a convention for people to name their programs starting with their own first initial. So I called the program SPART, for “Stephen’s Particle Program”. (Thinking about it today, maybe that name reflected some aspiration of relating this to particle physics.)

It was the most complicated program I had ever written. And it was hard to test, because, after all, I didn’t really know what to expect it to do. Over the course of several months, it went through many versions. Rather often the program would just mysteriously crash before producing any output (and, yes, there weren’t real debugging tools yet). But eventually I got it to systematically produce output. But to my disappointment the output never looked much like the book cover.

I didn’t know why, but I assumed it was because I was simplifying things too much, putting everything on a grid, etc. A decade later I realized that in writing my program I’d actually ended up inventing a form of 2D cellular automaton. And I now rather suspect that this cellular automaton—like rule 30—was actually intrinsically generating randomness, and in some sense showing what I now understand to be the core phenomenon of the Second Law. But at the time I absolutely wasn’t ready for this, and instead I just assumed that what I was seeing was something wrong and irrelevant. (In past years, I had suspected that what went wrong had to do with details of particle behavior on square—as opposed to other—grids. But I now suspect it was instead that the system was in a sense generating too much randomness, making the intended “molecular dynamics” unrecognizable.)

I’d love to “bring SPART back to life”, but I don’t seem to have a copy anymore, and I’m pretty sure the printouts I got as output back in 1973 seemed so “wrong” I didn’t keep them. I do still have quite a few paper tapes from around that time, but as of now I’m not sure what’s on them—not least because I wrote my own “advanced” paper-tape loader, which used what I later learned were error-correcting codes to try to avoid problems with pieces of “confetti” getting stuck in the holes that had been punched in the tape:

Click to enlarge

Becoming a Physicist

I don’t know what would have happened if I’d thought my program was more successful in reproducing “Second Law” behavior back in 1973 when I was 13 years old. But as it was, in the summer of 1973 I was away from “my” computer, and spending all my time on particle physics. And between that summer and early 1974 I wrote a book-length summary of what I called “The Physics of Subatomic Particles”:

Click to enlarge

I don’t think I’d looked at this in any detail in 48 years. But reading it now I am a bit shocked to find history and explanations that I think are often better than I would immediately give today—even if they do bear definite signs of coming from a British early teenager writing “scientific prose”.

Did I talk about statistical mechanics and the Second Law? Not directly, though there’s a curious passage where I speculate about the possibility of antimatter galaxies, and their (rather un-Second-Law-like) segregation from ordinary, matter galaxies:

Click to enlarge

By the next summer I was writing the 230-page, much more technical “Introduction to the Weak Interaction”. Lots of quantum mechanics and quantum field theory. No statistical mechanics. The closest it gets is a chapter on CP violation (AKA time-reversal violation)—a longtime favorite topic of mine—but from a very particle-physics point of view. By the next year I was publishing papers about particle physics, with no statistical mechanics in sight—though in a picture of me (as a “lanky youth”) from that time, the Statistical Physics book is right there on my shelf, albeit surrounded by particle physics books:

Click to enlarge

But despite my focus on particle physics, I still kept thinking about statistical mechanics and the Second Law, and particularly its implications for the large-scale structure of the universe, and things like the possibility of matter-antimatter separation. And in early 1977, now 17 years old, and (briefly) a college student in Oxford, my archives record that I gave a talk to the newly formed (and short-lived) Oxford Natural Science Club entitled “Whither Physics” in which I talked about “large, small, many” as the main frontiers of physics, and presented the visual

Click to enlarge

with a dash of “unsolved purple” impinging on statistical mechanics, particularly in connection with non-equilibrium situations. Meanwhile, looking at my archives today, I find some “back of the envelope” equilibrium statistical mechanics from that time (though I have no idea now what this was about):

Click to enlarge

But then, in the fall of 1977 I ended up for the first time really needing to use statistical mechanics “in production”. I had gotten interested in what would later become a hot area: the intersection between particle physics and the early universe. One of my interests was neutrino background radiation (the neutrino analog of the cosmic microwave background); another was early-universe production of stable charged particles heavier than the proton. And it turned out that to study these I needed all three of cosmology, particle physics, and statistical mechanics:

Click to enlarge

In the couple of years that followed, I worked on all sorts of topics in particle physics and in cosmology. Quite often ideas from statistical mechanics would show up, like when I worked on the hadronization of quarks and gluons, or when I worked on phase transitions in the early universe. But it wasn’t until 1979 that the Second Law made its first explicit appearance by name in my published work.

I was studying how there could be a net excess of matter over antimatter throughout the universe (yes, I’d by then given up on the idea of matter-antimatter separation). It was a subtle story of quantum field theory, time reversal violation, General Relativity—and non-equilibrium statistical mechanics. And in the paper we wrote we included a detailed appendix about Boltzmann’s H theorem and the Second Law—and the generalization we needed for relativistic quantum time-reversal-violating systems in an expanding universe:

Click to enlarge

All this got me thinking again about the foundations of the Second Law. The physicists I was around mostly weren’t too interested in such topics—though Richard Feynman was something of an exception. And indeed when I did my PhD thesis defense in November 1979 it ended up devolving into a spirited multi-hour debate with Feynman about the Second Law. He maintained that the Second Law must ultimately cause everything to randomize, and that the order we see in the universe today must be some kind of temporary fluctuation. I took the point of view that there was something else going on, perhaps related to gravity. Today I would have more strongly made the rather Feynmanesque point that if you have a theory that says everything we observe today is an exception to your theory, then the theory you have isn’t terribly useful.

Statistical Mechanics and Simple Programs

Back in 1973 I never really managed to do much science on the very first computer I used. But by 1976 I had access to much bigger and faster computers (as well as to the ARPANET—forerunner of the internet). And soon I was routinely using computers as powerful tools for physics, and particularly for symbolic manipulation. But by late 1979 I had basically outgrown the software systems that existed, and within weeks of getting my PhD I embarked on the project of building my own computational system.

It’s a story I’ve told elsewhere, but one of the important elements for our purposes here is that in designing the system I called SMP (for “Symbolic Manipulation Program”) I ended up digging deeply into the foundations of computation, and its connections to areas like mathematical logic. But even as I was developing the critical-to-Wolfram-Language-to-this-day paradigm of basing everything on transformations for symbolic expressions, as well as leading the software engineering to actually build SMP, I was also continuing to think about physics and its foundations.

There was often something of a statistical mechanics orientation to what I did. I worked on cosmology where even the collection of possible particle species had to be treated statistically. I worked on the quantum field theory of the vacuum—or effectively the “bulk properties of quantized fields”. I worked on what amounts to the statistical mechanics of cosmological strings. And I started working on the quantum-field-theory-meets-statistical-mechanics problem of “relativistic matter” (where my unfinished notes contain questions like “Does causality forbid relativistic solids?”):

Click to enlarge

But hovering around all of this was my old interest in the Second Law, and in the seemingly opposing phenomenon of the spontaneous emergence of complex structure.

SMP Version 1.0 was ready in mid-1981. And that fall, as a way to focus my efforts, I taught a “Topics in Theoretical Physics” course at Caltech (supposedly for graduate students but actually almost as many professors came too) on what, for want of a better name, I called “non-equilibrium statistical mechanics”. My notes for the first lecture dived right in:

Click to enlarge

Echoing what I’d seen on that book cover back in 1972 I talked about the example of the expansion of a gas, noting that even in this case “Many features [are] still far from understood”:

Click to enlarge

I talked about the Boltzmann transport equation and its elaboration in the BBGKY hierarchy, and explored what might be needed to extend it to things like self-gravitating systems. And then—in what must have been a very overstuffed first lecture—I launched into a discussion of “Possible origins of irreversibility”. I began by talking about things like ergodicity, but soon made it clear that this didn’t go the distance, and there was much more to understand—saying that “with a bit of luck” the material in my later lectures might help:

Click to enlarge

I continued by noting that some systems can “develop order and considerable organization”—which non-equilibrium statistical mechanics should be able to explain:

Click to enlarge

I then went quite “cosmological”:

Click to enlarge

The first candidate explanation I listed was the fluctuation argument Feynman had tried to use:

Click to enlarge

I discussed the possibility of fundamental microscopic irreversibility—say associated with time-reversal violation in gravity—but largely dismissed this. I talked about the possibility that the universe could have started in a special state in which “the matter is in thermal equilibrium, but the gravitational field is not.” And finally I gave what the 22-year-old me thought at the time was the most plausible explanation:

Click to enlarge

All of this was in a sense rooted in a traditional mathematical physics style of thinking. But the second lecture gave a hint of a quite different approach:

Click to enlarge

In my first lecture, I had summarized my plans for subsequent lectures:

Click to enlarge

But discovery intervened. People had discussed reaction-diffusion patterns as examples of structure being formed “away from equilibrium”. But I was interested in more dramatic examples, like galaxies, or snowflakes, or turbulent flow patterns, or forms of biological organisms. What kinds of models could realistically be made for these? I started from neural networks, self-gravitating gases and spin systems, and just kept on simplifying and simplifying. It was rather like language design, of the kind I’d done for SMP. What were the simplest primitives from which I could build up what I wanted?

Before long I came up with what I’d soon learn could be called one-dimensional cellular automata. And immediately I started running them on a computer to see what they did:

Click to enlarge

And, yes, they were “organizing themselves”—even from random initial conditions—to make all sorts of structures. By December I was beginning to frame how I would write about what was going on:

Click to enlarge

And by May 1982 I had written my first long paper about cellular automata (published in 1983 under the title “Statistical Mechanics of Cellular Automata”):

Click to enlarge

The Second Law featured prominently, even in the first sentence:

Click to enlarge

I made quite a lot out of the fundamentally irreversible character of most cellular automaton rules, pretty much assuming that this was the fundamental origin of their ability to “generate complex structures”—as the opening transparencies of two talks I gave at the time suggested:

Click to enlarge

It wasn’t that I didn’t know there could be reversible cellular automata. And a footnote in my paper even records the fact these can generate nested patterns with a certain fractal dimension—as computed in a charmingly manual way on a couple of pages I now find in my archives:

Click to enlarge

But somehow I hadn’t quite freed myself from the assumption that microscopic irreversibility was what was “causing” structures to be formed. And this was related to another important—and ultimately incorrect—assumption: that all the structure I was seeing was somehow the result of the “filtering” random initial conditions. Right there in my paper is a picture of rule 30 starting from a single cell:

Click to enlarge

And, yes, the printout from which that was made is still in my archives, if now a little worse for wear:

Click to enlarge

Of course, it probably didn’t help that with my “display” consisting of an array of printed characters I couldn’t see too much of the pattern—though my archives do contain a long “imitation-high-resolution” printout of the conveniently narrow, and ultimately nested, pattern from rule 225:

Click to enlarge

But I think the more important point was that I just didn’t have the necessary conceptual framework to absorb what I was seeing in rule 30—and I wasn’t ready for the intuitional shock that it takes only simple rules with simple initial conditions to produce highly complex behavior.

My motivation for studying the behavior of cellular automata had come from statistical mechanics. But I soon realized that I could discuss cellular automata without any of the “baggage” of statistical mechanics, or the Second Law. And indeed even as I was finishing my long statistical-mechanics-themed paper on cellular automata, I was also writing a short paper that described cellular automata essentially as purely computational systems (even though I still used the term “mathematical models”) without talking about any kind of Second Law connections:

Click to enlarge

Through much of 1982 I was alternating between science, technology and the startup of my first company. I left Caltech in October 1982, and after stops at Los Alamos and Bell Labs, started working at the Institute for Advanced Study in Princeton in January 1983, equipped with a newly obtained Sun workstation computer whose (“one megapixel”) bitmap display let me begin to see in more detail how cellular automata behave:

Click to enlarge

It had very much the flavor of classic observational science—looking not at something like mollusc shells, but instead at images on a screen—and writing down what I saw in a “lab notebook”:

Click to enlarge

What did all those rules do? Could I somehow find a way to classify their behavior?

Click to enlarge

Mostly I was looking at random initial conditions. But in a near miss of the rule 30 phenomenon I wrote in my lab notebook: “In irregular cases, appears that patterns starting from small initial states are not self-similar (e.g. code 10)”. I even looked again at asymmetric “elementary” rules (of which rule 30 is an example)—but only from random initial conditions (though noting the presence of “class 4” rules, which would include rule 110):

Click to enlarge

My technology stack at the time consisted of printing screen dumps of cellular automaton behavior

Click to enlarge

then using repeated photocopying to shrink them—and finally cutting out the images and assembling arrays of them using Scotch tape:

Click to enlarge

And looking at these arrays I was indeed able to make an empirical classification, identifying initially five—but in the end four—basic classes of behavior. And although I sometimes made analogies with solids, liquids and gases—and used the mathematical concept of entropy—I was now mostly moving away from thinking in terms of statistical mechanics, and was instead using methods from areas like dynamical systems theory, and computation theory:

Click to enlarge

Even so, when I summarized the significance of investigating the computational characteristics of cellular automata, I reached back to statistical mechanics, suggesting that much as information theory provided a mathematical basis for equilibrium statistical mechanics, so similarly computation theory might provide a foundation for non-equilibrium statistical mechanics:

Click to enlarge

Computational Irreducibility and Rule 30

My experiments had shown that cellular automata could “spontaneously produce structure” even from randomness. And I had been able to characterize and measure various features of this structure, notably using ideas like entropy. But could I get a more complete picture of what cellular automata could make? I turned to formal language theory, and started to work out the “grammar of possible states”. And, yes, a quarter century before Graph in Wolfram Language, laying out complicated finite state machines wasn’t easy:

Click to enlarge

But by November 1983 I was writing about “self-organization as a computational process”:

Click to enlarge

The introduction to my paper again led with the Second Law, though now talked about the idea that computation theory might be what could characterize non-equilibrium and self-organizing phenomena:

Click to enlarge

The concept of equilibrium in statistical mechanics makes it natural to ask what will happen in a system after an infinite time. But computation theory tells one that the answer to that question can be non-computable or undecidable. I talked about this in my paper, but then ended by discussing the ultimately much richer finite case, and suggesting (with a reference to NP completeness) that it might be common for there to be no computational shortcut to cellular automaton evolution. And rather presciently, I made the statement that “One may speculate that [this phenomenon] is widespread in physical systems” so that “the consequences of their evolution could not be predicted, but could effectively be found only by direct simulation or observation.”:

Click to enlarge

These were the beginnings of powerful ideas, but I was still tying them to somewhat technical things like ensembles of all possible states. But in early 1984, that began to change. In January I’d been asked to write an article for the then-top popular science magazine Scientific American on the subject of “Computers in Science and Mathematics”. I wrote about the general idea of computer experiments and simulation. I wrote about SMP. I wrote about cellular automata. But then I wanted to bring it all together. And that was when I came up with the term “computational irreducibility”.

By May 26, the concept was pretty clearly laid out in my draft text:

Click to enlarge

But just a few days later something big happened. On June 1 I left Princeton for a trip to Europe. And in order to “have something interesting to look at on the plane” I decided to print out pictures of some cellular automata I hadn’t bothered to look at much before. The first one was rule 30:

Click to enlarge

And it was then that it all clicked. The complexity I’d been seeing in cellular automata wasn’t the result of some kind of “self-organization” or “filtering” of random initial conditions. Instead, here was an example where it was very obviously being “generated intrinsically” just by the process of evolution of the cellular automaton. This was computational irreducibility up close. No need to think about ensembles of states or statistical mechanics. No need to think about elaborate programming of a universal computer. From just a single black cell rule 30 could produce immense complexity, and showed what seemed very likely to be clear computational irreducibility.

Why hadn’t I figured out before that something like this could happen? After all, I’d even generated a small picture of rule 30 more than two years earlier. But at the time I didn’t have a conceptual framework that made me pay attention to it. And a small picture like that just didn’t have the same in-your-face “complexity from nothing” character as my larger picture of rule 30.

Of course, as is typical in the history of ideas, there’s more to the story. One of the key things that had originally let me start “scientifically investigating” cellular automata is that out of all the infinite number of possible constructible rules, I’d picked a modest number on which I could do exhaustive experiments. I’d started by considering only “elementary” cellular automata, in one dimension, with k = 2 colors, and with rules of range r = 1. There are 256 such “elementary rules”. But many of them had what seemed to me “distracting” features—like backgrounds alternating between black and white on successive steps, or patterns that systematically shifted to the left or right. And to get rid of these “distractions” I decided to focus on what I (somewhat foolishly in retrospect) called “legal rules”: the 32 rules that leave blank states blank, and are left-right symmetric.

When one uses random initial conditions, the legal rules do seem—at least in small pictures—to capture the most obvious behaviors one sees across all the elementary rules. But it turns out that’s not true when one looks at simple initial conditions. Among the “legal” rules, the most complicated behavior one sees with simple initial conditions is nesting.

But even though I concentrated on “legal” rules, I still included in my first major paper on cellular automata pictures of a few “illegal” rules starting from simple initial conditions—including rule 30. And what’s more, in a section entitled “Extensions”, I discussed cellular automata with more than 2 colors, and showed—though without comment—the pictures:

Click to enlarge

These were low-resolution pictures, and I think I imagined that if one ran them further, the behavior would somehow resolve into something simple. But by early 1983, I had some clues that this wouldn’t happen. Because by then I was generating fairly high-resolution pictures—including ones of the k = 2, r = 2 totalistic rule with code 10 starting from a simple initial condition:

Click to enlarge

In early drafts of my 1983 paper on “Universality and Complexity in Cellular Automata” I noted the generation of “irregularity”, and speculated that it might be associated with class 4 behavior. But later I just stated as an observation without “cause” that some rules—like code 10—generate “irregular patterns”. I elaborated a little, but in a very “statistical mechanics” kind of way, not getting the main point:

Click to enlarge

In September 1983 I did a little better:

Click to enlarge

But in the end it wasn’t until June 1, 1984, that I really grokked what was going on. And a little over a week later I was in a scenic area of northern Sweden

Click to enlarge

at a posh “Nobel Symposium” conference on “The Physics of Chaos and Related Problems”—talking for the first time about the phenomenon I’d seen in rule 30 and code 10. And from June 15 there’s a transcript of a discussion session where I bring up the never-before-mentioned-in-public concept of computational irreducibility—and, unsurprisingly, leave the other participants (who were basically all traditional mathematically oriented physicists) at best slightly bemused:

Click to enlarge

I think I was still a bit prejudiced against rule 30 and code 10 as specific rules: I didn’t like the asymmetry of rule 30, and I didn’t like the rapid growth of code 10. (Rule 73—while symmetric—I also didn’t like because of its alternating background.) But having now grokked the rule 30 phenomenon I knew it also happened in “more aesthetic” “legal” rules with more than 2 colors. And while even 3 colors led to a rather large total space of rules, it was easy to generate examples of the phenomenon there.

A few days later I was back in the US, working on finishing my article for Scientific American. A photographer came to help get pictures from the color display I now had:

Click to enlarge

And, yes, those pictures included multicolor rules that showed the rule 30 phenomenon:

Click to enlarge

The caption I wrote commented: “Even in this case the patterns generated can be complex, and they sometimes appear quite random. The complex patterns formed in such physical processes as the flow of a turbulent fluid may well arise from the same mechanism.”

The article went on to describe computational irreducibility and its implications in quite a lot of detail— illustrating it rather nicely with a diagram, and commenting that “It seems likely that many physical and mathematical systems for which no simple description is now known are in fact computationally irreducible”:

Click to enlarge

I also included an example—that would show up almost unchanged in A New Kind of Science nearly 20 years later—indicating how computational irreducibility could lead to undecidability (back in 1984 the picture was made by stitching together many screen photographs, yes, with strange artifacts from long-exposure photography of CRTs):

Click to enlarge

In a rather newspaper-production-like experience, I spent the evening of July 18 at the offices of Scientific American in New York City putting finishing touches to the article, which at the end of the night—with minutes to spare—was dispatched for final layout and printing.

But already by that time, I was talking about computational irreducibility and the rule 30 phenomenon all over the place. In July I finished “Twenty Problems in the Theory of Cellular Automata” for the proceedings of the Swedish conference, including what would become a rather standard kind of picture:

Click to enlarge

Problem 15 talks specifically about rule 30, and already asks exactly what would—35 years later—become Problem #2 in my 2019 Rule 30 Prizes

Click to enlarge

while Problem 18 asks the (still largely unresolved) question of what the ultimate frequency of computational irreducibility is:

Click to enlarge

Very late in putting together the Scientific American article I’d added to the caption of the picture showing rule-30-like behavior the statement “Complex patterns generated by cellular automata can also serve as a source of effectively random numbers, and they can be applied to encrypt messages by converting a text into an apparently random form.” I’d realized both that cellular automata could act as good random generators (we used rule 30 as the default in Wolfram Language for more than 25 years), and that their evolution could effectively encrypt things, much as I’d later describe the Second Law as being about “encrypting” initial conditions to produce effective irreversibility.

Back in 1984 it was a surprising claim that something as simple and “science-oriented” as a cellular automaton could be useful for encryption. Because at the time practical encryption was basically always done by what at least seemed like arbitrary and complicated engineering solutions, whose security relied on details or explanations that were often considered military or commercial secrets.

I’m not sure when I first became aware of cryptography. But back in 1973 when I first had access to a computer there were a couple of kids (as well as a teacher who’d been a friend of Alan Turing’s) who were programming Enigma-like encryption systems (perhaps fueled by what were then still officially just rumors of World War II goings-on at Bletchley Park). And by 1980 I knew enough about encryption that I made a point of encrypting the source code of SMP (using a modified version of the Unix crypt program). (As it happens, we lost the password, and it was only in 2015 that we got access to the source again.)

My archives record a curious interaction about encryption in May 1982—right around when I’d first run (though didn’t appreciate) rule 30. A rather colorful physicist I knew named Brosl Hasslacher (who we’ll encounter again later) was trying to start a curiously modern-sounding company named Quantum Encryption Devices (or QED for short)—that was actually trying to market a quite hacky and definitively classical (multiple-shift-register-based) encryption system, ultimately to some rather shady customers (and, yes, the “expected” funding did not materialize):

Click to enlarge

But it was 1984 before I made a connection between encryption and cellular automata. And the first thing I imagined was giving input as the initial condition of the cellular automaton, then running the cellular automaton rule to produce “encrypted output”. The most straightforward way to make encryption was then to have the cellular automaton rule be reversible, and to run the inverse rule to do the decryption. I’d already done a little bit of investigation of reversible rules, but this led to a big search for reversible rules—which would later come in handy for thinking about microscopically reversible processes and thermodynamics.

Just down the hall from me at the Institute for Advanced Study was a distinguished mathematician named John Milnor, who got very interested in what I was doing with cellular automata. My archives contain all sorts of notes from Jack, like:

Click to enlarge

There’s even a reversible (“one-to-one”) rule, with nice, minimal BASIC code, along with lots of “real math”:

Click to enlarge

But by the spring of 1984 Jack and I were talking a lot about encryption in cellular automata—and we even began to draft a paper about it

Click to enlarge

complete with outlines of how encryption schemes could work:

Click to enlarge

The core of our approach involved reversible rules, and so we did all sorts of searches to find these (and by 1984 Jack was—like me—writing C code):

Click to enlarge

I wondered how random the output from cellular automata was, and I asked people I knew at Bell Labs about randomness testing (and, yes, email headers haven’t changed much in four decades, though then I was swolf@ias.uucp; research!ken was Ken Thompson of Unix fame):

Click to enlarge

But then came my internalization of the rule 30 phenomenon, which led to a rather different way of thinking about encryption with cellular automata. Before, we’d basically been assuming that the cellular automaton rule was the encryption key. But rule 30 suggested one could instead have a fixed rule, and have the initial condition define the key. And this is what led me to more physics-oriented thinking about cryptography—and to what I said in Scientific American.

In July I was making “encryption-friendly” pictures of rule 30:

Click to enlarge

But what Jack and I were most interested in was doing something more “cryptographically sophisticated”, and in particular inventing a practical public-key cryptosystem based on cellular automata. Pretty much the only public-key cryptosystems known then (or even now) are based on number theory. But we thought maybe one could use something like products of rules instead of products of numbers. Or maybe one didn’t need exact invertibility. Or something. But by the late summer of 1984, things weren’t looking good:

Click to enlarge

And eventually we decided we just couldn’t figure it out. And it’s basically still not been figured out (and maybe it’s actually impossible). But even though we don’t know how to make a public-key cryptosystem with cellular automata, the whole idea of encrypting initial data and turning it into effective randomness is a crucial part of the whole story of the computational foundations of thermodynamics as I think I now understand them.

Where Does Randomness Come From?

Right from when I first formulated it, I thought computational irreducibility was an important idea. And in the late summer of 1984 I decided I’d better write a paper specifically about it. The result was:

Click to enlarge

It was a pithy paper, arranged to fit in the 4-page limit of Physical Review Letters, with a rather clear description of computational irreducibility and its immediate implications (as well as the relation between physics and computation, which it footnoted as a “physical form of the Church–Turing thesis”). It illustrated computational reducibility and irreducibility in a single picture, here in its original Scotch-taped form:

Click to enlarge

The paper contains all sorts of interesting tidbits, like this run of footnotes:

Click to enlarge

In the paper itself I didn’t mention the Second Law, but in my archives I find some notes I made in preparing the paper, about candidate irreducible or undecidable problems (with many still unexplored)

Click to enlarge

which include “Will a hard sphere gas started from a particular state ever exhibit some specific anti-thermodynamic behaviour?”

In November 1984 the then-editor of Physics Today asked if I’d write something for them. I never did, but my archives include a summary of a possible article—which among other things promises to use computational ideas to explain “why the Second Law of thermodynamics holds so widely”:

Click to enlarge

So by November 1984 I was already aware of the connection between computational irreducibility and the Second Law (and also I didn’t believe that the Second Law would necessarily always hold). And my notes—perhaps from a little later—make it clear that actually I was thinking about the Second Law along pretty much the same lines as I do now, except that back then I didn’t yet understand the fundamental significance of the observer:

Click to enlarge

And spelunking now in my old filesystem (retrieved from a 9-track backup tape) I find from November 17, 1984 (at 2:42am), troff source for a putative paper (which, yes, we even now can run through troff):

Click to enlarge

This is all that’s in my filesystem. So, yes, in effect, I’m finally (more or less) finishing this 38 years later.

But in 1984 one of the hot—if not new—ideas of the time was “chaos theory”, which talked about how “randomness” could “deterministically arise” from progressive “excavation” of higher and higher-order digits in the initial conditions for a system. But having seen rule 30 this whole phenomenon of what was often (misleadingly) called “deterministic chaos” seemed to me at best like a sideshow—and definitely not the main effect leading to most randomness seen in physical systems.

I began to draft a paper about this

Click to enlarge

including for the first time an anchor picture of rule 30 intrinsically generating randomness—to be contrasted with pictures of randomness being generated (still in cellular automata) from sensitive dependence on random initial conditions:

Click to enlarge

It was a bit of a challenge to find an appropriate publishing venue for what amounted to a rather “interdisciplinary” piece of physics-meets-math-meets-computation. But Physical Review Letters seemed like the best bet, so on November 19, 1984, I submitted a version of the paper there, shortened to fit in its 4-page limit.

A couple of months later the journal said it was having trouble finding appropriate reviewers. I revised the paper a bit (in retrospect I think not improving it), then on February 1, 1985, sent it in again, with the new title “Origins of Randomness in Physical Systems”:

Click to enlarge

On March 8 the journal responded, with two reports from reviewers. One of the reviewers completely missed the point (yes, a risk in writing shift-the-paradigm papers). The other sent a very constructive two-page report:

Click to enlarge

I didn’t know it then, but later I found out that Bob Kraichnan had spent much of his life working on fluid turbulence (as well as that he was a very independent and think-for-oneself physicist who’d been one of Einstein’s last assistants at the Institute for Advanced Study). Looking at his report now it’s a little charming to see his statement that “no one who has looked much at turbulent flows can easily doubt [that they intrinsically generate randomness]” (as opposed to getting randomness from noise, initial conditions, etc.). Even decades later, very few people seem to understand this.

There were several exchanges with the journal, leaving it controversial whether they would publish the paper. But then in May I visited Los Alamos, and Bob Kraichnan invited me to lunch. He’d also invited a then-young physicist from Los Alamos who I’d known fairly well a few years earlier—and who’d once paid me the unintended compliment that it wasn’t fair for me to work on science because I was “too efficient”. (He told me he’d “intended to work on cellular automata”, but before he’d gotten around to it, I’d basically figured everything out.) Now he was riding the chaos theory bandwagon hard, and insofar as my paper threatened that, he wanted to do anything he could to kill the paper.

I hadn’t seen this kind of “paradigm attack” before. Back when I’d been doing particle physics, it had been a hot and cutthroat area, and I’d had papers plagiarized, sometimes even egregiously. But there wasn’t really any “paradigm divergence”. And cellular automata—being quite far from the fray—were something I could just peacefully work on, without anyone really paying much attention to whatever paradigm I might be developing.

At lunch I was treated to a lecture about why what I was doing was nonsense, or even if it wasn’t, I shouldn’t talk about it, at least now. Eventually I got a chance to respond, I thought rather effectively—causing my “opponent” to leave in a huff, with the parting line “If you publish the paper, I’ll ruin your career”. It was a strange thing to say, given that in the pecking order of physics, he was quite junior to me. (A decade and half later there were nevertheless a couple of “incidents”.) Bob Kraichnan turned to me, cracked a wry smile and said “OK, I’ll go right now and tell [the journal] to publish your paper”:

Click to enlarge

Kraichnan was quite right that the paper was much too short for what it was trying to say, and in the end it took a long book—namely A New Kind of Science—to explain things more clearly. But the paper was where a high-resolution picture of rule 30 first appeared in print. And it was the place where I first tried to explain the distinction between “randomness that’s just transcribed from elsewhere” and the fundamental phenomenon one sees in rule 30 where randomness is intrinsically generated by computational processes within a system.

I wanted words to describe these two different cases. And reaching back to my years of learning ancient Greek in school I invented the terms “homoplectic” and “autoplectic”, with the noun “autoplectism” to describe what rule 30 does. In retrospect, I think these terms are perhaps “too Greek” (or too “medical sounding”), and I’ve tended to just talk about “intrinsic randomness generation” instead of autoplectism. (Originally, I’d wanted to avoid the term “intrinsic” to prevent confusion with randomness that’s baked into the rules of a system.)

The paper (as Bob Kraichnan pointed out) talks about many things. And at the end, having talked about fluid turbulence, there’s a final sentence—about the Second Law:

Click to enlarge

In my archives, I find other mentions of the Second Law too. Like an April 1985 proto-paper that was never completed

Click to enlarge

but included the statement:

Click to enlarge

My main reason for working on cellular automata was to use them as idealized models for systems in nature, and as a window into foundational issues. But being quite involved in the computer industry, I couldn’t help wondering whether they might be directly useful for practical computation. And I talked about the possibility of building a “metachip” in which—instead of having predefined “meaningful” opcodes like in an ordinary microprocessor—everything would be built up “purely in software” from an underlying universal cellular automaton rule. And various people and companies started sending me possible designs:

Click to enlarge

But in 1984 I got involved in being a consultant to an MIT-spinoff startup called Thinking Machines Corporation that was trying to build a massively parallel “Connection Machine” computer with 65536 processors. The company had aspirations around AI (hence the name, which I’d actually been involved in suggesting), but their machine could also be put to work simulating cellular automata, like rule 30. In June 1985, hot off my work on the origins of randomness, I went to spend some of the summer at Thinking Machines, and decided it was time to do whatever analysis—or, as I would call it now, ruliology—I could on rule 30.

My filesystem from 1985 records that it was fast work. On June 24 I printed a somewhat-higher-resolution image of rule 30 (my login was “swolf” back then, so that’s how my printer output was labeled):

Click to enlarge

By July 2 a prototype Connection Machine had generated 2000 steps of rule 30 evolution:

Click to enlarge

With a large-format printer normally used to print integrated circuit layouts I got an even larger “piece of rule 30”—that I laid out on the floor for analysis, for example trying to measure (with meter rules, etc.) the slope of the border between regularity and irregularity in the pattern.

Richard Feynman was also a consultant at Thinking Machines, and we often timed our visits to coincide:

Click to enlarge

Feynman and I had talked about randomness quite a bit over the years, most recently in connection with the challenges of making a “quantum randomness chip” as a minimal example of quantum computing. Feynman at first didn’t believe that rule 30 could really be “producing randomness”, and that there must be some way to “crack” it. He tried, both by hand and with a computer, particularly using statistical mechanics methods to try to compute the slope of the border between regularity and irregularity:

Click to enlarge

But in the end, he gave up, telling me “OK, Wolfram, I think you’re on to something”.

Meanwhile, I was throwing all the methods I knew at rule 30. Combinatorics. Dynamical systems theory. Logic minimization. Statistical analysis. Computational complexity theory. Number theory. And I was pulling in all sorts of hardware and software too. The Connection Machine. A Cray supercomputer. A now-long-extinct Celerity C1200 (which successfully computed a length-40,114,679,273 repetition period). A LISP machine for graph layout. A circuit-design logic minimization program. As well as my own SMP system. (The Wolfram Language was still a few years in the future.)

But by July 21, there it was: a 50-page “ruliological profile” of rule 30, in a sense showing what one could of the “anatomy” of its randomness:

Click to enlarge

A month later I attended in quick succession a conference in California about cryptography, and one in Japan about fluid turbulence—with these two fields now firmly connected through what I’d discovered.

Hydrodynamics, and a Turbulent Tale

Back from when I first saw it at the age of 14 it was always my favorite page in The Feynman Lectures on Physics. But how did the phenomenon of turbulence that it showed happen, and what really was it?

Click to enlarge

In late 1984, the first version of the Connection Machine was nearing completion, and there was a question of what could be done with it. I agreed to analyze its potential uses in scientific computation, and in my resulting (never ultimately completed) report

Click to enlarge

the very first section was about fluid turbulence (others sections were about quantum field theory, n-body problems, number theory, etc.):

Click to enlarge

The traditional computational approach to studying fluids was to start from known continuum fluid equations, then to try to construct approximations to these suitable for numerical computation. But that wasn’t going to work well for the Connection Machine. Because in optimizing for parallelism, its individual processors were quite simple, and weren’t set up to do fast (e.g. floating-point) numerical computation.

I’d been saying for years that cellular automata should be relevant to fluid turbulence. And my recent study of the origins of randomness made me all the more convinced that they would for example be able to capture the fundamental randomness associated with turbulence (which I explained as being a bit like encryption):

Click to enlarge

I sent a letter to Feynman expressing my enthusiasm:

Click to enlarge

I had been invited to a conference in Japan that summer on “High Reynolds Number Flow Computation” (i.e. computing turbulent fluid flow), and on May 4 I sent an abstract which explained a little more of my approach:

Click to enlarge

My basic idea was to start not from continuum equations, but instead from a cellular automaton idealization of molecular dynamics. It was the same kind of underlying model as I’d tried to set up in my SPART program in 1973. But now instead of using it to study thermodynamic phenomena and the microscopic motions associated with heat, my idea was to use it to study the kind of visible motion that occurs in fluid dynamics—and in particular to see whether it could explain the apparent randomness of fluid turbulence.

I knew from the beginning that I needed to rely on “Second Law behavior” in the underlying cellular automaton—because that’s what would lead to the randomness necessary to “wash out” the simple idealizations I was using in the cellular automaton, and allow standard continuum fluid behavior to emerge. And so it was that I embarked on the project of understanding not only thermodynamics, but also hydrodynamics and fluid turbulence, with cellular automata—on the Connection Machine.

I’ve had the experience many times in my life of entering a field and bringing in new tools and new ideas. Back in 1985 I’d already done that several times, and it had always been a pretty much uniformly positive experience. But, sadly, with fluid turbulence, it was to be, at best, a turbulent experience.

The idea that cellular automata might be useful in studying fluid turbulence definitely wasn’t obvious. The year before, for example, at the Nobel Symposium conference in Sweden, a French physicist named Uriel Frisch had been summarizing the state of turbulence research. Fittingly for the topic of turbulence, he and I first met after a rather bumpy helicopter ride to a conference event—where Frisch told me in no uncertain terms that cellular automata would never be relevant to turbulence, and talked about how turbulence was better thought of as being associated (a bit like in the mathematical theory of phase transitions) with “singularities getting close to the real line”. (Strangely, I just now looked at Frisch’s paper in the proceedings of the conference: “Ou en est la Turbulence Developpée?” [roughly: “Fully Developed Turbulence: Where Do We Stand?”], and was surprised to discover that its last paragraph actually mentions cellular automata, and its acknowledgements thank me for conversations—even though the paper says it was received June 11, 1984, a couple of days before I had met Frisch. And, yes, this is the kind of thing that makes accurately reconstructing history hard.)

Los Alamos had always been a hotbed of computational fluid dynamics (not least because of its importance in simulating nuclear explosions)—and in fact of computing in general—and, starting in the late fall of 1984, on my visits there I talked to many people about using cellular automata to do fluid dynamics on the Connection Machine. Meanwhile, Brosl Hasslacher (mentioned above in connection with his 1982 encryption startup) had—after a rather itinerant career as a physicist—landed at Los Alamos. And in fact I had been asked by the Los Alamos management for a letter about him in December 1984 (yes, even though he was 18 years older than me), and ended what I wrote with: “He has considerable ability in identifying promising areas of research. I think he would be a significant addition to the staff at Los Alamos.”

Well, in early 1985 Brosl identified cellular automaton fluid dynamics as a promising area, and started energetically talking to me about it. Meanwhile, the Connection Machine was just starting to work, and a young software engineer named Jim Salem was assigned to help me get cellular automaton fluid dynamics running on it. I didn’t know it at the time, but Brosl—ever the opportunist—had also made contact with Uriel Frisch, and now I find the curious document in French dated May 10, 1985, with the translated title “A New Concept for Supercomputers: Cellular Automata”, laying out a grand international multiyear plan, and referencing the (so far as I know, nonexistent) B. Hasslacher and U. Frisch (1985), “The Cellular Automaton Turbulence Machine”, Los Alamos:

Click to enlarge

I visited Los Alamos again in May, but for much of the summer I was at Thinking Machines, and on July 18 Uriel Frisch came to visit there, along with a French physicist named Yves Pomeau, who had done some nice work in the 1970s on applying methods of traditional statistical mechanics to “lattice gases”.

But what about realistic fluid dynamics, and turbulence? I wasn’t sure how easy it would be to “build up from the (idealized) molecules” to get to pictures of recognizable fluid flows. But we were starting to have some success in generating at least basic results. It wasn’t clear how seriously anyone else was taking this (especially given that at the time I hadn’t seen the material Frisch had already written), but insofar as anything was “going on”, it seemed to be a perfectly collegial interaction—where perhaps Los Alamos or the French government or both would buy a Connection Machine computer. But meanwhile, on the technical side, it had become clear that the most obvious square-lattice model (that Pomeau had used in the 1970s, and that was basically what my SPART program from 1973 was supposed to implement) was fine for diffusion processes, but couldn’t really represent proper fluid flow.

When I first started working on cellular automata in 1981 the minimal 1D case in which I was most interested had barely been studied, but there had been quite a bit of work done in previous decades on the 2D case. By the 1980s, however, it had mostly petered out—with the exception of a group at MIT led by Ed Fredkin, who had long had the belief that one might in effect be able to “construct all of physics” using cellular automata. Tom Toffoli and Norm Margolus, who were working with him, had built a hardware 2D cellular automaton simulator—that I happened to photograph in 1982 when visiting Fredkin’s island in the Caribbean:

Click to enlarge

But while “all of physics” was elusive (and our Physics Project suggests that a cellular automaton with a rigid lattice is not the right place to start), there’d been success in making for example an idealized gas, using essentially a block cellular automaton on a square grid. But mostly the cellular automaton machine was used in a maddeningly “Look at this cool thing!” mode, often accompanied by rapid physical rewiring.

In early 1984 I visited MIT to use the machine to try to do what amounted to natural science, systematically studying 2D cellular automata. The result was a paper (with Norman Packard) on 2D cellular automata. We restricted ourselves to square grids, though mentioned hexagonal ones, and my article in Scientific American in late 1984 opened with a full-page hexagonal cellular automaton simulation of a snowflake made by Packard (and later in 1984 turned into one of a set of cellular automaton cards for sale):

Click to enlarge

In any case, in the summer of 1985, with square lattices not doing what was needed, it was time to try hexagonal ones. I think Yves Pomeau already had a theoretical argument for this, but as far as I was concerned, it was (at least at first) just a “next thing to try”. Programming the Connection Machine was at that time a rather laborious process (which, almost unprecedentedly for me, I wasn’t doing myself), and mapping a hexagonal grid onto its basically square architecture was a little fiddly, as my notes record:

Click to enlarge

Meanwhile, at Los Alamos, I’d introduced a young and very computer-savvy protege of mine named Tsutomu Shimomura (who had a habit of getting himself into computer security scrapes, though would later become famous for taking down a well-known hacker) to Brosl Hasslacher, and now Tsutomu jumped into writing optimized code to implement hexagonal cellular automata on a Cray supercomputer.

In my archives I now find a draft paper from September 7 that starts with a nice (if not entirely correct) discussion of what amounts to computational irreducibility, and then continues by giving theoretical symmetry-based arguments that a hexagonal cellular automaton should be able to reproduce fluid mechanics:

Click to enlarge

Click to enlarge

Near the end, the draft says (misspelling Tsutomu Shimomura’s name):

Click to enlarge

Meanwhile, we (as well as everyone else) were starting to get results that looked at least suggestive:

Click to enlarge

By November 15 I had drafted a paper

Click to enlarge

that included some more detailed pictures

Click to enlarge

and that at the end (I thought, graciously) thanked Frisch, Hasslacher, Pomeau and Shimomura for “discussions and for sharing their unpublished results with us”, which by that point included a bunch of suggestive, if not obviously correct, pictures of fluid-flow-like behavior.

To me, what was important about our paper is that, after all these years, it filled in with more detail just how computational systems like cellular automata could lead to Second-Law-style thermodynamic behavior, and it “proved” the physicality of what was going on by showing easy-to-recognize fluid-dynamics-like behavior.

Just four days later, though, there was a big surprise. The Washington Post ran a front-page story—alongside the day’s characteristic-Cold-War-era geopolitical news—about the “Hasslacher–Frisch model”, and about how it might be judged so important that it “should be classified to keep it out of Soviet hands”:

Click to enlarge

At that point, things went crazy. There was talk of Nobel Prizes (I wasn’t buying it). There were official complaints from the French embassy about French scientists not being adequately recognized. There was upset at Thinking Machines for not even being mentioned. And, yes, as the originator of the idea, I was miffed that nobody seemed to have even suggested contacting me—even if I did view the rather breathless and “geopolitical” tenor of the article as being pretty far from immediate reality.

At the time, everyone involved denied having been responsible for the appearance of the article. But years later it emerged that the source was a certain John Gage, former political operative and longtime marketing operative at Sun Microsystems, who I’d known since 1982, and had at some point introduced to Brosl Hasslacher. Apparently he’d called around various government contacts to help encourage open (international) sharing of scientific code, quoting this as a test case.

But as it was, the article had pretty much exactly the opposite effect, with everyone now out for themselves. In Princeton, I’d interacted with Steve Orszag, whose funding for his new (traditional) computational fluid dynamics company, Nektonics, now seemed at risk, and who pulled me into an emergency effort to prove that cellular automaton fluid dynamics couldn’t be competitive. (The paper he wrote about this seemed interesting, but I demurred on being a coauthor.) Meanwhile, Thinking Machines wanted to file a patent as quickly as possible. Any possibility of the French government getting a Connection Machine evaporated and soon Brosl Hasslacher was claiming that “the French are faking their data”.

And then there was the matter of the various academic papers. I had been sent the Frisch–Hasslacher–Pomeau paper to review, and checking my 1985 calendar for my whereabouts I must have received it the very day I finished my paper. I told the journal they should publish the paper, suggesting some changes to avoid naivete about computing and computer technology, but not mentioning its very thin recognition of my work.

Our paper, on the other hand, triggered a rather indecorous competitive response, with two “anonymous reviewers” claiming that the paper said nothing more than its “reference 5” (the Frisch–Hasslacher–Pomeau paper). I patiently pointed out that that wasn’t the case, not least because our paper had actual simulations, but also that actually I happened to have “been there first” with the overall idea. The journal solicited other opinions, which were mostly supportive. But in the end a certain Leo Kadanoff swooped in to block it, only to publish his own a few months later.

It felt corrupt, and distasteful. I was at that point a successful and increasingly established academic. And some of the people involved were even longtime friends. So was this kind of thing what I had to look forward to in a life in academia? That didn’t seem attractive, or necessary. And it was what began the process that led me, a year and a half later, to finally choose to leave academia behind, never to return.

Still, despite the “turbulence”—and in the midst of other activities—I continued to work hard on cellular automaton fluids, and by January 1986 I had the first version of a long (and, I thought, rather good) paper on their basic theory (that was finished and published later that year):

Click to enlarge

As it turns out, the methods I used in that paper provide some important seeds for our Physics Project, and even in recent times I’ve often found myself referring to the paper, complete with its SMP open-code appendix:

Click to enlarge

But in addition to developing the theory, I was also getting simulations done on the Connection Machine, and getting actual experimental data (particularly on flow past cylinders) to compare them to. By February 1986, we had quite a few results:

Click to enlarge

But by this point there was a quite industrial effort, particularly in France, that was churning out papers on cellular automaton fluids at a high rate. I’d called my theory paper “Cellular Automaton Fluid 1: Basic Theory”. But was it really worth finishing part 2? There was a veritable army of perfectly good physicists “competing” with me. And, I thought, “I have other things to do. Just let them do this. This doesn’t need me”.

And so it was that in the middle of 1986 I stopped working on cellular automaton fluids. And, yes, that freed me up to work on lots of other interesting things. But even though methods derived from cellular automaton fluids have become widely used in practical fluid dynamics computations, the key basic science that I thought could be addressed with cellular automaton fluids—about things like the origin of randomness in turbulence—has still, even to this day, not really been further explored.

Getting to the Continuum

In June 1986 I was about to launch both a research center (the Center for Complex Systems Research at the University of Illinois) and a journal (Complex Systems)—and I was also organizing a conference called CA ’86 (which was held at MIT). The core of the conference was poster presentations, and a few days before the conference was to start I decided I should find a “nice little project” that I could quickly turn into a poster.

In studying cellular automaton fluids I had found that cellular automata with rules based on idealized physical molecular dynamics could on a large scale approximate the continuum behavior of fluids. But what if one just started from continuum behavior? Could one derive underlying rules that would reproduce it? Or perhaps even find the minimal such rules?

By mid-1985 I felt I’d made decent progress on the science of cellular automata. But what about their engineering? What about constructing cellular automata with particular behavior? In May 1985 I had given a conference talk about “Cellular Automaton Engineering”, which turned into a paper about “Approaches to Complexity Engineering”—that in effect tried to set up “trainable cellular automata” in what might still be a powerful simple-programs-meet-machine-learning scheme that deserves to be explored:

Click to enlarge

But so it was that a few days before the CA ’86 conference I decided to try to find a minimal “cellular automaton approximation” to a simple continuum process: diffusion in one dimension.

I explained

Click to enlarge

and described as my objective:

Click to enlarge

I used block cellular automata, and tried to find rules that were reversible and also conserved something that could serve as “microscopic density” or “particle number”. I quickly determined that there were no such rules with 2 colors and blocks of sizes 2 or 3 that achieved any kind of randomization.

To go to 3 colors, I used SMP to generate candidate rules

Click to enlarge

where for example the function Apper can be literally be translated into Wolfram Language as

or, more idiomatically, just

then did what I have done so many times and just printed out pictures of their behavior:

Click to enlarge

Some clearly did not show randomization, but a couple did. And soon I was studying what I called the “winning rule”, which—like rule 30—went from simple initial conditions to apparent randomness:

Click to enlarge

I analyzed what the rule was “microscopically doing”

Click to enlarge

and explored its longer-time behavior:

Click to enlarge

Then I did things like analyze its cycle structure in a finite-size region by running C programs I’d basically already developed back in 1982 (though now they were modified to automatically generate troff code for typesetting):

Click to enlarge

And, like rule 30, the “winning rule” that I found back in June 1986 has stayed with me, essentially as a minimal example of reversible, number-conserving randomness. It appeared in A New Kind of Science, and it appears now in my recent work on the Second Law—and, of course, the patterns it makes are always the same:

Click to enlarge

Back in 1986 I wanted to know just how efficiently a simple rule like this could reproduce continuum behavior. And in a portent of observer theory my notes from the time talk about “optimal coarse graining, where the 2nd law is ‘most true’”, then go on to compare the distributed character of the cellular automaton with traditional “collect information into numerical value” finite-difference approximations:

Click to enlarge

In a talk I gave I summarized my understanding:

Click to enlarge

The phenomenon of randomization is generic in computational systems (witness rule 30, the “winning rule”, etc.) This leads to the genericity of thermodynamics. And this in turn leads to the genericity of continuum behavior, with diffusion and fluid behavior being two examples.

It would take another 34 years, but these basic ideas would eventually be what underlies our Physics Project, and our understanding of the emergence of things like spacetime. As well as now being crucial to our whole understanding of the Second Law.

The Second Law in A New Kind of Science

By the end of 1986 I had begun the development of Mathematica, and what would become the Wolfram Language, and for most of the next five years I was submerged in technology development. But in 1991 I started to use the technology I now had, and began the project that became A New Kind of Science.

Much of the first couple of years was spent exploring the computational universe of simple programs, and discovering that the phenomena I’d discovered in cellular automata were actually much more general. And it was seeing that generality that led me to the Principle of Computational Equivalence. In formulating the concept of computational irreducibility I’d in effect been thinking about trying to “reduce” the behavior of systems using an external as-powerful-as-possible universal computer. But now I’d realized I should just be thinking about all systems as somehow computationally equivalent. And in doing that I was pulling the conception of the “observer” and their computational ability closer to the systems they were observing.

But the further development of that idea would have to wait nearly three more decades, until the arrival of our Physics Project. In A New Kind of Science, Chapter 7 on “Mechanisms in Programs and Nature” describes the concept of intrinsic randomness generation, and how it’s distinguished from other sources of randomness. Chapter 8 on “Implications for Everyday Systems” then has a section on fluid flow, where I describe the idea that randomness in turbulence could be intrinsically generated, making it, for example, repeatable, rather than inevitably different every time an experiment is run.

And then there’s Chapter 9, entitled “Fundamental Physics”. The majority of the chapter—and its “most famous” part—is the presentation of the direct precursor to our Physics Project, including the concept of graph-rewriting-based computational models for the lowest-level structure of spacetime and the universe.

But there’s an earlier part of Chapter 9 as well, and it’s about the Second Law. There’s a precursor about “The Notion of Reversibility”, and then we’re on to a section about “Irreversibility and the Second Law of Thermodynamics”, followed by “Conserved Quantities and Continuum Phenomena”, which is where the “winning rule” I discovered in 1996 appears again:

Click to enlarge

My records show I wrote all of this—and generated all the pictures—between May 2 and July 11, 1995. I felt I already had a pretty good grasp of how the Second Law worked, and just needed to write it down. My emphasis was on explaining how a microscopically reversible rule—through its intrinsic ability to generate randomness—could lead to what appears to be irreversible behavior.

Mostly I used reversible 1D cellular automata as my examples, showing for example randomization both forwards and backwards in time:

Click to enlarge

I soon got to the nub of the issue with irreversibility and the Second Law:

Click to enlarge

I talked about how “typical textbook thermodynamics” involves a bunch of details about energy and motion, and to get closer to this I showed a simple example of an “ideal gas” 2D cellular automaton:

Click to enlarge

But despite my early exposure to hard-sphere gases, I never went as far as to use them as examples in A New Kind of Science. We did actually take some photographs of the mechanics of real-life billiards:

Click to enlarge

But cellular automata always seemed like a much clearer way to understand what was going on, free from issues like numerical precision, or their physical analogs. And by looking at cellular automata I felt as if I could really see down the foundations of the Second Law, and why it was true.

And mostly it was a story of computational irreducibility, and intrinsic randomness generation. But then there was rule 37R. I’ve often said that in studying the computational universe we have to remember that the “computational animals” are at least as smart as we are—and they’re always up to tricks we don’t expect.

And so it is with rule 37R. In 1986 I’d published a book of cellular automaton papers, and as an appendix I’d included lots of tables of properties of cellular automata. Almost all the tables were about the ordinary elementary cellular automata. But as a kind of “throwaway” at the very end I gave a table of the behavior of the 256 second-order reversible versions of the elementary rules, including 37R starting both from completely random initial conditions

Click to enlarge

and from single black cells:

Click to enlarge

So far, nothing remarkable. And years go by. But then—apparently in the middle of working on the 2D systems section of A New Kind of Science—at 4:38am on February 21, 1994 (according to my filesystem records), I generate pictures of all the reversible elementary rules again, but now from initial conditions that are slightly more complicated than a single black cell. Opening the notebook from that time (and, yes, Wolfram Language and our notebook format have been stable enough that 28 years later that still works) it shows up tiny on a modern screen, but there it is: rule 37R doing something “interesting”:

Click to enlarge

Clearly I noticed it. Because by 4:47am I’ve generated lots of pictures of rule 37R, like this one evolving from a block of 21 black cells, and showing only every other step

Click to enlarge

and by 4:54am I’ve got things like:

Click to enlarge

My guess is that I was looking for class 4 behavior in reversible cellular automata. And with rule 37R I’d found it. And at the time I moved on to other things. (On March 1, 1994, I slipped on some ice and broke my ankle, and was largely out of action for several weeks.)

And that takes us back to May 1995, when I was working on writing about the Second Law. My filesystem records that I did quite a few more experiments on rule 37R then, looking at different initial conditions, and running it as long as I could, to see if its strange neither-simple-nor-randomizing—and not very Second-Law-like—behavior would somehow “resolve”.

Up to that moment, for nearly a quarter of a century, I had always fundamentally believed in the Second Law. Yes, I thought there might be exceptions with things like self-gravitating systems. But I’d always assumed that—perhaps with some pathological exceptions—the Second Law was something quite universal, whose origins I could even now understand through computational irreducibility.

But seeing rule 37R this suddenly didn’t seem right. In A New Kind of Science I included a long run of rule 37R (here colorized to emphasize the structure)

Click to enlarge

then explained:

Click to enlarge

How could one describe what was happening in rule 37R? I discussed the idea that it was effectively forming “membranes” which could slowly move, but keep things “modular” and organized inside. I summarized at the time, tagging it as “something I wanted to explore in more detail one day”:

Click to enlarge

Rounding out the rest of A New Kind of Science takes another seven years of intense work. But finally in May 2002 it was published. The book talked about many things. And even within Chapter 9 my discussion of the Second Law was overshadowed by the outline I gave of an approach to finding a truly fundamental theory of physics—and of the ideas that evolved into our Physics Project.

The Physics Project—and the Second Law Again

After A New Kind of Science was finished I spent many years working mainly on technology—building Wolfram|Alpha, launching the Wolfram Language and so on. But “follow up on Chapter 9” was always on my longterm to-do list. The biggest—and most difficult—part of that had to do with fundamental physics. But I still had a great intellectual attachment to the Second Law, and I always wanted to use what I’d then understood about the computational paradigm to “tighten up” and “round out” the Second Law.

I’d mention it to people from time to time. Usually the response was the same: “Wasn’t the Second Law understood a century ago? What more is there to say?” Then I’d explain, and it’d be like “Oh, yes, that is interesting”. But somehow it always seemed like people felt the Second Law was “old news”, and that whatever I might do would just be “dotting an i or crossing a t”. And in the end my Second Law project never quite made it onto my active list, despite the fact that it was something I always wanted to do.

Occasionally I would write about my ideas for finding a fundamental theory of physics. And, implicitly I’d rely on the understanding I’d developed of the foundations and generalization of the Second Law. In 2015, for example, celebrating the centenary of General Relativity, I wrote about what spacetime might really be like “underneath”

Click to enlarge

and how a perceived spacetime continuum might emerge from discrete underlying structure like fluid behavior emerges from molecular dynamics—in effect through the operation of a generalized Second Law:

Click to enlarge

It was 17 years after the publication of A New Kind of Science that (as I’ve described elsewhere) circumstances finally aligned to embark on what became our Physics Project. And after all those years, the idea of computational irreducibility—and its immediate implications for the Second Law—had come to seem so obvious to me (and to the young physicists with whom I worked) that they could just be taken for granted as conceptual building blocks in constructing the tower of ideas we needed.

One of the surprising and dramatic implications of our Physics Project is that General Relativity and quantum mechanics are in a sense both manifestations of the same fundamental phenomenon—but played out respectively in physical space and in branchial space. But what really is this phenomenon?

What became clear is that ultimately it’s all about the interplay between underlying computational irreducibility and our nature as observers. It’s a concept that had its origins in my thinking about the Second Law. Because even in 1984 I’d understood that the Second Law is about our inability to “decode” underlying computationally irreducible behavior.

In A New Kind of Science I’d devoted Chapter 10 to “Processes of Perception and Analysis”, and I’d recognized that we should view such processes—like any processes in nature or elsewhere—as being fundamentally computational. But I still thought of processes of perception and analysis as being separated from—and in some sense “outside”—actual processes we might be studying. But in our Physics Project we’re studying the whole universe, so inevitably we as observers are “inside” and part of the system.

And what then became clear is the emergence of things like General Relativity and quantum mechanics depends on certain characteristics of us as observers. “Alien observers” might perceive quite different laws of physics (or no systematic laws at all). But for “observers like us”, who are computationally bounded and believe we are persistent in time, General Relativity and quantum mechanics are inevitable.

In a sense, therefore, General Relativity and quantum mechanics become “abstractly derivable” given our nature as observers. And the remarkable thing is that at some level the story is exactly the same with the Second Law. To me it’s a surprising and deeply beautiful scientific unification: that all three of the great foundational theories of physics—General Relativity, quantum mechanics and statistical mechanics—are in effect manifestations of the same core phenomenon: an interplay between computational irreducibility and our nature as observers.

Back in the 1970s I had no inkling of all this. And even when I chose to combine my discussions of the Second Law and of my approach to a fundamental theory of physics into a single chapter of A New Kind of Science, I didn’t know how deeply these would be connected. It’s been a long and winding path, that’s needed to pass through many different pieces of science and technology. But in the end the feeling I had when I first studied that book cover when I was 12 years old that “this was something fundamental” has played out on a scale almost incomprehensibly beyond what I had ever imagined.

Click to enlarge

Discovering Class 4

Most of my journey with the Second Law has had to do with understanding origins of randomness, and their relation to “typical Second-Law behavior”. But there’s another piece—still incompletely worked out—which has to do with surprises like rule 37R, and, more generally, with large-scale versions of class 4 behavior, or what I’ve begun to call the “mechanoidal phase”.

I first identified class 4 behavior as part of my systematic exploration of 1D cellular automata at the beginning of 1983—with the “code 20” k = 2, r = 2 totalistic rule being my first clear example:

Click to enlarge

Very soon my searches had identified a whole variety of localized structures in this rule:

Click to enlarge

Click to enlarge

At the time, the most significant attribute of class 4 cellular automata as far as I was concerned was that they seemed likely to be computation universal—and potentially provably so. But from the beginning I was also interested in what their “thermodynamics” might be. If you start them off from random initial conditions, will their patterns die out, or will some arrangement of localized structures persist, and perhaps even grow?

In most cellular automata—and indeed most systems with local rules—one expects that at least their statistical properties will somehow stabilize when one goes to the limit of infinite size. But, I asked, does that infinite-size limit even “exist” for class 4 systems—or if you progressively increase the size, will the results you get keep on jumping around forever, perhaps as you succeed in sampling progressively more exotic structures?

Click to enlarge

A paper I wrote in September 1983 talks about the idea that in a sufficiently large class 4 cellular automaton one would eventually get self-reproducing structures, which would end up “taking over everything”:

Click to enlarge

The idea that one might be able to see “biology-like” self-reproduction in cellular automata has a long history. Indeed, one of the multiple ways that cellular automata were invented (and the one that led to their name) was through John von Neumann’s 1952 effort to construct a complicated cellular automaton in which there could be a complicated configuration capable of self-reproduction.

But could self-reproducing structures ever “occur naturally” in cellular automata? Without the benefit of intuition from things like rule 30, von Neumann assumed that something like self-reproduction would need an incredibly complicated setup, as it seems to have, for example, in biology. But having seen rule 30—and more so class 4 cellular automata—it didn’t seem so implausible to me that even with very simple underlying rules, there could be fairly simple configurations that would show phenomena like self-reproduction.

But for such a configuration to “occur naturally” in a random initial condition might require a system with exponentially many cells. And I wondered if in the oceans of the early Earth there might have been only “just enough” molecules for something like a self-reproducing lifeform to occur.

Back in 1983 I already had pretty efficient code for searching for structures in class 4 cellular automata. But even running for days at a time, I never found anything more complicated than purely periodic (if sometimes moving) structures. And in March 1985, following an article about my work in Scientific American, I appealed to the public to find “interesting structures”—like “glider guns” that would “shoot out” moving structures:

Click to enlarge

As it happened, right before I made my “public appeal”, a student at Princeton working with a professor I knew had sent me a glider gun he’d found the k = 2, r = 3 totalistic code 88 rule:

Click to enlarge

At the time, though, with computer displays only large enough to see behavior like

I wasn’t convinced this was an “ordinary class 4 rule”—even though now, with the benefit of higher display resolution, it seems more convincing:

The “public appeal” generated a lot of interesting feedback—but no glider guns or other exotic structures in the rules I considered “obviously class 4”. And it wasn’t until after I started working on A New Kind of Science that I got back to the question. But then, on the evening of December 31, 1991, using exactly the same code as in 1983, but now with faster computers, there it was: in an ordinary class 4 rule (k = 3, r = 1 code 1329), after finding several localized structures, there was one that grew without bound (albeit not in the most obvious “glider gun” way):

Click to enlarge

But that wasn’t all. Exemplifying the principle that in the computational universe there are always surprises, searching a little further revealed yet other unexpected structures:

Click to enlarge

Every few years something else would come up with class 4 rules. In 1994, lots of work on rule 110. In 1995, the surprise of rule 37R. In 1998 efforts to find analogs of particles that might carry over to my graph-based model of space.

After A New Kind of Science was published in 2002, we started our annual Wolfram Summer School (at first called the NKS Summer School)—and in 2010 our High School Summer Camp. Some years we asked students to pick their “favorite cellular automaton”. Often they were class 4:

Click to enlarge

And occasionally someone would do a project to explore the world of some particular class 4 rule. But beyond those specifics—and statements about computation universality—it’s never been clear quite what one could say about class 4.

Back in 1984 in the series of cellular automaton postcards I’d produced, there were a couple of class 4 examples:

Click to enlarge

And even then the typical response to these images was that they looked “organic”—like the kind of thing living organisms might produce. A decade later—for A New Kind of Science—I studied “organic forms” quite a bit, trying to understand how organisms get their overall shapes, and surface patterns. Mostly that didn’t end up being a story of class 4 behavior, though.

Since the early 1980s I’ve been interested in molecular computing, and in how computation might be done at the level of molecules. My discoveries in A New Kind of Science (and specifically the Principle of Computational Equivalence) convinced me that it should be possible to get even fairly simple collections of molecules to “do arbitrary computations” or even build more or less arbitrary structures (in a more general and streamlined way than happens with the whole protein synthesis structure in biology). And over the years, I sometimes thought about trying to do practical work in this area. But it didn’t feel as if the ambient technology was quite ready. So I never jumped in.

Meanwhile, I’d long understood the basic correspondence between multiway systems and patterns of possible pathways for chemical reactions. And after our Physics Project was announced in 2020 and we began to develop the general multicomputational paradigm, I immediately considered molecular computing a potential application. But just what might the “choreography” of molecules be like? What causal relationships might there be, for example, between different interactions of the same molecule? That’s not something ordinary chemistry—dealing for example with liquid-phase reactions—tends to consider important.

But what I increasingly started to wonder is whether in molecular biology it might actually be crucial. And even in the 20 years since A New Kind of Science was published, it’s become increasingly clear that in molecular biology things are extremely “orchestrated”. It’s not about molecules randomly moving around, like in a liquid. It’s about molecules being carefully channeled and actively transported from one “event” to another.

Class 3 cellular automata seem to be good “metamodels” for things like liquids, and readily give Second-Law-like behavior. But what about the kind of situation that seems to exist in molecular biology? It’s something I’ve been thinking about only recently, but I think this is a place where class 4 cellular automata can contribute. I’ve started calling the “bulk limit” of class 4 systems the “mechanoidal phase”. It’s a place where the ordinary Second Law doesn’t seem to apply.

Four decades ago when I was trying to understand how structure could arise “in violation of the Second Law” I didn’t yet even know about computational irreducibility. But now we’ve come a lot further, in particular with the development of the multicomputational paradigm, and the recognition of the importance of the characteristics of the observer in defining what perceived overall laws there will be. It’s an inevitable feature of computational irreducibility that there will always be an infinite sequence of new challenges for science, and new pieces of computational reducibility to be found. So, now, yes, a challenge is to understand the mechanoidal phase. And with all the tools and ideas we’ve developed, I’m hoping the process will happen more than it has for the ordinary Second Law.

The End of a 50-Year Journey

I began my quest to understand the Second Law a bit more than 50 years ago. And—even though there’s certainly more to say and figure out—it’s very satisfying now to be able to bring a certain amount of closure to what has been the single longest-running piece of intellectual “unfinished business” in my life. It’s been an interesting journey—that’s very much relied on, and at times helped drive, the tower of science and technology that I’ve spent my life building. There are many things that might not have happened as they did. And in the end it’s been a story of longterm intellectual tenacity—stretching across much of my life so far.

For a long time I’ve kept (automatically when possible) quite extensive archives. And now these archives allow one to reconstruct in almost unprecedented detail my journey with the Second Law. One sees the gradual formation of intellectual frameworks over the course of years, then the occasional discovery or realization that allows one to take the next step in what is sometimes mere days. There’s a curious interweaving of computational and essentially philosophical methodologies—with an occasional dash of mathematics.

Sometimes there’s general intuition that’s significantly ahead of specific results. But more often there’s a surprise computational discovery that seeds the development of new intuition. And, yes, it’s a little embarrassing how often I managed to generate in a computer experiment something that I completely failed to interpret or even notice at first because I didn’t have the right intellectual framework or intuition.

And in the end, there’s an air of computational irreducibility to the whole process: there really wasn’t a way to shortcut the intellectual development; one just had to live it. Already in the 1990s I had taken things a fair distance, and I had even written a little about what I had figured out. But for years it hung out there as one of a small collection of unfinished projects: to finally round out the intellectual story of the Second Law, and to write down an exposition of it. But the arrival of our Physics Project just over two years ago brought both a cascade of new ideas, and for me personally a sense that even things that had been out there for a long time could in fact be brought to closure.

And so it is that I’ve returned to the quest I began when I was 12 years old—but now with five decades of new tools and new ideas. The wonder and magic of the Second Law is still there. But now I’m able to see it in a much broader context, and to realize that it’s not just a law about thermodynamics and heat, but instead a window into a very general computational phenomenon. None of this I could know when I was 12 years old. But somehow the quest I was drawn to all those years ago has turned out to be deeply aligned with the whole arc of intellectual development that I have followed in my life. And no doubt it’s no coincidence.

But for now I’m just grateful to have had the quest to understand Second Law as one of my guiding forces through so much of my life, and now to realize that my quest was part of something so broad and so deep.

Appendix: The Backstory of the Book Cover That Started It All

Click to enlarge

What is the backstory of the book cover that launched my long journey with the Second Law? The book was published in 1965, and inside its front flap we find:

Click to enlarge

On page 7 we then find:

Click to enlarge

In 2001—as I was putting the finishing touches to the historical notes for A New Kind of Science—I tracked down Berni Alder (who died in 2020 at the age of 94) to ask him the origin of the pictures. It turned out to be a complex story, reaching back to the earliest serious uses of computers for basic science, and even beyond.

The book had been born out the sense of urgency around science education in the US that followed the launch of Sputnik by the Soviet Union—with a group of professors from Berkeley and Harvard believing that the teaching of freshman college physics was in need of modernization, and that they should write a series of textbooks to enable this. (It was also the time of the “new math”, and a host of other STEM-related educational initiatives.) Fred Reif (who died at the age of 92 in 2019) was asked to write the statistical physics volume. As he explained in the preface to the book

Click to enlarge

ending with:

Click to enlarge

Well, it’s taken me 50 years to get to the point where I think I really understand the Second Law that is at the center of the book. And in 2001 I was able to tell Fred Reif that, yes, his book had indeed been useful. He said he was pleased to learn that, adding “It is all too rare that one’s educational efforts seem to bear some fruit.”

He explained to me that when he was writing the book he thought that “the basic ideas of irreversibility and fluctuations might be very vividly illustrated by the behavior of a gas of particles spreading through a box”. He added: “It then occurred to me that Berni Alder might actually show this by a computer generated film since he had worked on molecular dynamics simulations and had also good computer facilities available to him. I was able to enlist Berni’s interest in this project, with the results shown in my book.”

The acknowledgements in the book report:

Click to enlarge

Berni Alder and and Fred Reif did indeed create a “film loop”, which “could be bought separately from the book and viewed in the physics lab”, as Alder told me, adding that “I understand the students liked it very much, but the venture was not a commercial success.” Still, he sent me a copy of a videotape version:

Click to enlarge

The film (which has no sound) begins:

Click to enlarge

Soon it’s showing an actual process of “coming to equilibrium”:

“However”, as Alder explained it to me, “if a large number of particles are put in the corner and the velocities of all the particles are reversed after a certain time, the audience laughs or is supposed to after all the particles return to their original positions.” (One suspects that particularly in the 1960s this might have been reminiscent of various cartoon-film gags.)

OK, so how were the pictures (and the film) made? It was done in 1964 at what’s now Lawrence Livermore Lab (that had been created in 1952 as a spinoff of the Berkeley Radiation Lab, which had initiated some key pieces for the Manhattan Project) on a computer called the LARC (“Livermore Advanced Research Computer”), first made in 1960, that was probably the most advanced scientific computer of the time. Alder explained to me, however: “We could not run the problem much longer than about 10 collision times with 64 bits [sic] arithmetic before the round-off error prevented the particles from returning.”

Why did they start the particles off in a somewhat random configuration? (The randomness, Alder told me, had been created by a middle-square random number generator.) Apparently if they’d been in a regular array—which would have made the whole process of randomization much easier to see—the roundoff errors would have been too obvious. (And it’s issues like this that made it so hard to recognize the rule 30 phenomenon in systems based on real numbers—and without the idea of just studying simple programs not tied to traditional equation-based formulations of physics.)

The actual code for the molecular dynamics simulation was written in assembler and run by Mary Ann Mansigh (Karlsen), who had a degree in math and chemistry and worked as a programmer at Livermore from 1955 until the 1980s, much of the time specifically with Alder. Here she is at the console of the LARC (yes, computers had built-in desks in those days):

Click to enlarge

The program that was used was called STEP, and the original version of it had actually been written (by a certain Norm Hardy, who ended up having a long Silicon Valley career) to run on a previous generation of computer. (A still-earlier program was called APE, for “Approach to Equilibrium”.) But it was only with the LARC—and STEP—that things were fast enough to run substantial simulations, at the rate of about 200,000 collisions per hour (the simulation for the book cover involved 40 particles and about 500 collisions). At the time of the book STEP used an n2 algorithm where all pairs of particles were tested for collisions; later a neighborhood-based linked list method was used.

The standard method of getting output from a computer back in 1964—and basically until the 1980s—was to print characters on paper. But the LARC could also drive an oscilloscope, and it was with this that the graphics for the book were created (capturing them from the oscilloscope screen with a Polaroid instant camera).

But why was Berni Alder studying molecular dynamics and “hard sphere gases” in the first place? Well, that’s another long story. But ultimately it was driven by the effort to develop a microscopic theory of liquids.

The notion that gases might consist of discrete molecules in motion had arisen in the 1700s (and even to some extent in antiquity), but it was only in the mid-1800s that serious development of the “kinetic theory” idea began. Pretty immediately it was clear how to derive the ideal gas law P V = R T for essentially non-interacting molecules. But what analog of this “equation of state” might apply to gases with significant interactions between molecules, or, for that matter, liquids? In 1873 Johannes Diderik van der Waals proposed, on essentially empirical grounds, the formula (P + a/V2)(Vb) = RT—where the parameter b represented “excluded volume” taken up by molecules, that were implicitly being viewed as hard spheres. But could such a formula be derived—like the ideal gas law—from a microscopic kinetic theory of molecules? At the time, nobody really knew how to start, and the problem languished for more than half a century.

(It’s worth pointing out, by the way, that the idea of modeling gases, as opposed to liquids, as collections of hard spheres was extensively pursued in the mid-1800s, notably by Maxwell and Boltzmann—though with their traditional mathematical analysis methods, they were limited to studying average properties of what amount to dilute gases.)

Meanwhile, there was increasing interest in the microscopic structure of liquids, particularly among chemists concerned for example with how chemical solutions might work. And at the end of the 1920s the technique of x-ray diffraction, which had originally been used to study the microscopic structure of crystals, was applied to liquids—allowing in particular the experimental determination of the radial distribution function (or pair correlation function) g(r), which gives the probability to find another molecule a distance r from a given one.

But how might this radial distribution function be computed? By the mid-1930s there were several proposals based on looking at the statistics of random assemblies of hard spheres:

Click to enlarge

Some tried to get results by mathematical methods; others did physical experiments with ball bearings and gelatin balls, getting at least rough agreement with actual experiments on liquids:

Click to enlarge

But then in 1939 a physical chemist named John Kirkwood gave an actual probabilistic derivation (using a variety of simplifying assumptions) that fairly closely reproduced the radial distribution function:

Click to enlarge

But what about just computing from first principles, on the basis of the mechanics of colliding molecules? Back in 1872 Ludwig Boltzmann had proposed a statistical equation (the “Boltzmann transport equation”) for the behavior of collections of molecules, that was based on the approximation of independent probabilities for individual molecules. By the 1940s the independence assumption had been overcome, but at the cost of introducing an infinite hierarchy of equations (the “BBGKY hierarchy”, where the “K” stood for Kirkwood). And although the full equations were intractable, approximations were suggested that—while themselves mathematically sophisticated—seemed as if they should, at least in principle, be applicable to liquids.

Meanwhile, in 1948, Berni Alder, fresh from a master’s degree in chemical engineering, and already interested in liquids, went to Caltech to work on a PhD with John Kirkwood—who suggested that he look at a couple of approximations to the BBGKY hierarchy for the case of hard spheres. This led to some nasty integro-differential equations which couldn’t be solved by analytical techniques. Caltech didn’t yet have a computer in the modern sense, but in 1949 they acquired an IBM 604 Electronic Calculating Punch, which could be wired to do calculations with input and output specified on punched cards—and it was on this machine that Alder got the calculations he needed done (the paper records that “[this] … was calculated … with the use of IBM equipment and the file of punched cards of sin(ut) employed in these laboratories for electron diffraction calculation”):

Click to enlarge

Our story now moves to Los Alamos, where in 1947 Stan Ulam had suggested the Monte Carlo method as a way to study neutron diffusion. In 1949 the method was implemented on the ENIAC computer. And in 1952 Los Alamos got its own MANIAC computer. Meanwhile, there was significant interest at Los Alamos in computing equations of state for matter, especially in extreme conditions such as those in a nuclear explosion. And by 1953 the idea had arisen of using the Monte Carlo method to do this.

The concept was to take a collection of hard spheres (or actually 2D disks), and move them randomly in a series of steps with the constraint that they could not overlap—then look at the statistics of the resulting “equilibrium” configurations. This was done on the MANIAC, with the resulting paper now giving “Monte Carlo results” for things like the radial distribution function:

Click to enlarge

Kirkwood and Alder had been continuing their BBGKY hierarchy work, now using more realistic Lennard-Jones forces between molecules. But by 1954 Alder was also using the Monte Carlo method, implementing it partly (rather painfully) on the IBM Electronic Calculating Punch, and partly on the Manchester Mark II computer in the UK (whose documentation had been written by Alan Turing):

Click to enlarge

In 1955 Alder started working full-time at Livermore, recruited by Edward Teller. Another Livermore recruit—fresh from a physics PhD—was Thomas Wainwright. And soon Alder and Wainwright came up with an alternative to the Monte Carlo method—that would eventually give the book cover pictures: just explicitly compute the dynamics of colliding hard spheres, with the expectation that after enough collisions the system would come to equilibrium and allow things like equations of state to be obtained.

In 1953 Livermore had obtained its first computer: a Remington Rand Univac I. And it was on this computer that Alder and Wainwright did a first proof of concept of their method, tracing 100 hard spheres with collisions computed at the rate of about 100 per hour. Then in 1955 Livermore got IBM 704 computers, which, with their hardware floating-point capabilities, were able to compute about 2000 collisions per hour.

Alder and Wainwright reported their first results at a statistical mechanics conference in Brussels in August 1956 (organized by Ilya Prigogine). The published version appeared in 1958:

Click to enlarge

It gives evidence—that they tagged as “provisional”—for the emergence of a Maxwell–Boltzmann velocity distribution “after the system reached equilibrium”

Click to enlarge

as well as things like the radial distribution function—and the equation of state:

Click to enlarge

It was notable that there seemed to be a discrepancy between the results for the equation of state computed by explicit molecular dynamics and by the Monte Carlo method. And what is more, there seemed to be evidence of some kind of discontinuous phase-transition-like behavior as the density of spheres changed (an effect which Kirkwood had predicted in 1949).

Given the small system sizes and short runtimes it was all a bit muddy. But by August 1957 Alder and Wainwright announced that they’d found a phase transition, presumably between a high-density phase where the spheres were packed together like in a crystalline solid, and a low-density phase, where they were able to more freely “wander around” like in a liquid or gas. Meanwhile, the group at Los Alamos had redone their Monte Carlo calculations, and they too now claimed a phase transition. Their papers were published back to back:

Click to enlarge

But at this point no actual pictures of molecular trajectories had yet been published, or, I believe, made. All there was were traditional plots of aggregated quantities. And in 1958, these plots made their first appearance in a textbook. Tucked into Appendix C of Elementary Statistical Physics by Berkeley physics professor Charles Kittel (who would later be chairman of the group developing the Berkeley Physics Course book series) were two rather confusing plots about the approach to the Maxwell–Boltzmann distribution taken from a pre-publication version of Alder and Wainwright’s paper:

Click to enlarge

Alder and Wainwright’s phase transition result had created enough of a stir that they were asked to write a Scientific American article about it. And in that article—entitled “Molecular Motions”, from October 1959—there were finally pictures of actual trajectories, with their caption explaining that the “paths of particles … appear as bright lines on the face of a cathode-ray tube hooked to a computer” (the paths are of the centers of the colliding disks):

Click to enlarge

A technical article published at the same time gave a diagram of the logic for the dynamical computation:

Click to enlarge

Then in 1960 Livermore (after various delays) took delivery of the LARC computer—arguably the first scientific supercomputer—which allowed molecular dynamics computations to be done perhaps 20 times times faster. A 1962 picture shows Berni Alder (left) and Thomas Wainwright (right) looking at outputs from the LARC with Mary Ann Mansigh (yes, in those days it was typical for male physicists to wear ties):

Click to enlarge

And in 1964, the pictures for the Statistical Physics book (and film loop) got made, with Mary Ann Mansigh painstakingly constructing images of disks on the oscilloscope display.

Work on molecular dynamics continued, though to do it required the most powerful computers, so for many years it was pretty much restricted to places like Livermore. And in 1967, Alder and Wainwright made another discovery about hard spheres. Even in their first paper about molecular dynamics they’d plotted the velocity autocorrelation function, and noted that it decayed roughly exponentially with time. But by 1967 they had much more precise data, and realized that there was a deviation from exponential decay: a definite “long-time tail”. And soon they had figured out that this power-law tail was basically the result of a continuum hydrodynamic effect (essentially a vortex) operating even on the scale of a few molecules. (And—though it didn’t occur to me at the time—this should have suggested that even with fairly small numbers of cells cellular automaton fluid simulations had a good chance of giving recognizable hydrodynamic results.)

It’s never been entirely easy to do molecular dynamics, even with hard spheres, not least because in standard computations one’s inevitably confronted with things like numerical roundoff errors. And no doubt this is why some of the obvious foundational questions about the Second Law weren’t really explored there, and why intrinsic randomness generation and the rule 30 phenomenon weren’t identified.

Incidentally, even before molecular dynamics emerged, there was already one computer study of what could potentially have been Second Law behavior. Visiting Los Alamos in the early 1950s Enrico Fermi had gotten interested in using computers for physics, and wondered what would happen if one simulated the motion of an array of masses with nonlinear springs between them. The results of running this on the MANIAC computer were reported in 1955 (after Fermi had died)

Click to enlarge

and it was noted that there wasn’t just exponential approach to equilibrium, but instead something more complicated (later connected to solitons). Strangely, though, instead of plotting actual particle trajectories, what were given were mode energies—but these still exhibited what, if it hadn’t been obscured by continuum issues, might have been recognized as something like the rule 30 phenomenon:

Click to enlarge

But I knew none of this history when I saw the Statistical Physics book cover in 1972. And indeed, for all I knew, it could have been a “standard statistical physics cover picture”. I didn’t know it was the first of its kind—and a leading-edge example of the use of computers for basic science, accessible only with the most powerful computers of the time. Of course, had I known those things, I probably wouldn’t have tried to reproduce the picture myself and I wouldn’t have had that early experience in trying to use a computer to do science. (Curiously enough, looking at the numbers now, I realize that the base speed of the LARC was only 20x the Elliott 903C, though with floating point, etc.—a factor that pales in comparison with the 500x speedup in computers in the 40 years since I started working on cellular automata.)

But now I know the history of that book cover, and where it came from. And what I only just discovered now is that actually there’s a bigger circle than I knew. Because the path from Berni Alder to that book cover to my work on cellular automaton fluids came full circle—when in 1988 Alder wrote a paper based on cellular automaton fluids (though through the vicissitudes of academic behavior I don’t think he knew these had been my idea—and now it’s too late to tell him his role in seeding them):

Click to enlarge

Notes & Thanks

There are many people who’ve contributed to the 50-year journey I’ve described here. Some I’ve already mentioned by name, but others not—including many who doubtless wouldn’t even be aware that they contributed. The longtime store clerk at Blackwell’s bookstore who in 1972 sold college physics books to a 12-year old without batting an eye. (I learned his name—Keith Clack—30 years later when he organized a book signing for A New Kind of Science at Blackwell’s.) John Helliwell and Lawrence Wickens who in 1977 invited me to give the first talk where I explicitly discussed the foundations of the Second Law. Douglas Abraham who in 1977 taught a course on mathematical statistical mechanics that I attended. Paul Davies who wrote a book on The Physics of Time Asymmetry that I read around that time. Rocky Kolb who in 1979 and 1980 worked with me on cosmology that used statistical mechanics. The students (including professors like Steve Frautschi and David Politzer) who attended my 1981 class at Caltech about “nonequilibrium statistical mechanics”. David Pines and Elliott Lieb who in 1983 were responsible for publishing my breakout paper on “Statistical Mechanics of Cellular Automata”. Charles Bennett (curiously, a student of Berni Alder’s) with whom in the early 1980s I discussed applying computation theory (notably the ideas of Greg Chaitin) to physics. Brian Hayes who commissioned my 1984 Scientific American article, and Peter Brown who edited it. Danny Hillis and Sheryl Handler who in 1984 got me involved with Thinking Machines. Jim Salem and Bruce Nemnich (Walker) who worked on fluid dynamics on the Connection Machine with me. Then—36 years later—Jonathan Gorard and Max Piskunov, who catalyzed the doing of our Physics Project.

In the last 50 years, there’ve been surprisingly few people with whom I’ve directly discussed the foundations of the Second Law. Perhaps one reason is that back when I was a “professional physicist” statistical mechanics as a whole wasn’t a prominent area. But, more important, as I’ve described elsewhere, for more than a century most physicists have effectively assumed that the foundations of the Second Law are a solved (or at least merely pedantic) problem.

Probably the single person with whom I had the most discussions about the foundations of the Second Law is Richard Feynman. But there are others with whom at one time or another I’ve discussed related issues, including: Bruce Boghosian, Richard Crandall, Roger Dashen, Mitchell Feigenbaum, Nigel Goldenfeld, Theodore Gray, Bill Hayes, Joel Lebowitz, David Levermore, Ed Lorenz, John Maddox, Roger Penrose, Ilya Prigogine, Rudy Rucker, David Ruelle, Rob Shaw, Yakov Sinai, Michael Trott, Léon van Hove and Larry Yaffe. (There are also many others with whom I’ve discussed general issues about origins of randomness.)

Finally, one technical note about the presentation here: in an effort to maintain a clearer timeline, I’ve typically shown the earliest drafts or preprint versions of papers that I have. Their final published versions (if indeed they were ever published) appeared anything from weeks to years later, sometimes with changes.

How Did We Get Here? The Tangled History of the Second Law of Thermodynamics

1 février 2023 à 03:14

How Did We Get Here? The Tangled History of the Second Law of Thermodynamics

The Basic Arc of the Story

As I’ve explained elsewhere, I think I now finally understand the Second Law of thermodynamics. But it’s a new understanding, and to get to it I’ve had to overcome a certain amount of conventional wisdom about the Second Law that I at least have long taken for granted. And to check myself I’ve been keen to know just where this conventional wisdom came from, how it’s been validated, and what might have made it go astray.

And from this I’ve been led into a rather detailed examination of the origins and history of thermodynamics. All in all, it’s a fascinating story, that both explains what’s been believed about thermodynamics, and provides some powerful examples of the complicated dynamics of the development and acceptance of ideas.

The basic concept of the Second Law was first formulated in the 1850s, and rather rapidly took on something close to its modern form. It began partly as an empirical law, and partly as something abstractly constructed on the basis of the idea of molecules, that nobody at the time knew for sure existed. But by the end of the 1800s, with the existence of molecules increasingly firmly established, the Second Law began to often be treated as an almost-mathematically-proven necessary law of physics. There were still mathematical loose ends, as well as issues such as its application to living systems and to systems involving gravity. But the almost-universal conventional wisdom became that the Second Law must always hold, and if it didn’t seem to in a particular case, then that must just be because there was something one didn’t yet understand about that case.

There was also a sense that regardless of its foundations, the Second Law was successfully used in practice. And indeed particularly in chemistry and engineering it’s often been in the background, justifying all the computations routinely done using entropy. But despite its ubiquitous appearance in textbooks, when it comes to foundational questions, there’s always been a certain air of mystery around the Second Law. Though after 150 years there’s typically an assumption that “somehow it must all have been worked out”. I myself have been interested in the Second Law now for a little more than 50 years, and over that time I’ve had a growing awareness that actually, no, it hasn’t all been worked out. Which is why, now, it’s wonderful to see the computational paradigm—and ideas from our Physics Project—after all these years be able to provide solid foundations for understanding the Second Law, as well as seeing its limitations.

And from the vantage point of the understanding we now have, we can go back and realize that there were precursors of it even from long ago. In some ways it’s all an inspiring tale—of how there were scientists with ideas ahead of their time, blocked only by the lack of a conceptual framework that would take another century to develop. But in other ways it’s also a cautionary tale, of how the forces of “conventional wisdom” can blind people to unanswered questions and—over a surprisingly long time—inhibit the development of new ideas.

But, first and foremost, the story of the Second Law is the story of a great intellectual achievement of the mid-19th century. It’s exciting now, of course, to be able to use the latest 21st-century ideas to take another step. But to appreciate how this fits in with what’s already known we have to go back and study the history of what originally led to the Second Law, and how what emerged as conventional wisdom about it took shape.

What Is Heat?

Once it became clear what heat is, it actually didn’t take long for the Second Law to be formulated. But for centuries—and indeed until the mid-1800s—there was all sorts of confusion about the nature of heat.

That there’s a distinction between hot and cold is a matter of basic human perception. And seeing fire one might imagine it as a disembodied form of heat. In ancient Greek times Heraclitus (~500 BC) talked about everything somehow being “made of fire”, and also somehow being intrinsically “in motion”. Democritus (~460–~370 BC) and the Epicureans had the important idea (that also arose independently in other cultures) that everything might be made of large numbers of a few types of tiny discrete atoms. They imagined these atoms moving around in the “void” of space. And when it came to heat, they seem to have correctly associated it with the motion of atoms—though they imagined it came from particular spherical “fire” atoms that could slide more quickly between other atoms, and they also thought that souls were the ultimate sources of motion and heat (at least in warm-blooded animals?), and were made of fire atoms.

And for two thousand years that’s pretty much where things stood. And indeed in 1623 Galileo (1564–1642) (in his book The Assayer, about weighing competing world theories) was still saying:

Those materials which produce heat in us and make us feel warmth, which are known by the general name of “fire,” would then be a multitude of minute particles having certain shapes and moving with certain velocities. Meeting with our bodies, they penetrate by means of their extreme subtlety, and their touch as felt by us when they pass through our substance is the sensation we call “heat.”

He goes on:

Since the presence of fire-corpuscles alone does not suffice to excite heat, but their motion is needed also, it seems to me that one may very reasonably say that motion is the cause of heat… But I hold it to be silly to accept that proposition in the ordinary way, as if a stone or piece of iron or a stick must heat up when moved. The rubbing together and friction of two hard bodies, either by resolving their parts into very subtle flying particles or by opening an exit for the tiny fire-corpuscles within, ultimately sets these in motion; and when they meet our bodies and penetrate them, our conscious mind feels those pleasant or unpleasant sensations which we have named heat…

And although he can tell there’s something different about it, he thinks of heat as effectively being associated with a substance or material:

The tenuous material which produces heat is even more subtle than that which causes odor, for the latter cannot leak through a glass container, whereas the material of heat makes its way through any substance.

In 1620, Francis Bacon (1561–1626) (in his “update on Aristotle”, The New Organon) says, a little more abstractly, if obscurely—and without any reference to atoms or substances:

[It is not] that heat generates motion or that motion generates heat (though both are true in certain cases), but that heat itself, its essence and quiddity, is motion and nothing else.

But real progress in understanding the nature of heat had to wait for more understanding about the nature of gases, with air being the prime example. (It was actually only in the 1640s that any kind of general notion of gas began to emerge—with the word “gas” being invented by the “anti-Galen” physician Jan Baptista van Helmont (1580–1644), as a Dutch rendering of the Greek word “chaos”, that meant essentially “void”, or primordial formlessness.) Ever since antiquity there’d been Aristotle-style explanations like “nature abhors a vacuum” about what nature “wants to do”. But by the mid-1600s the idea was emerging that there could be more explicit and mechanical explanations for phenomena in the natural world.

And in 1660 Robert Boyle (1627–1691)—now thoroughly committed to the experimental approach to science—published New Experiments Physico-mechanicall, Touching the Spring of the Air and its Effects in which he argued that air has an intrinsic pressure associated with it, which pushes it to fill spaces, and for which he effectively found Boyle’s Law PV = constant.

But what was air actually made of? Boyle had two basic hypotheses that he explained in rather flowery terms:

Click to enlarge

His first hypothesis was that air might be like a “fleece of wool” made of “aerial corpuscles” (gases were later often called “aeriform fluids”) with a “power or principle of self-dilatation” that resulted from there being “hairs” or “little springs” between these corpuscles. But he had a second hypothesis too—based, he said, on the ideas of “that most ingenious gentleman, Monsieur Descartes”: that instead air consists of “flexible particles” that are “so whirled around” that “each corpuscle endeavors to beat off all others”. In this second hypothesis, Boyle’s “spring of the air” was effectively the result of particles bouncing off each other.

And, as it happens, in 1668 there was quite an effort to understand the “laws of impact” (that would for example be applicable to balls in games like croquet and billiards, that had existed since at least the 1300s, and were becoming popular), with John Wallis (1616–1703), Christopher Wren (1632–1723) and Christiaan Huygens (1629–1695) all contributing, and Huygens producing diagrams like:

Click to enlarge

But while some understanding developed of what amount to impacts between pairs of hard spheres, there wasn’t the mathematical methodology—or probably the idea—to apply this to large collections of spheres.

Meanwhile, in his 1687 Principia Mathematica, Isaac Newton (1642–1727), wanting to analyze the properties of self-gravitating spheres of fluid, discussed the idea that fluids could in effect be made up of arrays of particles held apart by repulsive forces, as in Boyle’s first hypothesis. Newton had of course had great success with his 1/r2 universal attractive force for gravity. But now he noted (writing originally in Latin) that with a 1/r repulsive force between particles in a fluid, he could essentially reproduce Boyle’s law:

Click to enlarge

Newton discussed questions like whether one particle would “shield” others from the force, but then concluded:

But whether elastic fluids do really consist of particles so repelling each other, is a physical question. We have here demonstrated mathematically the property of fluids consisting of particles of this kind, that hence philosophers may take occasion to discuss that question.

Well, in fact, particularly given Newton’s authority, for well over a century people pretty much just assumed that this was how gases worked. There was one major exception, however, in 1738, when—as part of his eclectic mathematical career spanning probability theory, elasticity theory, biostatistics, economics and more—Daniel Bernoulli (1700–1782) published his book on hydrodynamics. Mostly he discusses incompressible fluids and their flow, but in one section he considers “elastic fluids”—and along with a whole variety of experimental results about atmospheric pressure in different places—draws the picture

Click to enlarge

and says

Let the space ECDF contain very small particles in rapid motion; as they strike against the piston EF and hold it up by their impact, they constitute an elastic fluid which expands as the weight P is removed or reduced; but if P is increased it becomes denser and presses on the horizontal case CD just as if it were endowed with no elastic property.

Then—in a direct and clear anticipation of the kinetic theory of heat—he goes on:

The pressure of the air is increased not only by reduction in volume but also by rise in temperature. As it is well known that heat is intensified as the internal motion of the particles increases, it follows that any increase in the pressure of air that has not changed its volume indicates more intense motion of its particles, which is in agreement with our hypothesis…

But at the time, and in fact for more than a century thereafter, this wasn’t followed up.

A large part of the reason seems to have been that people just assumed that heat ultimately had to have some kind of material existence; to think that it was merely a manifestation of microscopic motion was too abstract an idea. And then there was the observation of “radiant heat” (i.e. infrared radiation)—that seemed like it could only work by explicitly transferring some kind of “heat material” from one body to another.

But what was this “heat material”? It was thought of as a fluid—called caloric—that could suffuse matter, and for example flow from a hotter body to a colder. And in an echo of Democritus, it was often assumed that caloric consisted of particles that could slide between ordinary particles of matter. There was some thought that it might be related to the concept of phlogiston from the mid-1600s, that was effectively a chemical substance, for example participating in chemical reactions or being generated in combustion (through the “principle of fire”). But the more mainstream view was that there were caloric particles that would collect around ordinary particles of matter (often called “molecules”, after the use of that term by Descartes (1596–1650) in 1620), generating a repulsive force that would for example expand gases—and that in various circumstances these caloric particles would move around, corresponding to the transfer of heat.

To us today it might seem hacky and implausible (perhaps a little like dark matter, cosmological inflation, etc.), but the caloric theory lasted for more than two hundred years and managed to explain plenty of phenomena—and indeed was certainly going strong in 1825 when Laplace wrote his A Treatise of Celestial Mechanics, which included a successful computation of properties of gases like the speed of sound and the ratio of specific heats, on the basis of a somewhat elaborated and mathematicized version of caloric theory (that by then included the concept of “caloric rays” associated with radiant heat).

But even though it wasn’t understood what heat ultimately was, one could still measure its attributes. Already in antiquity there were devices that made use of heat to produce pressure or mechanical motion. And by the beginning of the 1600s—catalyzed by Galileo’s development of the thermoscope (in which heated liquid could be seen to expand up a tube)—the idea quickly caught on of making thermometers, and of quantitatively measuring temperature.

And given a measurement of temperature, one could correlate it with effects one saw. So, for example, in the late 1700s the French balloonist Jacques Charles (1746–1823) noted the linear increase of volume of a gas with temperature. Meanwhile, at the beginning of the 1800s Joseph Fourier (1768–1830) (science advisor to Napoleon) developed what became his 1822 Analytical Theory of Heat, and in it he begins by noting that:

Heat, like gravity, penetrates every substance of the universe, its rays occupy all parts of space. The object of our work is to set forth the mathematical laws which this element obeys. The theory of heat will hereafter form one of the most important branches of general physics.

Later he describes what he calls the “Principle of the Communication of Heat”. He refers to “molecules”—though basically just to indicate a small amount of substance—and says

When two molecules of the same solid are extremely near and at unequal temperatures, the most heated molecule communicates to that which is less heated a quantity of heat exactly expressed by the product of the duration of the instant, of the extremely small difference of the temperatures, and of certain function of the distance of the molecules.

then goes on to develop what’s now called the heat equation and all sorts of mathematics around it, all the while effectively adopting a caloric theory of heat. (And, yes, if you think of heat as a fluid it does lead you to describe its “motion” in terms of differential equations just like Fourier did. Though it’s then ironic that Bernoulli, even though he studied hydrodynamics, seemed to have a less “fluid-based” view of heat.)

Heat Engines and the Beginnings of Thermodynamics

At the beginning of the 1800s the Industrial Revolution was in full swing—driven in no small part by the availability of increasingly efficient steam engines. There had been precursors of steam engines even in antiquity, but it was only in 1712 that the first practical steam engine was developed. And after James Watt (1736–1819) produced a much more efficient version in 1776, the adoption of steam engines began to take off.

Over the years that followed there were all sorts of engineering innovations that increased the efficiency of steam engines. But it wasn’t clear how far it could go—and whether for example there was a limit to how much mechanical work could ever, even in principle, be derived from a given amount of heat. And it was the investigation of this question—in the hands of a young French engineer named Sadi Carnot (1796–1832)—that began the development of an abstract basic science of thermodynamics, and to the Second Law.

The story really begins with Sadi Carnot’s father, Lazare Carnot (1753–1823), who was trained as an engineer but ascended to the highest levels of French politics, and was involved with both the French Revolution and Napoleon. Particularly in years when he was out of political favor, Lazare Carnot worked on mathematics and mathematical engineering. His first significant work—in 1778—was entitled Memoir on the Theory of Machines. The mathematical and geometrical science of mechanics was by then fairly well developed; Lazare Carnot’s objective was to understand its consequences for actual engineering machines, and to somehow abstract general principles from the mechanical details of the operation of those machines. In 1803 (alongside works on the geometrical theory of fortifications) he published his Fundamental Principles of [Mechanical] Equilibrium and Movement, which argued for what was at one time called (in a strange foreshadowing of reversible thermodynamic processes) “Carnot’s Principle”: that useful work in a machine will be maximized if accelerations and shocks of moving parts are minimized—and that a machine with perpetual motion is impossible.

Sadi Carnot was born in 1796, and was largely educated by his father until he went to college in 1812. It’s notable that during the years when Sadi Carnot was a kid, one of his father’s activities was to give opinions on a whole range of inventions—including many steam engines and their generalizations. Lazare Carnot died in 1823. Sadi Carnot was by that point a well-educated but professionally undistinguished French military engineer. But in 1824, at the age of 28, he produced his one published work, Reflections on the Motive Power of Fire, and on Machines to Develop That Power (where by “fire” he meant what we would call heat):

Click to enlarge

The style and approach of the younger Carnot’s work is quite similar to his father’s. But the subject matter turned out to be more fruitful. The book begins:

Everyone knows that heat can produce motion. That it possesses vast motive-power none can doubt, in these days when the steam-engine is everywhere so well known… The study of these engines is of the greatest interest, their importance is enormous, their use is continually increasing, and they seem destined to produce a great revolution in the civilized world. Already the steam-engine works our mines, impels our ships, excavates our ports and our rivers, forges iron, fashions wood, grinds grain, spins and weaves our cloths, transports the heaviest burdens, etc. It appears that it must some day serve as a universal motor, and be substituted for animal power, water-falls, and air currents. …

Notwithstanding the work of all kinds done by steam-engines, notwithstanding the satisfactory condition to which they have been brought to-day, their theory is very little understood, and the attempts to improve them are still directed almost by chance. …

The question has often been raised whether the motive power of heat is unbounded, whether the possible improvements in steam-engines have an assignable limit, a limit which the nature of things will not allow to be passed by any means whatever; or whether, on the contrary, these improvements may be carried on indefinitely. We propose now to submit these questions to a deliberate examination.

Carnot operated very much within the framework of caloric theory, and indeed his ideas were crucially based on the concept that one could think about “heat itself” (which for him was caloric fluid), independent of the material substance (like steam) that was hot. But—like his father’s efforts with mechanical machines—his goal was to develop an abstract “metamodel” of something like a steam engine, crucially assuming that the generation of unbounded heat or mechanical work (i.e. perpetual motion) in the closed cycle of the operation of the machine was impossible, and noting (again with a reflection of his father’s work) that the system would necessarily maximize efficiency if it operated reversibly. And he then argued that:

The production of motive power is then due in steam-engines not to an actual consumption of caloric, but to its transportation from a warm body to a cold body, that is, to its re-establishment of equilibrium…

In other words, what was important about a steam engine was that it was a “heat engine”, that “moved heat around”. His book is mostly words, with just a few formulas related to the behavior of ideal gases, and some tables of actual parameters for particular materials. But even though his underlying conceptual framework—of caloric theory—was not correct, the abstract arguments that he made (that involved essentially logical consequences of reversibility and of operating in a closed cycle) were robust enough that it didn’t matter, and in particular he was able to successfully show that there was a theoretical maximum efficiency for a heat engine, that depended only on the temperatures of its hot and cold reservoirs of heat. But what’s important for our purposes here is that in the setup Carnot constructed he basically ended up introducing the Second Law.

At the time it appeared, however, Carnot’s book was basically ignored, and Carnot died in obscurity from cholera in 1832 (about 9 months after Évariste Galois (1811–1832)) at the age of 36. (The Sadi Carnot who would later become president of France was his nephew.) But in 1834, Émile Clapeyron (1799–1864)—a rather distinguished French engineering professor (and steam engine designer)—wrote a paper entitled “Memoir on the Motive Power of Heat”. He starts off by saying about Carnot’s book:

The idea which serves as a basis of his researches seems to me to be both fertile and beyond question; his demonstrations are founded on the absurdity of the possibility of creating motive power or heat out of nothing. …

This new method of demonstration seems to me worthy of the attention of theoreticians; it seems to me to be free of all objection …

I believe that it is of some interest to take up this theory again; S. Carnot, avoiding the use of mathematical analysis, arrives by a chain of difficult and elusive arguments at results which can be deduced easily from a more general law which I shall attempt to prove…

Clapeyron’s paper doesn’t live up to the claims of originality or rigor expressed here, but it served as a more accessible (both in terms of where it was published and how it was written) exposition of Carnot’s work, featuring, for example, for the first time a diagrammatic representation of a Carnot cycle

Click to enlarge

as well as notations like Q-for-heat that are still in use today:

Click to enlarge

The Second Law Is Formulated

One of the implications of Newton’s Laws of Motion is that momentum is conserved. But what else might also be conserved? In the 1680s Gottfried Leibniz (1646–1716) suggested the quantity m v2, which he called, rather grandly, vis viva—or, in English, “life force”. And yes, in things like elastic collisions, this quantity did seem to be conserved. But in plenty of situations it wasn’t. By 1807 the term “energy” had been introduced, but the question remained of whether it could in any sense globally be thought of as conserved.

It had seemed for a long time that heat was something a bit like mechanical energy, but the relation wasn’t clear—and the caloric theory of heat implied that caloric (i.e. the fluid corresponding to heat) was conserved, and so certainly wasn’t something that for example could be interconverted with mechanical energy. But in 1798 Benjamin Thompson (Count Rumford) (1753–1814) measured the heat produced by the mechanical process of boring a cannon, and began to make the argument that, in contradiction to the caloric theory, there was actually some kind of correspondence between mechanical energy and amount of heat.

It wasn’t a very accurate experiment, and it took until the 1840s—with new experiments by the English brewer and “amateur” scientist James Joule (1818–1889) and the German physician Robert Mayer (1814–1878)—before the idea of some kind of equivalence between heat and mechanical work began to look more plausible. And in 1847 this was something William Thomson (1824–1907) (later Lord Kelvin)—a prolific young physicist recently graduated from the Mathematical Tripos in Cambridge and now installed as a professor of “natural philosophy” (i.e. physics) in Glasgow—began to be curious about.

But first we have to go back a bit in the story. In 1845 Kelvin (as we’ll call him) had spent some time in Paris (primarily at at a lab that was measuring properties of steam for the French government), and there he’d learned about Carnot’s work from Clapeyron’s paper (at first he couldn’t get a copy of Carnot’s actual book). Meanwhile, one of the issues of the time was a proliferation of different temperature scales based on using different kinds of thermometers based on different substances. And in 1848 Kelvin realized that Carnot’s concept of a “pure heat engine”—assumed at the time to be based on caloric—could be used to define an “absolute” scale of temperature in which, for example, at absolute zero all caloric would have been removed from all substances:

Click to enlarge

Having found Carnot’s ideas useful, Kelvin in 1849 wrote a 33-page summary of them (small world that it was then, the immediately preceding paper in the journal is “On the Theory of Rolling Curves”, written by the then-17-year-old James Clerk Maxwell (1831–1879), while the one that follows is “Theoretical Considerations on the Effect of Pressure in Lowering the Freezing Point of Water” by James Thomson (1822–1892), engineering-oriented older brother of William):

Click to enlarge

He characterizes Carnot’s work as being based not so much on physics and experiment, but on the “strictest principles of philosophy”:

Click to enlarge

He doesn’t immediately mention “caloric” (though it does slip in later), referring instead to a vaguer concept of “thermal agency”:

Click to enlarge

In keeping with the idea that this is more philosophy than experimental science, he refers to “Carnot’s fundamental principle”—that after a complete cycle an engine can be treated as back in the “same state”—while adding the footnote that “this is tacitly assumed as an axiom”:

Click to enlarge

In actuality, to say that an engine comes back to the same state is a nontrivial statement of the existence of some kind of unique equilibrium in the system, related to the Second Law. But in 1848 Kelvin brushes this off by saying that the “axiom” has “never, so far as I am aware, been questioned by practical engineers”.

His next page is notable for the first-ever use of the term “thermo-dynamic” (then hyphenated) to discuss systems where what matters is “the dynamics of heat”:

Click to enlarge

That same page has a curious footnote presaging what will come, and making the statement that “no energy can be destroyed”, and considering it “perplexing” that this seems incompatible with Carnot’s work and its caloric theory framework:

Click to enlarge

After going through Carnot’s basic arguments, the paper ends with an appendix in which Kelvin basically says that even though the theory seems to just be based on a formal axiom, it should be experimentally tested:

Click to enlarge

He proceeds to give some tests, which he claims agree with Carnot’s results—and finally ends with a very practical (but probably not correct) table of theoretical efficiencies for steam engines of his day:

Click to enlarge

But now what of Joule’s and Mayer’s experiments, and their apparent disagreement with the caloric theory of heat? By 1849 a new idea had emerged: that perhaps heat was itself a form of energy, and that, when heat was accounted for, the total energy of a system would always be conserved. And what this suggested was that heat was somehow a dynamical phenomenon, associated with microscopic motion—which in turn suggested that gases might indeed consist just of molecules in motion.

And so it was that in 1850 Kelvin (then still “William Thomson”) wrote a long exposition “On the Dynamical Theory of Heat”, attempting to reconcile Carnot’s ideas with the new concept that heat was dynamical in origin:

Click to enlarge

He begins by quoting—presumably for some kind of “British-based authority”—an “anti-caloric” experiment apparently done by Humphry Davy (1778–1829) as a teenager, involving melting pieces of ice by rubbing them together, and included anonymously in a 1799 list of pieces of knowledge “principally from the west of England”:

Click to enlarge

But soon Kelvin is getting to the main point:

Click to enlarge

And then we have it: a statement of the Second Law (albeit with some hedging to which we’ll come back later):

Click to enlarge

And there’s immediately a footnote that basically asserts the “absurdity” of a Second-Law-violating perpetual motion machine:

Click to enlarge

But by the next page we find out that Kelvin admits he’s in some sense been “scooped”—by a certain Rudolf Clausius (1822–1888), who we’ll be discussing soon. But what’s remarkable is that Clausius’s “axiom” turns out to be exactly equivalent to Kelvin’s statement:

Click to enlarge

And what this suggests is that the underlying concept—the Second Law—is something quite robust. And indeed, as Kelvin implies, it’s the main thing that ultimately underlies Carnot’s results. And so even though Carnot is operating on the now-outmoded idea of caloric theory, his main results are still correct, because in the end all they really depend on is a certain amount of “logical structure”, together with the Second Law (and a version of the First Law, but that’s a slightly trickier story).

Kelvin recognized, though, that Carnot had chosen to look at the particular (“equilibrium thermodynamics”) case of processes that occur reversibly, effectively at an infinitesimal rate. And at the end of the first installment of his exposition, he explains that things will be more complicated if finite rates are considered—and that in particular the results one gets in such cases will depend on things like having a correct model for the nature of heat.

Kelvin’s exposition on the “dynamical nature of heat” runs to four installments, and the next two dive into detailed derivations and attempted comparison with experiment:

Click to enlarge

But before Kelvin gets to publish part four of his exposition he publishes two other pieces. In the first, he’s talking about sources of energy for human use (now that he believes energy is conserved):

Click to enlarge

He emphasizes that the Sun is—directly or indirectly—the main source of energy on Earth (later he’ll argue that coal will run out, etc.):

Click to enlarge

But he wonders how animals actually manage to produce mechanical work, noting that “the animal body does not act as a thermo-dynamic engine; and [it is] very probable that the chemical forces produce the external mechanical effects through electrical means”:

Click to enlarge

And then, by April 1852, he’s back to thinking directly about the Second Law, and he’s cut through the technicalities, and is stating the Second Law in everyday (if slightly ponderous) terms:

Click to enlarge

It’s interesting to see his apparently rather deeply held Presbyterian beliefs manifest themselves here in his mention that “Creative Power” is what must set the total energy of the universe. He ends his piece with:

Click to enlarge

In (2) the hedging is interesting. He makes the definitive assertion that what amounts to a violation of the Second Law “is impossible in inanimate material processes”. And he’s pretty sure the same is true for “vegetable life” (recognizing that in his previous paper he discussed the harvesting of sunlight by plants). But what about “animal life”, like us humans? Here he says that “by our will” we can’t violate the Second Law—so we can’t, for example, build a machine to do it. But he leaves it open whether we as humans might have some innate (“God-given”?) ability to overcome the Second Law.

And then there’s his (3). It’s worth realizing that his whole paper is less than 3 pages long, and right before his conclusions we’re seeing triple integrals:

Click to enlarge

So what is (3) about? It’s presumably something like a Second-Law-implies-heat-death-of-the-universe statement (but what’s this stuff about the past?)—but with an added twist that there’s something (God?) beyond the “known operations going on at present in the material world” that might be able to swoop in to save the world for us humans.

It doesn’t take people long to pick up on the “cosmic significance” of all this. But in the fall of 1852, Kelvin’s colleague, the Glasgow engineering professor William Rankine (1820–1872) (who was deeply involved with the First Law of thermodynamics), is writing about a way the universe might save itself:

Click to enlarge

After touting the increasingly solid evidence for energy conservation and the First Law

Click to enlarge

he goes on to talk about dissipation of energy and what we now call the Second Law

Click to enlarge

and the fact that it implies an “end of all physical phenomena”, i.e. heat death of the universe. He continues:

Click to enlarge

But now he offers a “ray of hope”. He believes that there must exist a “medium capable of transmitting light and heat”, i.e. an aether, “[between] the heavenly bodies”. And if this aether can’t itself acquire heat, he concludes that all energy must be converted into a radiant form:

Click to enlarge

Now he supposes that the universe is effectively a giant drop of aether, with nothing outside, so that all this radiant energy will get totally internally reflected from its surface, allowing the universe to “[reconcentrate] its physical energies, and [renew] its activity and life”—and save it from heat death:

Click to enlarge

He ends with the speculation that perhaps “some of the luminous objects which we see in distant regions of space may be, not stars, but foci in the interstellar aether”.

But independent of cosmic speculations, Kelvin himself continues to study the “dynamical theory of gases”. It’s often a bit unclear what’s being assumed. There’s the First Law (energy conservation). And the Second Law. But there’s also reversibility. Equilibrium. And the ideal gas law (P V = R T). But it soon becomes clear that that’s not always correct for real gases—as the Joule–Thomson effect demonstrates:

Click to enlarge

Kelvin soon returned to more cosmic speculations, suggesting that perhaps gravitation—rather than direct “Creative Power”—might “in reality [be] the ultimate created antecedent of all motion…”:

Click to enlarge

Not long after these papers Kelvin got involved with the practical “electrical” problem of laying a transatlantic telegraph cable, and in 1858 was on the ship that first succeeded in doing this. (His commercial efforts soon allowed him to buy a 126-ton yacht.) But he continued to write physics papers, which ranged over many different areas, occasionally touching thermodynamics, though most often in the service of answering a “general science” question—like how old the Sun is (he estimated 32,000 years from thermodynamic arguments, though of course without knowledge of nuclear reactions).

Kelvin’s ideas about the inevitable dissipation of “useful energy” spread quickly—by 1854, for example, finding their way into an eloquent public lecture by Hermann von Helmholtz (1821–1894). Helmholtz had trained as a doctor, becoming in 1843 a surgeon to a German military regiment. But he was also doing experiments and developing theories about “animal heat” and how muscles manage to “do mechanical work”, for example publishing an 1845 paper entitled “On Metabolism during Muscular Activity”. And in 1847 he was one of the inventors of the law of conservation of energy—and the First Law of thermodynamics—as well as perhaps its clearest expositor at the time (the word “force” in the title is what we now call “energy”):

Click to enlarge

By 1854 Helmholtz was a physiology professor, beginning a distinguished career in physics, psychophysics and physiology—and talking about the Second Law and its implications. He began his lecture by saying that “A new conquest of very general interest has been recently made by natural philosophy”—and what he’s referring to here is the Second Law:

Click to enlarge

Having discussed the inability of “automata” (he uses that word) to reproduce living systems, he starts talking about perpetual motion machines:

Click to enlarge

First he disposes of the idea that perpetual motion can be achieved by generating energy from nothing (i.e. violating the First Law), charmingly including the anecdote:

Click to enlarge

And then he’s on to talking about the Second Law

Click to enlarge

and discussing how it implies the heat death of the universe:

Click to enlarge

He notes, correctly, that the Second Law hasn’t been “proved”. But he’s impressed at how Kelvin was able to go from a “mathematical formula” to a global fact about the fate of the universe:

Click to enlarge

He ends the whole lecture quite poetically:

Click to enlarge

We’ve talked quite a bit about Kelvin and how his ideas spread. But let’s turn now to Rudolf Clausius, who in 1850 at least to some extent “scooped” Kelvin on the Second Law. At that time Clausius was a freshly minted German physics PhD. His thesis had been on an ingenious but ultimately incorrect theory of why the sky is blue. But he’d also worked on elasticity theory, and there he’d been led to start thinking about molecules and their configurations in materials. By 1850 caloric theory had become fairly elaborate, complete with concepts like “latent heat” (bound to molecules) and “free heat” (able to be transferred). Clausius’s experience in elasticity theory made him skeptical, and knowing Mayer’s and Joule’s results he decided to break with the caloric theory—writing his career-launching paper (translated from German in 1851, with Carnot’s puissance motrice [“motive power”] being rendered as “moving force”):

Click to enlarge

The first installment of the English version of the paper gives a clear description of the ideal gas laws and the Carnot cycle, having started from a statement of the “caloric-busting” First Law:

Click to enlarge

The general discussion continues in the second installment, but now there’s a critical side comment that describes the “general deportment of heat, which every-where exhibits the tendency to annul differences of temperature, and therefore to pass from a warmer body to a colder one”:

Click to enlarge

Clausius “has” the Second Law, as Carnot basically did before him. But when Kelvin quotes Clausius he does so much more forcefully:

Click to enlarge

But there it is: by 1852 the Second Law is out in the open, in at least two different forms. The path to reach it has been circuitous and quite technical. But in the end, stripped of its technical origins, the law seems somehow unsurprising and even obvious. For it’s a matter of common experience that heat flows from hotter bodies to colder ones, and that motion is dissipated by friction into heat. But the point is that it wasn’t until basically 1850 that the overall scientific framework existed to make it useful—or even really possible—to enunciate such observations as a formal scientific law.

Of course the fact that a law “seems true” based on common experience doesn’t mean it’ll always be true, and that there won’t be some special circumstance or elaborate construction that will evade it. But somehow the very fact that the Second Law had in a sense been “technically hard won”—yet in the end seemed so “obvious”—appears to have given it a sense of inevitability and certainty. And it didn’t hurt that somehow it seemed to have emerged from Carnot’s work, which had a certain air of “logical necessity”. (Of course, in reality, the Second Law entered Carnot’s logical structure as an “axiom”.) But all this helped set the stage for some of the curious confusions about the Second Law that would develop over the century that followed.

The Concept of Entropy

In the first half of the 1850s the Second Law had in a sense been presented in two ways. First, as an almost “footnote-style” assumption needed to support the “pure thermodynamics” that had grown out of Carnot’s work. And second, as an explicitly-stated-for-the-first-time—if “obvious”—“everyday” feature of nature, that was now realized as having potentially cosmic significance. But an important feature of the decade that followed was a certain progressive at-least-phenomenological “mathematicization” of the Second Law—pursued most notably by Rudolf Clausius.

In 1854 Clausius was already beginning this process. Perhaps confusingly, he refers to the Second Law as the “second fundamental theorem [Hauptsatz]” in the “mechanical theory of heat”—suggesting it’s something that is proved, even though it’s really introduced just as an empirical law of nature, or perhaps a theoretical axiom:

Click to enlarge

He starts off by discussing the “first fundamental theorem”, i.e. the First Law. And he emphasizes that this implies that there’s a quantity U (which we now call “internal energy”) that is a pure “function of state”—so that its value depends only on the state of a system, and not the path by which that state was reached. And as an “application” of this, he then points out that the overall change in U in a cyclic process (like the one executed by Carnot’s heat engine) must be zero.

And now he’s ready to tackle the Second Law. He gives a statement that at first seems somewhat convoluted:

Click to enlarge

But soon he’s deriving this from a more “everyday” statement of the Second Law (which, notably, is clearly not a “theorem” in any normal sense):

Click to enlarge

After giving a Carnot-style argument he’s then got a new statement (that he calls “the theorem of the equivalence of transformations”) of the Second Law:

Click to enlarge

And there it is: basically what we now call entropy (even with the same notation of Q for heat and T for temperature)—together with the statement that this quantity is a function of state, so that its differences are “independent of the nature of the process by which the transformation is effected”.

Pretty soon there’s a familiar expression for entropy change:

Click to enlarge

And by the next page he’s giving what he describes as “the analytical expression” of the Second Law, for the particular case of reversible cyclic processes:

Click to enlarge

A bit later he backs out of the assumption of reversibility, concluding that:

Click to enlarge

(And, yes, with modern mathematical rigor, that should be “non-negative” rather than “positive”.)

He goes on to say that if something has changed after going around a cycle, he’ll call that an “uncompensated transformation”—or what we would now refer to as an irreversible change. He lists a few possible (now very familiar) examples:

Click to enlarge

Earlier in his paper he’s careful to say that T is “a function of temperature”; he doesn’t say it’s actually the quantity we measure as temperature. But now he wants to determine what it is:

Click to enlarge

He doesn’t talk about the ultimately critical assumption (effectively the Zeroth Law of thermodynamics) that the system is “in equilibrium”, with a uniform temperature. But he uses an ideal gas as a kind of “standard material”, and determines that, yes, in that case T can be simply the absolute temperature.

So there it is: in 1854 Clausius has effectively defined entropy and described its relation to the Second Law, though everything is being done in a very “heat-engine” style. And pretty soon he’s writing about “Theory of the Steam-Engine” and filling actual approximate steam tables into his theoretical formulas:

Click to enlarge

After a few years “off” (working, as we’ll discuss later, on the kinetic theory of gases) Clausius is back in 1862 talking about the Second Law again, in terms of his “theorem of the equivalence of transformations”:

Click to enlarge

He’s slightly tightened up his 1854 discussion, but, more importantly, he’s now stating a result not just for reversible cyclic processes, but for general ones:

Click to enlarge

But what does this result really mean? Clausius claims that this “theorem admits of strict mathematical proof if we start from the fundamental proposition above quoted”—though it’s not particularly clear just what that proposition is. But then he says he wants to find a “physical cause”:

Click to enlarge

A little earlier in the paper he said:

Click to enlarge

So what does he think the “physical cause” is? He says that even from his first investigations he’d assumed a general law:

Click to enlarge

What are these “resistances”? He’s basically saying they are the forces between molecules in a material (which from his work on the kinetic theory of gases he now imagines exist):

Click to enlarge

He introduces what he calls the “disgregation” to represent the microscopic effect of adding heat:

Click to enlarge

For ideal gases things are straightforward, including the proportionality of “resistance” to absolute temperature. But in other cases, it’s not so clear what’s going on. A decade later he identifies “disgregation” with average kinetic energy per molecule—which is indeed proportional to absolute temperature. But in 1862 it’s all still quite muddy, with somewhat curious statements like:

Click to enlarge

And then the main part of the paper ends with what seems to be an anticipation of the Third Law of thermodynamics:

Click to enlarge

There’s an appendix entitled “On Terminology” which admits that between Clausius’s own work, and other people’s, it’s become rather difficult to follow what’s going on. He agrees that the term “energy” that Kelvin is using makes sense. He suggests “energy of the body” for what he calls U and we now call “internal energy”. He suggests “heat of the body” or “thermal content of the body” for Q. But then he talks about the fact that these are measured in thermal units (say the amount of heat needed to increase the temperature of water by 1°), while mechanical work is measured in units related to kilograms and meters. He proposes therefore to introduce the concept of “ergon” for “work measured in thermal units”:

Click to enlarge

And pretty soon he’s talking about the “interior ergon” and “exterior ergon”, as well as concepts like “ergonized heat”. (In later work he also tries to introduce the concept of “ergal” to go along with his development of what he called—in a name that did stick—the “virial theorem”.)

But in 1865 he has his biggest success in introducing a term. He’s writing a paper, he says, basically to clarify the Second Law, (or, as he calls it, “the second fundamental theorem”—rather confidently asserting that he will “prove this theorem”):

Click to enlarge

Part of the issue he’s trying to address is how the calculus is done:

Click to enlarge

The partial derivative symbol ∂ had been introduced in the late 1700s. He doesn’t use it, but he does introduce the now-standard-in-thermodynamics subscript notation for variables that are kept constant:

Click to enlarge

A little later, as part of the “notational cleanup”, we see the variable S:

Click to enlarge

And then—there it is—Clausius introduces the term “entropy”, “Greekifying” his concept of “transformation”:

Click to enlarge

His paper ends with his famous crisp statements of the First and Second Laws of thermodynamics—manifesting the parallelism he’s been claiming between energy and entropy:

Click to enlarge

The Kinetic Theory of Gases

We began above by discussing the history of the question of “What is heat?” Was it like a fluid—the caloric theory? Or was it something more dynamical, and in a sense more abstract? But then we saw how Carnot—followed by Kelvin and Clausius—managed in effect to sidestep the question, and come up with all sorts of “thermodynamic conclusions”, by talking just about “what heat does” without ever really having to seriously address the question of “what heat is”. But to be able to discuss the foundations of the Second Law—and what it says about heat—we have to know more about what heat actually is. And the crucial development that began to clarify the nature of heat was the kinetic theory of gases.

Central to the kinetic theory of gases is the idea that gases are made up of discrete molecules. And it’s important to remember that it wasn’t until the beginning of the 1900s that anyone knew for sure that molecules existed. Yes, something like them had been discussed ever since antiquity, and in the 1800s there was increasing “circumstantial evidence” for them. But nobody had directly “seen a molecule”, or been able, for example, until about 1870, to even guess what the size of molecules might be. Still, by the mid-1800s it had become common for physicists to talk and reason in terms of ordinary matter at least effectively being made of up molecules.

But if a gas was made of molecules bouncing off each other like billiard balls according to the laws of mechanics, what would its overall properties be? Daniel Bernoulli had in 1738 already worked out the basic answer that pressure would vary inversely with volume, or in his notation, π = P/s (and he even also gave formulas for molecules of nonzero size—in a precursor of van der Waals):

Click to enlarge

Results like Bernouilli’s would be rediscovered several times, for example in 1820 by John Herapath (1790–1868), a math teacher in England, who developed a fairly elaborate theory that purported to describe gravity as well as heat (but for example implied a P V = a T2 gas law):

Click to enlarge

Then there was the case of John Waterston (1811–1883), a naval instructor for the East India company, who in 1843 published a book called Thoughts on the Mental Functions, which included results on what he called the “vis viva theory of heat”—that he developed in more detail in a paper he wrote in 1846. But when he submitted the paper to the Royal Society it was rejected as “nonsense”, and its manuscript was “lost” until 1891 when it was finally published (with an “explanation” of the “delay”):

Click to enlarge

The paper had included a perfectly sensible mathematical analysis that included a derivation of the kinetic theory relation between pressure and mean-square molecular velocity:

Click to enlarge

But with all these pieces of work unknown, it fell to a German high-school chemistry teacher (and sometime professor and philosophical/theological writer) named August Krönig (1822–1879) to publish in 1856 yet another “rediscovery”, that he entitled “Principles of a Theory of Gases”. He said it was going to analyze the “mechanical theory of heat”, and once again he wanted to compute the pressure associated with colliding molecules. But to simplify the math, he assumed that molecules went only along the coordinate directions, at a fixed speed—almost anticipating a cellular automaton fluid:

Click to enlarge

What ultimately launched the subsequent development of the kinetic theory of gases, however, was the 1857 publication by Rudolf Clausius (by then an increasingly established German physics professor) of a paper entitled rather poetically “On the Nature of the Motion Which We Call Heat” (“Über die Art der Bewegung die wir Wärme nennen”):

Click to enlarge

It’s a clean and clear paper, with none of the mathematical muddiness around Clausius’s work on the Second Law (which, by the way, isn’t even mentioned in this paper even though Clausius had recently worked on it). Clausius figures out lots of the “obvious” implications of his molecular theory, outlining for example what happens in different phases of matter:

Click to enlarge

It takes him only a couple of pages of very light mathematics to derive the standard kinetic theory formula for the pressure of an ideal gas:

Click to enlarge

He’s implicitly assuming a certain randomness to the motions of the molecules, but he barely mentions it (and this particular formula is robust enough that average values are actually all that matter):

Click to enlarge

But having derived the formula for pressure, he goes on to use the ideal gas law to derive the relation between average molecular kinetic energy (which he still calls “vis viva”) and absolute temperature:

Click to enlarge

From this he can do things like work out the actual average velocities of molecules in different gases—which he does without any mention of the question of just how real or not molecules might be. By knowing experimental results about specific heats of gases he also manages to determine that not all the energy (“heat”) in a gas is associated with “translatory motion”: he realizes that for molecules involving several atoms there can be energy associated with other (as we would now say) internal degrees of freedom:

Click to enlarge

Clausius’s paper was widely read. And it didn’t take long before the Dutch meteorologist (and effectively founder of the World Meteorological Organization) Christophorus Buys Ballot (1817–1890) asked why—if molecules were moving as quickly as Clausius suggested—gases didn’t mix much more quickly than they’re observed to do:

Click to enlarge

Within a few months, Clausius published the answer: the molecules didn’t just keep moving in straight lines; they were constantly being deflected, to follow what we would now call a random walk. He invented the concept of a mean free path to describe how far on average a molecule goes before it hits another molecule:

Click to enlarge

As a capable theoretical physicist, Clausius quickly brings in the concept of probability

Click to enlarge

and is soon computing the average number of molecules which will survive undeflected for a certain distance:

Click to enlarge

Then he works out the mean free path λ (and it’s often still called λ):

Click to enlarge

And he concludes that actually there’s no conflict between rapid microscopic motion and large-scale “diffusive” motion:

Click to enlarge

Of course, he could have actually drawn a sample random walk, but drawing diagrams wasn’t his style. And in fact it seems as if the first published drawing of a random walk was something added by John Venn (1834–1923) in the 1888 edition of his Logic of Chance—and, interestingly, in alignment with my computational irreducibility concept from a century later he used the digits of π to generate his “randomness”:

Click to enlarge

In 1859, Clausius’s paper came to the attention of the then-28-year-old James Clerk Maxwell, who had grown up in Scotland, done the Mathematical Tripos in Cambridge, and was now back in Scotland as professor of “natural philosophy” at Aberdeen. Maxwell had already worked on things like elasticity theory, color vision, the mechanics of tops, the dynamics of the rings of Saturn and electromagnetism—having published his first paper (on geometry) at age 14. And, by the way, Maxwell was quite a “diagrammist”—and his early papers include all sorts of pictures that he drew:

Click to enlarge

But in 1859 Maxwell applied his talents to what he called the “dynamical theory of gases”:

Click to enlarge

He models molecules as hard spheres, and sets about computing the “statistical” results of their collisions:

Click to enlarge

And pretty soon he’s trying to compute distribution of their velocities:

Click to enlarge

It’s a somewhat unconvincing (or, as Maxwell himself later put it, “precarious”) derivation (how does it work in 1D, for example?), but somehow it manages to produce what’s now known as the Maxwell distribution:

Click to enlarge

Maxwell observes that the distribution is the same as for “errors … in the ‘method of least squares’”:

Click to enlarge

Maxwell didn’t get back to the dynamical theory of gases until 1866, but in the meantime he was making a “dynamical theory” of something else: what he called the electromagnetic field:

Click to enlarge

Even though he’d worked extensively with the inverse square law of gravity he didn’t like the idea of “action at a distance”, and for example he wanted magnetic field lines to have some underlying “material” manifestation

Click to enlarge

imagining that they might be associated with arrays of “molecular vortices”:

Click to enlarge

We now know, of course, that there isn’t this kind of “underlying mechanics” for the electromagnetic field. But—with shades of the story of Carnot—even though the underlying framework isn’t right, Maxwell successfully derives correct equations for the electromagnetic field—that are now known as Maxwell’s equations:

Click to enlarge

His statement of how the electromagnetic field “works” is highly reminiscent of the dynamical theory of gases:

Click to enlarge

But he quickly and correctly adds:

Click to enlarge

And a few sections later he derives the idea of general electromagnetic waves

Click to enlarge

noting that there’s no evidence that the medium through which he assumes they’re propagating has elasticity:

Click to enlarge

By the way, when it comes to gravity he can’t figure out how to make his idea of a “mechanical medium” work:

Click to enlarge

But in any case, after using it as an inspiration for thinking about electromagnetism, Maxwell in 1866 returns to the actual dynamical theory of gases, still feeling that he needs to justify looking at a molecular theory:

Click to enlarge

And now he gives a recognizable (and correct, so far as it goes) derivation of the Maxwell distribution:

Click to enlarge

He goes on to try to understand experimental results on gases, about things like diffusion, viscosity and conductivity. For some reason, Maxwell doesn’t want to think of molecules, as he did before, as hard spheres. And instead he imagines that they have “action at a distance” forces, which basically work like hard squares if it’s r-5 force law:

Click to enlarge

In the years that followed, Maxwell visited the dynamical theory of gases several more times. In 1871, a few years before he died at age 48, he wrote a textbook entitled Theory of Heat, which begins, in erudite fashion, discussing what “thermodynamics” should even be called:

Click to enlarge

Most of the book is concerned with the macroscopic “theory of heat”—though, as we’ll discuss later, in the very last chapter Maxwell does talk about the “molecular theory”, if in somewhat tentative terms.

“Deriving” the Second Law from Molecular Dynamics

The Second Law was in effect originally introduced as a formalization of everyday observations about heat. But the development of kinetic theory seemed to open up the possibility that the Second Law could actually be proved from the underlying mechanics of molecules. And this was something that Ludwig Boltzmann (1844–1906) embarked on towards the end of his physics PhD at the University of Vienna. In 1865 he’d published his first paper (“On the Movement of Electricity on Curved Surfaces”), and in 1866 he published his second paper, “On the Mechanical Meaning of the Second Law of Thermodynamics”:

Click to enlarge

The introduction promises “a purely analytical, perfectly general proof of the Second Law”. And what he seemed to imagine was that the equations of mechanics would somehow inevitably lead to motion that would reproduce the Second Law. And in a sense what computational irreducibility, rule 30, etc. now show is that in the end that’s indeed basically how things work. But the methods and conceptual framework that Boltzmann had at his disposal were very far away from being able to see that. And instead what Boltzmann did was to use standard mathematical methods from mechanics to compute average properties of cyclic mechanical motions—and then made the somewhat unconvincing claim that combinations of these averages could be related (e.g. via temperature as average kinetic energy) to “Clausius’s entropy”:

Click to enlarge

It’s not clear how much this paper was read, but in 1871 Boltzmann (now a professor of mathematical physics in Graz) published another paper entitled simply “On the Priority of Finding the Relationship between the Second Law of Thermodynamics and the Principle of Least Action” that claimed (with some justification) that Clausius’s then-newly-announced virial theorem was already contained in Boltzmann’s 1866 paper.

But back in 1868—instead of trying to get all the way to Clausius’s entropy—Boltzmann instead uses mechanics to get a generalization of Maxwell’s law for the distribution of molecular velocities. His paper “Studies on the Equilibrium of [Kinetic Energy] between [Point Masses] in Motion” opens by saying that while analytical mechanics has in effect successfully studied the evolution of mechanical systems “from a given state to another”, it’s had little to say about what happens when such systems “have been left moving on their own for a long time”. He intends to remedy that, and spends 47 pages—complete with elaborate diagrams and formulas about collisions between hard spheres—in deriving an exponential distribution of energies if one assumes “equilibrium” (or, more specifically, balance between forward and backward processes):

Click to enlarge

It’s notable that one of the mathematical approaches Boltzmann uses is to discretize (i.e. effectively quantize) things, then look at the “combinatorial” limit. (Based on his later statements, he didn’t want to trust “purely continuous” mathematics—at least in the context of discrete molecular processes—and wanted to explicitly “watch the limits happening”.) But in the end it’s not clear that Boltzmann’s 1868 arguments do more than the few-line functional-equation approach that Maxwell had already used. (Maxwell would later complain about Boltzmann’s “overly long” arguments.)

Boltzmann’s 1868 paper had derived what the distribution of molecular energies should be “in equilibrium”. (In 1871 he was talking about “equipartition” not just of kinetic energy, but also of energies associated with “internal motion” of polyatomic molecules.) But what about the approach to equilibrium? How would an initial distribution of molecular energies evolve over time? And would it always end up at the exponential (“Maxwell–Boltzmann”) distribution? These are questions deeply related to a microscopic understanding of the Second Law. And they’re what Boltzmann addressed in 1872 in his 22nd published paper “Further Studies on the Thermal Equilibrium of Gas Molecules”:

Click to enlarge

Boltzmann explains that:

Maxwell already found the value Av2 eBv2 [for the distribution of velocities] … so that the probability of different velocities is given by a formula similar to that for the probability of different errors of observation in the theory of the method of least squares. The first proof which Maxwell gave for this formula was recognized to be incorrect even by himself. He later gave a very elegant proof that, if the above distribution has once been established, it will not be changed by collisions. He also tries to prove that it is the only velocity distribution that has this property. But the latter proof appears to me to contain a false inference. It has still not yet been proved that, whatever the initial state of the gas may be, it must always approach the limit found by Maxwell. It is possible that there may be other possible limits. This proof is easily obtained, however, by the method which I am about to explain…

(He gives a long footnote explaining why Maxwell might be wrong, talking about how a sequence of collisions might lead to a “cycle of velocity states”—which Maxwell hasn’t proved will be traversed with equal probability in each direction. Ironically, this is actually already an analog of where things are going to go wrong with Boltzmann’s own argument.)

The main idea of Boltzmann’s paper is not to assume equilibrium, but instead to write down an equation (now called the Boltzmann Transport Equation) that explicitly describes how the velocity (or energy) distribution of molecules will change as a result of collisions. He begins by defining infinitesimal changes in time:

Click to enlarge

He then goes through a rather elaborate analysis of velocities before and after collisions, and how to integrate over them, and eventually winds up with a partial differential equation for the time variation of the energy distribution (yes, he confusingly uses x to denote energy)—and argues that Maxwell’s exponential distribution is a stationary solution to this equation:

Click to enlarge

A few paragraphs further on, something important happens: Boltzmann introduces a function that here he calls E, though later he’ll call it H:

Click to enlarge

Ten pages of computation follow

Click to enlarge

and finally Boltzmann gets his main result: if the velocity distribution evolves according to his equation, H can never increase with time, becoming zero for the Maxwell distribution. In other words, he is saying that he’s proved that a gas will always (“monotonically”) approach equilibrium—which seems awfully like some kind of microscopic proof of the Second Law.

But then Boltzmann makes a bolder claim:

It has thus been rigorously proved that, whatever the initial distribution of kinetic energy may be, in the course of a very long time it must always necessarily approach the one found by Maxwell. The procedure used so far is of course nothing more than a mathematical artifice employed in order to give a rigorous proof of a theorem whose exact proof has not previously been found. It gains meaning by its applicability to the theory of polyatomic gas molecules. There one can again prove that a certain quantity E can only decrease as a consequence of molecular motion, or in a limiting case can remain constant. One can also prove that for the atomic motion of a system of arbitrarily many material points there always exists a certain quantity which, in consequence of any atomic motion, cannot increase, and this quantity agrees up to a constant factor with the value found for the well-known integral ∫dQ/T in my [1871] paper on the “Analytical proof of the 2nd law, etc.”. We have therefore prepared the way for an analytical proof of the Second Law in a completely different way from those previously investigated. Up to now the object has been to show that ∫dQ/T = 0 for reversible cyclic processes, but it has not been proved analytically that this quantity is always negative for irreversible processes, which are the only ones that occur in nature. The reversible cyclic process is only an ideal, which one can more or less closely approach but never completely attain. Here, however, we have succeeded in showing that ∫dQ/T is in general negative, and is equal to zero only for the limiting case, which is of course the reversible cyclic process (since if one can go through the process in either direction, ∫dQ/T cannot be negative).

In other words, he’s saying that the quantity H that he’s defined microscopically in terms of velocity distributions can be identified (up to a sign) with the entropy that Clausius defined as dQ/T. He says that he’ll show this in the context of analyzing the mechanics of polyatomic molecules.

But first he’s going to take a break and show that his derivation doesn’t need to assume continuity. In a pre-quantum-mechanics pre-cellular-automaton-fluid kind of way he replaces all the integrals by limits of sums of discrete quantities (i.e. he’s quantizing kinetic energy, etc.):

Click to enlarge

He says that this discrete approach makes everything clearer, and quotes Lagrange’s derivation of vibrations of a string as an example of where this has happened before. But then he argues that everything works out fine with the discrete approach, and that H still decreases, with the Maxwell distribution as the only possible end point. As an aside, he makes a jab at Maxwell’s derivation, pointing out that with Maxwell’s functional equation:

… there are infinitely many other solutions, which are not useful however since ƒ(x) comes out negative or imaginary for some values of x. Hence, it follows very clearly that Maxwell’s attempt to prove a priori that his solution is the only one must fail, since it is not the only one but rather it is the only one that gives purely positive probabilities, and therefore the only useful one.

But finally—after another aside about computing thermal conductivities of gases—Boltzmann digs into polyatomic molecules, and his claim about H being related to entropy. There’s another 26 pages of calculations, and then we get to a section entitled “Solution of Equation (81) and Calculation of Entropy”. More pages of calculation about polyatomic molecules ensue. But finally we’re computing H, and, yes, it agrees with the Clausius result—but anticlimactically he’s only dealing with the case of equilibrium for monatomic molecules, where we already knew we got the Maxwell distribution:

Click to enlarge

And now he decides he’s not talking about polyatomic molecules anymore, and instead:

In order to find the relation of the quantity [H] to the second law of thermodynamics in the form ∫dQ/T < 0, we shall interpret the system of mass points not, as previously, as a gas molecule, but rather as an entire body.

But then, in the last couple of pages of his paper, Boltzmann pulls out another idea. He’s discussed the concept that polyatomic molecules (or, now, whole systems) can be in many different configurations, or “phases”. But now he says: “We shall replace [our] single system by a large number of equivalent systems distributed over many different phases, but which do not interact with each other”. In other words, he’s introducing the idea of an ensemble of states of a system. And now he says that instead of looking at the distribution just for a single velocity, we should do it for all velocities, i.e. for the whole “phase” of the system.

[These distributions] may be discontinuous, so that they have large values when the variables are very close to certain values determined by one or more equations, and otherwise vanishingly small. We may choose these equations to be those that characterize visible external motion of the body and the kinetic energy contained in it. In this connection it should be noted that the kinetic energy of visible motion corresponds to such a large deviation from the final equilibrium distribution of kinetic energy
that it leads to an infinity in H, so that from the point of view of the Second Law of thermodynamics it acts like heat supplied from an infinite temperature.

There are a bunch of ideas swirling around here. Phase-space density (cf. Liouville’s equation). Coarse-grained variables. Microscopic representation of mechanical work. Etc. But the paper is ending. There’s a discussion about H for systems that interact, and how there’s an equilibrium value achieved. And finally there’s a formula for entropy

Click to enlarge

that Boltzmann said “agrees … with the expression I found in my previous [1871] paper”.

So what exactly did Boltzmann really do in his 1872 paper? He introduced the Boltzmann Transport Equation which allows one to compute at least certain non-equilibrium properties of gases. But is his ƒ log ƒ quantity really what we can call “entropy” in the sense Clausius meant? And is it true that he’s proved that entropy (even in his sense) increases? A century and a half later there’s still a remarkable level of confusion around both these issues.

But in any case, back in 1872 Boltzmann’s “minimum theorem” (now called his “H theorem”) created quite a stir. But after some time there was an objection raised, which we’ll discuss below. And partly in response to this, Boltzmann (after spending time working on microscopic models of electrical properties of materials—as well as doing some actual experiments) wrote another major paper on entropy and the Second Law in 1877:

Click to enlarge

The translated title of the paper is “On the Relation between the Second Law of Thermodynamics and Probability Theory with Respect to the Laws of Thermal Equilibrium”. And at the very beginning of the paper Boltzmann makes a statement that was pivotal for future discussions of the Second Law: he says it’s now clear to him that an “analytical proof” of the Second Law is “only possible on the basis of probability calculations”. Now that we know about computational irreducibility and its implications one could say that this was the point where Boltzmann and those who followed him went off track in understanding the true foundations of the Second Law. But Boltzmann’s idea of introducing probability theory was effectively what launched statistical mechanics, with all its rich and varied consequences.

Boltzmann makes his basic claim early in the paper

Click to enlarge

with the statement (quoting from a comment in a paper he’d written earlier the same year) that “it is clear” (always a dangerous thing to say!) that in thermal equilibrium all possible states of the system—say, spatially uniform and nonuniform alike—are equally probable

… comparable to the situation in the game of Lotto where every single quintet is as improbable as the quintet 12345. The higher probability that the state distribution becomes uniform with time arises only because there are far more uniform than nonuniform state distributions…

He goes on:

[Thus] it is possible to calculate the thermal equilibrium state by finding the probability of the different possible states of the system. The initial state will in most cases be highly improbable but from it the system will always rapidly approach a more probable state until it finally reaches the most probable state, i.e., that of thermal equilibrium. If we apply this to the Second Law we will be able to identify the quantity which is usually called entropy with the probability of the particular state…

He’s talked about thermal equilibrium, even in the title, but now he says:

… our main purpose here is not to limit ourselves to thermal equilibrium, but to explore the relationship of the probabilistic formulation to the [Second Law].

He says his goal is to calculate probability distribution for different states, and he’ll start with

as simple a case as possible, namely a gas of rigid absolutely elastic spherical molecules trapped in a container with absolutely elastic walls. (Which interact with central forces only within a certain small distance, but not otherwise; the latter assumption, which includes the former as a special case, does not change the calculations in the least).

In other words, yet again he’s going to look at hard sphere gases. But, he says:

Even in this case, the application of probability theory is not easy. The number of molecules is not infinite, in a mathematical sense, yet the number of velocities each molecule is capable of is effectively infinite. Given this last condition, the calculations are very difficult; to facilitate understanding, I will, as in earlier work, consider a limiting case.

And this is where he “goes discrete” again—allowing (“cellular-automaton-style”) only discrete possible velocities for each molecule:

Click to enlarge

He says that upon colliding, two molecules can exchange these discrete velocities, but nothing more. As he explains, though:

Even if, at first sight, this seems a very abstract way of treating the problem, it rapidly leads to the desired objective, and when you consider that in nature all infinities are but limiting cases, one assumes each molecule can behave in this fashion only in the limiting case where each molecule can assume more and more values of the velocity.

But now—much like in an earlier paper—he makes things even simpler, saying he’s going to ignore velocities for now, and just say that the possible energies of molecules are “in an arithmetic progression”:

Click to enlarge

He plans to look at collisions, but first he just wants to consider the combinatorial problem of distributing these energies among n molecules in all possible ways, subject to the constraint of having a certain fixed total energy. He sets up a specific example, with 7 molecules, total energy 7, and maximum energy per molecule 7—then gives an explicit table of all possible states (up to, as he puts it, “immaterial permutations of molecular labels”):

Click to enlarge

Tables like this had been common for nearly two centuries in combinatorial mathematics books like Jacob Bernoulli’s (1655–1705) Ars Conjectandi

Click to enlarge

but this might have been the first place such a table had appeared in a paper about fundamental physics.

And now Boltzmann goes into an analysis of the distribution of states—of the kind that’s now long been standard in textbooks of statistical physics, but will then have been quite unfamiliar to the pure-calculus-based physicists of the time:

Click to enlarge

He derives the average energy per molecule, as well as the fluctuations:

Click to enlarge

He says that “of course” the real interest is in the limit of an infinite number of molecules, but he still wants to show that for “moderate values” the formulas remain quite accurate. And then (even without Wolfram Language!) he’s off finding (using Newton’s method it seems) approximate roots of the necessary polynomials:

Click to enlarge

Just to show how it all works, he considers a slightly larger case as well:

Click to enlarge

Now he’s computing the probability that a given molecule has a particular energy

Click to enlarge

and determining that in the limit it’s an exponential

Click to enlarge

that is, as he says, “consistent with that known from gases in thermal equilibrium”.

He claims that in order to really get a “mechanical theory of heat” it’s necessary to take a continuum limit. And here he concludes that thermal equilibrium is achieved by maximizing the quantity Ω (where the “l” stands for log, so this is basically ƒ log ƒ):

Click to enlarge

He explains that Ω is basically the log of the number of possible permutations, and that it’s “of special importance”, and he’ll call it the “permutability measure”. He immediately notes that “the total permutability measure of two bodies is equal to the sum of the permutability measures of each body”. (Note that Boltzmann’s Ω isn’t the modern total-number-of-states Ω; confusingly, that’s essentially the exponential of Boltzmann’s Ω.)

He goes through some discussion of how to handle extra degrees of freedom in polyatomic molecules, but then he’s on to the main event: arguing that Ω is (essentially) the entropy. It doesn’t take long:

Click to enlarge

Basically he just says that in equilibrium the probability ƒ(…) for a molecule to have a particular velocity is given by the Maxwell distribution, then he substitutes this into the formula for Ω, and shows that indeed, up to a constant, Ω is exactly the “Clausius entropy” ∫dQ/T.

So, yes, in equilibrium Ω seems to be giving the entropy. But then Boltzmann makes a bit of a jump. He says that in processes that aren’t reversible both “Clausius entropy” and Ω will increase, and can still be identified—and enunciates the general principle, printed in his paper in special doubled-spaced form:

Click to enlarge

… [In] any system of bodies that undergoes state changes … even if the initial and final states are not in thermal equilibrium … the total permutability measure for the bodies will continually increase during the state changes, and can remain constant only so long as all the bodies during the state changes remain infinitely close to thermal equilibrium (reversible state changes).

In other words, he’s asserting that Ω behaves the same way entropy is said to behave according to the Second Law. He gives various thought experiments about gases in boxes with dividers, gases under gravity, etc. And finally concludes that, yes, the relationship of entropy to Ω “applies to the general case”.

There’s one final paragraph in the paper, though:

Up to this point, these propositions may be demonstrated exactly using the theory of gases. If one tries, however, to generalize to liquid drops and solid bodies, one must dispense with an exact treatment from the outset, since far too little is known about the nature of the latter states of matter, and the mathematical theory is barely developed. But I have already mentioned reasons in previous papers, in virtue of which it is likely that for these two aggregate states, the thermal equilibrium is achieved when Ω becomes a maximum, and that when thermal equilibrium exists, the entropy is given by the same expression. It can therefore be described as likely that the validity of the principle which I have developed is not just limited to gases, but that the same constitutes a general natural law applicable to solid bodies and liquid droplets, although the exact mathematical treatment of these cases still seems to encounter extraordinary difficulties.

Interestingly, Boltzmann is only saying that it’s “likely” that in thermal equilibrium his permutability measure agrees with Clausius’s entropy, and he’s implying that actually that’s really the only place where Clausius’s entropy is properly defined. But certainly his definition is more general (after all, it doesn’t refer to things like temperature that are only properly defined in equilibrium), and so—even though Boltzmann didn’t explicitly say it—one can imagine basically just using it as the definition of entropy for arbitrary cases. Needless to say, the story is actually more complicated, as we’ll see soon.

But this definition of entropy—crispened up by Max Planck (1858–1947) and with different notation—is what ended up years later “written in stone” at Boltzmann’s grave:

Click to enlarge

The Concept of Ergodicity

In his 1877 paper Boltzmann had made the claim that in equilibrium all possible microscopic states of a system would be equally probable. But why should this be true? One reason could be that in its pure “mechanical evolution” the system would just successively visit all these states. And this was an idea that Boltzmann seems to have had—with increasing clarity—from the time of his very first paper in 1866 that purported to “prove the Second Law” from mechanics.

In modern times—with our understanding of discrete systems and computational rules—it’s not difficult to describe the idea of “visiting all states”. But in Boltzmann’s time it was considerably more complicated. Did one expect to hit all the infinite possible infinitesimally separated configurations of a system? Or somehow just get close? The fact is that Boltzmann had certainly dipped his toe into thinking about things in terms of discrete quantities. But he didn’t make the jump to imagining discrete rules, even though he certainly did know about discrete iterative processes, like Newton’s method for finding roots.

Boltzmann knew about cases—like circular motion—where everything was purely periodic. But maybe when motion wasn’t periodic, it’d inevitably “visit all states”. Already in 1868 Boltzmann was writing a paper entitled “Solution to a Mechanical Problem” where he studies a single point mass moving in an α/r – β/r2 potential and bouncing elastically off a line—and manages to show that it visits every position with equal probability. In this paper he’s just got traditional formulas, but by 1871, in “Some General Theorems about Thermal Equilibrium”—computing motion in the same potential as before—he’s got a picture:

Click to enlarge

Boltzmann probably knew about Lissajous figures—cataloged in 1857

Click to enlarge

and the fact that in this case a rational ratio of x and y periods gives a periodic overall curve while an irrational one always gives a curve that visits every position might have led him to suspect that all systems would either be periodic, or would visit every possible configuration (or at least, as he identified in his paper, every configuration that had the same values of “constants of the motion”, like energy).

In early 1877 Boltzmann returned to the same question, including as one section in his “Remarks on Some Problems in the Mechanical Theory of Heat” more analysis of the same potential as before, but now showing a diversity of more complicated pictures that almost seem to justify his rule-30-before-its-time idea that there could be “pure mechanics” that would lead to “Second Law” behavior:

Click to enlarge

In modern times, of course, it’s easy to solve those equations of motion, and typical results obtained for an array of values of parameters are:

Boltzmann returned to these questions in 1884, responding to Helmholtz’s analysis of what he was calling “monocyclic systems”. Boltzmann used the same potential again, but now with a name for the “visit-all-states” property: isodic. Meanwhile, Boltzmann had introduced the name “ergoden” for the collection of all possible configurations of a system with a given energy (what would now be called the microcanonical ensemble). But somehow, quite a few years later, Boltzmann’s student Paul Ehrenfest (1880–1933) (along with Tatiana Ehrenfest-Afanassjewa (1876–1964)) would introduce the term “ergodic” for Boltzmann’s isodic. And “ergodic” is the term that caught on. And in the twentieth century there was all sorts of development of “ergodic theory”, as we’ll discuss a bit later.

But back in the 1800s people continued to discuss the possibility that what would become called ergodicity was somehow generic, and would explain why all states would somehow be equally probable, why the Maxwell distribution of velocities would be obtained, and ultimately why the Second Law was true. Maxwell worked out some examples. So did Kelvin. But it remained unclear how it would all work out, as Kelvin (now with many letters after his name) discussed in a talk he gave in 1900 celebrating the new century:

Click to enlarge

The dynamical theory of light didn’t work out. And about the dynamical theory of heat, he quotes Maxwell (following Boltzmann) in one of his very last papers, published in 1878, as saying, in reference to what amounts to a proof of the Second Law from underlying dynamics:

Click to enlarge

Kelvin talks about exploring test cases:

Click to enlarge

When, for example, is the motion of a single particle bouncing around in a fixed region ergodic? He considers first an ellipse, and proves that, no, there isn’t in general ergodicity there:

Click to enlarge

Then he goes on to the much more complicated case

Click to enlarge

and now he does an “experiment” (with a rather Monte Carlo flavor):

Click to enlarge

Kelvin considers a few other examples

Click to enlarge

but mostly concludes that he can’t tell in general about ergodicity—and that probably something else is needed, or as he puts it (somehow wrapping the theory of light into the story as well):

Click to enlarge

But What about Reversibility?

Had Boltzmann’s 1872 H theorem proved the Second Law? Was the Second Law—with its rather downbeat implication about the heat death of the universe—even true? One skeptic was Boltzmann’s friend and former teacher, the chemist Josef Loschmidt (1821–1895), who in 1866 had used kinetic theory to (rather accurately) estimate the size of air molecules. And in 1876 Loschmidt wrote a paper entitled “On the State of Thermal Equilibrium in a System of Bodies with Consideration of Gravity” in which he claimed to show that when gravity was taken into account, there wouldn’t be uniform thermal equilibrium, the Maxwell distribution, or the Second Law—and thus, as he poetically explained:

The terroristic nimbus of the Second Law is destroyed, a nimbus which makes that Second Law appear as the annihilating principle of all life in the universe—and at the same time we are confronted with the comforting perspective that, as far as the conversion of heat into work is concerned, mankind will not solely be dependent on the intervention of coal or of the Sun, but will have available an inexhaustible resource of convertible heat at all times.

His main argument revolves around a thought experiment involving molecules in a gravitational field:

Click to enlarge

Over the next couple of years, despite Loschmidt’s progressively more elaborate constructions

Click to enlarge

Boltzmann and Maxwell will debunk this particular argument—even though to this day the role of gravity in relation to the Second Law remains incompletely resolved.

But what’s more important for our narrative about Loschmidt’s original paper are a couple of paragraphs tucked away at the end of one section (that in fact Kelvin had basically anticipated in 1874):

[Consider what would happen if] after a time t sufficiently long for the stationary state to obtain, we suddenly reversed the velocities of all atoms. Initially we would be in a state that would look like the stationary state. This would be true for some time, but in the long run the stationary state would deteriorate and after the time t we would inevitably return to the initial state…

It is clear that in general in any system one can revert the entire course of events by suddenly inverting the velocities of all the elements of the system. This doesn’t give a solution to the problem of undoing everything that happens [in the universe] but it does give a simple prescription: just suddenly revert the instantaneous velocities of all atoms of the universe.

How did this relate to the H theorem? The underlying molecular equations of motion that Boltzmann had assumed in his proof were reversible in time. Yet Boltzmann claimed that H was always going to a minimum. But why couldn’t one use Loschmidt’s argument to construct an equally possible “reverse evolution” in which H was instead going to a maximum?

It didn’t take Boltzmann long to answer, in print, tucked away in a section of his paper “Remarks on Some Problems in the Mechanical Theory of Heat”. He admits that Loschmidt’s argument “has great seductiveness”. But he claims it is merely “an interesting sophism”—and then says he will “locate the source of the fallacy”. He begins with a classic setup: a collection of hard spheres in a box.

Suppose that at time zero the distribution of spheres in the box is not uniform; for example, suppose that the density of spheres is greater on the right than on the left … The sophism now consists in saying that, without reference to the initial conditions, it cannot be proved that the spheres will become uniformly mixed in the course of time.

But then he rather boldly claims that with the actual initial conditions described, the spheres will “almost always [become] uniform” at a future time t. Now he imagines (following Loschmidt) reversing all the velocities in this state at time t. Then, he says:

… the spheres would sort themselves out as time progresses, and at [the analog of] time 0, they would have a completely nonuniform distribution, even though the [new] initial distribution [one had used] was almost uniform.

But now he says that, yes—given this counterexample—it won’t be possible to prove that the final distribution of spheres will always be uniform.

This is in fact a consequence of probability theory, for any nonuniform distribution, no matter how improbable it may be, is still not absolutely impossible. Indeed it is clear that any individual uniform distribution, which might arise after a certain time from some particular initial state, is just as improbable as an individual nonuniform distribution; just as in the game of Lotto, any individual set of five numbers is as improbable as the set 1, 2, 3, 4, 5. It is only because there are many more uniform distributions than nonuniform ones that the distribution of states will become uniform in the course of time. One therefore cannot prove that, whatever may be the positions and velocities of the spheres at the beginning, the distribution must become uniform after a long time; rather one can only prove that infinitely many more initial states will lead to a uniform one after a definite length of time than to a nonuniform one.

He adds:

One could even calculate, from the relative numbers of the different state distributions, their probabilities, which might lead to an interesting method for the calculation of thermal equilibrium.

And indeed within a few months Boltzmann has followed up on that “interesting method” to produce his classic paper on the probabilistic interpretation of entropy.

But in his earlier paper he goes on to argue:

Since there are infinitely many more uniform than nonuniform distributions of states, the latter case is extraordinarily improbable [to arise] and can be considered impossible for practical purposes; just as it may be considered impossible that if one starts with oxygen and nitrogen mixed in a container, after a month one will find chemically pure oxygen in the lower half and nitrogen in the upper half, although according to probability theory this is merely very improbable but not impossible.

He talks about how interesting it is that the Second Law is intimately connected with probability while the First Law is not. But at the end he does admit:

Perhaps this reduction of the Second Law to the realm of probability makes its application to the entire universe appear dubious, but the laws of probability theory are confirmed by all experiments carried out in the laboratory.

At this point it’s all rather unconvincing. The H theorem had purported to prove the Second Law. But now he’s just talking about probability theory. He seems to have given up on proving the Second Law. And he’s basically just saying that the Second Law is true because it’s observed to be true—like other laws of nature, but not like something that can be “proved”, say from underlying molecular dynamics.

For many years not much attention was paid to these issues, but by the late 1880s there were attempts to clarify things, particularly among the rather active British circle of kinetic theorists. A published 1894 letter from the Irish mathematician Edward Culverwell (1855–1931) (who also wrote about ice ages and Montessori education) summed up some of the confusions that were circulating:

Click to enlarge

At a lecture in England the next year, Boltzmann countered (conveniently, in English):

Click to enlarge

He goes on, but doesn’t get much more specific:

Click to enlarge

He then makes an argument that will be repeated many times in different forms, saying that there will be fluctuations, where H deviates temporarily from its minimum value, but these will be rare:

Click to enlarge

Later he’s talking about what he calls the “H curve” (a plot of H as a function of time), and he’s trying to describe its limiting form:

Click to enlarge

And he even refers to Weierstrass’s recent work on nondifferentiable functions:

Click to enlarge

But he doesn’t pursue this, and instead ends his “rebuttal” with a more philosophical—and in some sense anthropic—argument that he attributes to his former assistant Ignaz Schütz (1867–1927):

Click to enlarge

It’s an argument that we’ll see in various forms repeated over the century and a half that follows. In essence what it’s saying is that, yes, the Second Law implies that the universe will end up in thermal equilibrium. But there’ll always be fluctuations. And in a big enough universe there’ll be fluctuations somewhere that are large enough to correspond to the world as we experience it, where “visible motion and life exist”.

But regardless of such claims, there’s a purely formal question about the H theorem. How exactly is it that from the Boltzmann transport equation—which is supposed to describe reversible mechanical processes—the H theorem manages to prove that the H function irreversibly decreases? It wasn’t until 1895—fully 25 years after Boltzmann first claimed to prove the H theorem—that this issue was even addressed. And it first came up rather circuitously through Boltzmann’s response to comments in a textbook by Gustav Kirchhoff (1824–1887) that had been completed by Max Planck.

The key point is that Boltzmann’s equation makes an implicit assumption, that’s essentially the same as Maxwell made back in 1860: that before each collision between molecules, the molecules are statistically uncorrelated, so that the probability for the collision has the factored form ƒ(v1) ƒ(v2). But what about after the collision? Inevitably the collision itself will lead to correlations. So now there’s an asymmetry: there are no correlations before each collision, but there are correlations after. And that’s why the behavior of the system doesn’t have to be symmetrical—and the H theorem can prove that H irreversibly decreases.

In 1895 Boltzmann wrote a 3-page paper (after half in footnotes) entitled “More about Maxwell’s Distribution Law for Speeds” where he explained what he thought was going on:

[The reversibility of the laws of mechanics] has been recently applied in judging the assumptions necessary for a proof of [the H theorem]. This proof requires the hypothesis that the state of the gas is and remains molecularly disordered, namely, that the molecules of a given class do not always or predominantly collide in a specific manner and that, on the contrary, the number of collisions of a given kind can be found by the laws of probability.

Now, if we assume that in general a state distribution never remains molecularly ordered for an unlimited time and also that for a stationary state-distribution every velocity is as probable as the reversed velocity, then it follows that by inversion of all the velocities after an infinitely long time every stationary state-distribution remains unchanged. After the reversal, however, there are exactly as many collisions occurring in the reversed way as there were collisions occurring in the direct way. Since the two state distributions are identical, the probability of direct and indirect collisions must be equal for each of them, whence follows Maxwell’s distribution of velocities.

Boltzmann is introducing what we’d now call the “molecular chaos” assumption (and what Ehrenfest would call the Stosszahlansatz)—giving a rather self-fulfilling argument for why the assumption should be true. In Boltzmann’s time there wasn’t really anything better to do. By the 1940s the BBGKY hierarchy at least let one organize the hierarchy of correlations between molecules—though it still didn’t give one a tractable way to assess what correlations should exist in practice, and what not.

Boltzmann knew these were all complicated issues. But he wrote about them at a technical level only a few more times in his life. The last time was in 1898 when, responding to a request from the mathematician Felix Klein (1849–1925), he wrote a paper about the H curve for mathematicians. He begins by saying that although this curve comes from the theory of gases, the essence of it can be reproduced by a process based on accumulating balls randomly picked from an urn. He then goes on to outline what amounts to a story of random walks and fractals. In another paper, he actually sketches the curve

Click to enlarge

saying that his drawing “should be taken with a large grain of salt”, noting—in a remarkably fractal-reminiscent way—that “a zincographer [i.e. an engraver of printing plates] would not have been able to produce a real figure since the H-curve has a very large number of maxima and minima on each finite segment, and hence defies representation as a line of continuously changing direction.”

Of course, in modern times it’s easy to produce an approximation to the H curve according to his prescription:

But at the end of his “mathematical” paper he comes back to talking about gases. And first he makes the claim that the effective reversibility seen in the H curve will never be seen in actual physical systems because, in essence, there are always perturbations from outside. But then he ends, in a statement of ultimate reversibility that casts our everyday observation of irreversibility as tautological:

There is no doubt that it is just as conceivable to have a world in which all natural processes take place in the wrong chronological order. But a person living in this upside-down world would have feelings no different than we do: they would just describe what we call the future as the past and vice versa.

The Recurrence Objection

Probably the single most prominent research topic in mathematical physics in the 1800s was the three-body problem—of solving for the motion under gravity of three bodies, such as the Earth, Moon and Sun. And in 1890 the French mathematician Henri Poincaré (1854–1912) (whose breakout work had been on the three-body problem) wrote a paper entitled “On the Three-Body Problem and the Equations of Dynamics” in which, as he said:

It is proved that there are infinitely many ways of choosing the initial conditions such that the system will return infinitely many times as close as one wishes to its initial position. There are also an infinite number of solutions that do not have this property, but it is shown that these unstable solutions can be regarded as “exceptional” and may be said to have zero probability.

This was a mathematical result. But three years later Poincaré wrote what amounted to a philosophy paper entitled “Mechanism and Experience” which expounded on its significance for the Second Law:

In the mechanistic hypothesis, all phenomena must be reversible; for example, the stars might traverse their orbits in the retrograde sense without violating Newton’s law; this would be true for any law of attraction whatever. This is therefore not a fact peculiar to astronomy; reversibility is a necessary consequence of all mechanistic hypotheses.

Experience provides on the contrary a number of irreversible phenomena. For example, if one puts together a warm and a cold body, the former will give up its heat to the latter; the opposite phenomenon never occurs. Not only will the cold body not return to the warm one the heat which it has taken away when it is in direct contact with it; no matter what artifice one may employ, using other intervening bodies, this restitution will be impossible, at least unless the gain thereby realized is compensated by an equivalent or large loss. In other words, if a system of bodies can pass from state A to state B by a certain path, it cannot return from B to A, either by the same path or by a different one. It is this circumstance that one describes by saying that not only is there not direct reversibility, but also there is not even indirect reversibility.

But then he continues:

A theorem, easy to prove, tells us that a bounded world, governed only by the laws of mechanics, will always pass through a state very close to its initial state. On the other hand, according to accepted experimental laws (if one attributes absolute validity to them, and if one is willing to press their consequences to the extreme), the universe tends toward a certain final state, from which it will never depart. In this final state, which will be a kind of death, all bodies will be at rest at the same temperature.

But in fact, he says, the recurrence theorem shows that:

This state will not be the final death of the universe, but a sort of slumber, from which it will awake after millions of millions of centuries. According to this theory, to see heat pass from a cold body to a warm one … it will suffice to have a little patience. [And we may] hope that some day the telescope will show us a world in the process of waking up, where the laws of thermodynamics are reversed.

By 1903, Poincaré was more strident in his critique of the formalism around the Second Law, writing (in English) in a paper entitled “On Entropy”:

Click to enlarge

But back in 1896, Boltzmann and the H theorem had another critic: Ernst Zermelo (1871–1953), a recent German math PhD who was then working with Max Planck on applied mathematics—though would soon turn to foundations of mathematics and become the “Z” in ZFC set theory. Zermelo’s attack on the H theorem began with a paper entitled “On a Theorem of Dynamics and the Mechanical Theory of Heat”. After explaining Poincaré’s recurrence theorem, Zermelo gives some “mathematician-style” conditions (the gas must be in a finite region, must have no infinite energies, etc.), then says that even though there must exist states that would be non-recurrent and could show irreversible behavior, there would necessarily be infinitely more states that “would periodically repeat themselves … with arbitrarily small variations”. And, he argues, such repetition would affect macroscopic quantities discernable by our senses. He continues:

In order to retain the general validity of the Second Law, we therefore would have to assume that just those initial states leading to irreversible processes are realized in nature, their small number notwithstanding, while the other ones, whose probability of existence is higher, mathematically speaking, do not actually occur.

And he concludes that the Poincaré recurrence phenomenon means that:

… it is certainly impossible to carry out a mechanical derivation of the Second Law on the basis of the existing theory without specializing the initial states.

Boltzmann responded promptly but quite impatiently:

I have pointed out particularly often, and as clearly as I possibly could … that the Second Law is but a principle of probability theory as far as the molecular-theoretic point of view is concerned. … While the theorem by Poincaré that Zermelo discusses in the beginning of his paper is of course correct, its application to heat theory is not.

Boltzmann talks about the H curve, and first makes rather a mathematician-style point about the order of limits:

If we first take the number of gas molecules to be infinite, as was clearly done in [my 1896 proof], and only then let the time grow very large, then, in the vast majority of cases, we obtain a curve asymptotically [always close to zero]. Moreover, as can easily be seen, Poincaré’s theorem is not applicable in this case. If, however, we take the time [span] to be infinitely great and, in contrast, the number of molecules to be very great but not absolutely infinite, then the H-curve has a different character. It almost always runs very close to [zero], but in rare cases it rises above that, in what we shall call a “hump” … at which significant deviations from the Maxwell velocity distribution can occur …

Boltzmann then argues that even if you start “at a hump”, you won’t stay there, and “over an enormously long period of time” you’ll see something infinitely close to “equilibrium behavior”. But, he says:

… it is [always] possible to reach again a greater hump of the H-curve by further extending the time … In fact, it is even the case that the original state must return, provided only that we continue to sufficiently extend the time…

He continues:

Mr. Zermelo is therefore right in claiming that, mathematically speaking, the motion is periodic. He has by no means succeeded, however, in refuting my theorems, which, in fact, are entirely consistent with this periodicity.

After giving arguments about the probabilistic character of his results, and (as we would now say it) the fact that a 1D random walk is certain to repeatedly return to the origin, Boltzmann says that:

… we must not conclude that the mechanical approach has to be modified in any way. This conclusion would be justified only if the approach had a consequence that runs contrary to experience. But this would be the case only if Mr. Zermelo were able to prove that the duration of the period within which the old state of the gas must recur in accordance with Poincaré’s theorem has an observable length…

He goes on to imagine “a trillion tiny spheres, each with a [certain initial velocity] … in the one corner of a box” (and by “trillion” he means million million million, or today’s quintillion) and then says that “after a short time the spheres will be distributed fairly evenly in the box”, but the period for a “Poincaré recurrence” in which they all will return to their original corner is “so great that nobody can live to see it happen”. And to make this point more forcefully, Boltzmann has an appendix in which he tries to get an actual approximation to the recurrence time, concluding that its numerical value “has many trillions of digits”.

He concludes:

If we consider heat as a motion of molecules that occurs in accordance with the general equations of mechanics and assume that the arrangement of bodies that we perceive is currently in a highly improbable state, then a theorem follows that is in agreement with the Second Law for all phenomena so far observed.

Of course, this theorem can no longer hold once we observe bodies of so small a scale that they only contain a few molecules. Since, however, we do not have at hand any experimental results on the behavior of bodies so small, this assumption does not run counter to previous experience. In fact, certain experiments conducted on very small bodies in gases seem rather to support the assumption, although we are still far from being able to assert its correctness on the basis of experimental proof.

But then he gives an important caveat—with a small philosophical flourish:

Of course, we cannot expect natural science to answer the question as to why the bodies surrounding us currently exist in a highly improbable state, just as we cannot expect it to answer the question as to why there are any phenomena at all and why they adhere to certain given principles.

Unsurprisingly—particularly in view of his future efforts in the foundations of mathematics—Zermelo is unconvinced by all of this. And six months later he replies again in print. He admits that a full Poincaré recurrence might take astronomically long, but notes that (where, by “physical state”, he means one that we perceive):

… we are after all always concerned only with the “physical state”, which can be realized by many different combinations, and hence can recur much sooner.

Zermelo zeroes in on many of the weaknesses in Boltzmann’s arguments, saying that the thing he particularly “contests … is the analogy that is supposed to exist between the properties of the H curve and the Second Law”. He claims that irreversibility cannot be explained from “mechanical suppositions” without “new physical assumptions”—and in particular criteria for choosing appropriate initial states. He ends by saying that:

From the great successes of the kinetic theory of gases in explaining the relationships among states we must not deduce its … applicability also to temporal processes. … [For in this case I am] convinced that it necessarily fails in the absence of entirely new assumptions.

Boltzmann replies again—starting off with the strangely weak argument:

The Second Law receives a mechanical explanation by virtue of the assumption, which is of course unprovable, that the universe, when considered as a mechanical system, or at least a very extensive part thereof surrounding us, started out in a highly improbable state and still is in such a state.

And, yes, there’s clearly something missing in the understanding of the Second Law. And even as Zermelo pushes for formal mathematician-style clarity, Boltzmann responds with physicist-style “reasonable arguments”. There’s lots of rhetoric:

The applicability of the calculus of probabilities to a particular case can of course never be proved with precision. If 100 out of 100,000 objects of a particular sort are consumed by fire per year, then we cannot infer with certainty that this will also be the case next year. On the contrary! If the same conditions continue to obtain for 1010 years, then it will often be the case during this period that the 100,000 objects are all consumed by fire at once on a single day, and even that not a single object suffers damage over the course of an entire year. Nevertheless, every insurance company places its faith in the calculus of probabilities.

Or, in justification of the idea that we live in a highly improbable “low-entropy” part of the universe:

I refuse to grant the objection that a mental picture requiring so great a number of dead parts of the universe for the explanation of so small a number of animated parts is wasteful, and hence inexpedient. I still vividly remember someone who adamantly refused to believe that the Sun’s distance from the Earth is 20 million miles on the ground that it would simply be foolish to assume so vast a space only containing luminiferous aether alongside so small a space filled with life.

Curiously—given his apparent reliance on “commonsense” arguments—Boltzmann also says:

I myself have repeatedly cautioned against placing excessive trust in the extension of our mental pictures beyond experience and issued reminders that the pictures of contemporary mechanics, and in particular the conception of the smallest particles of bodies as material points, will turn out to be provisional.

In other words, we don’t know that we can think of atoms (even if they exist at all) as points, and we can’t really expect our everyday intuition to tell us about how they work. Which presumably means that we need some kind of solid, “formal” argument if we’re going to explain the Second Law.

Zermelo didn’t respond again, and moved on to other topics. But Boltzmann wrote one more paper in 1897 about “A Mechanical Theorem of Poincaré” ending with two more why-it-doesn’t-apply-in-practice arguments:

Poincaré’s theorem is of course never applicable to terrestrial bodies which we can hold in our hands as none of them is entirely closed. Nor it is applicable to an entirely closed gas of the sort considered by the kinetic theory if first the number of molecules and only then the quotients of the intervals between two neighboring collisions in the observation time is allowed to become infinite.

Ensembles, and an Effort to Make Things Rigorous

Boltzmann—and Maxwell before him—had introduced the idea of using probability theory to discuss the emergence of thermodynamics and potentially the Second Law. But it wasn’t until around 1900—with the work of J. Willard Gibbs (1839–1903)—that a principled mathematical framework for thinking about this developed. And while we can now see that this framework distracts in some ways from several of the key issues in understanding the foundations of the Second Law, it’s been important in framing the discussion of what the Second Law really says—as well as being central in defining the foundations for much of what’s been done over the past century or so under the banner of “statistical mechanics”.

Gibbs seems to have first gotten involved with thermodynamics around 1870. He’d finished his PhD at Yale on the geometry of gears in 1863—getting the first engineering PhD awarded in the US. After traveling in Europe and interacting with various leading mathematicians and physicists, he came back to Yale (where he stayed for the remaining 34 years of his life) and in 1871 became professor of mathematical physics there.

His first papers (published in 1873 when he was already 34 years old) were in a sense based on taking seriously the formalism of equilibrium thermodynamics defined by Clausius and Maxwell—treating entropy and internal energy, just like pressure, volume and temperature, as variables that defined properties of materials (and notably whether they were solids, liquids or gases). Gibbs’s main idea was to “geometrize” this setup, and make it essentially a story of multivariate calculus:

Click to enlarge

Unlike the European developers of thermodynamics, Gibbs didn’t interact deeply with other scientists—with the possible exception of Maxwell, who (a few years before his death in 1879) made a 3D version of Gibbs’s thermodynamic surface out of clay—and supplemented his 2D thermodynamic diagrams after the first edition of his textbook Theory of Heat with renderings of 3D versions:

Click to enlarge

Three years later, Gibbs began publishing what would be a 300-page work defining what has become the standard formalism for equilibrium chemical thermodynamics. He began with a quote from Clausius:

Click to enlarge

In the years that followed, Gibbs’s work—stimulated by Maxwell—mostly concentrated on electrodynamics, and later quaternions and vector analysis. But Gibbs published a few more small papers on thermodynamics—always in effect taking equilibrium (and the Second Law) for granted.

In 1882—a certain Henry Eddy (1844–1921) (who in 1879 had written a book on thermodynamics, and in 1890 would become president of the University of Cincinnati), claimed that “radiant heat” could be used to violate the Second Law:

Click to enlarge

Gibbs soon published a 2-page rebuttal (in the 6th-ever issue of Science magazine):

Click to enlarge

Then in 1889 Clausius died, and Gibbs wrote an obituary—praising Clausius but making it clear he didn’t think the kinetic theory of gases was a solved problem:

Click to enlarge

That same year Gibbs announced a short course that he would teach at Yale on “The a priori Deduction of Thermodynamic Principles from the Theory of Probabilities”. After a decade of work, this evolved into Gibbs’s last publication—an original and elegant book that’s largely defined how the Second Law has been thought about ever since:

Click to enlarge

The book begins by explaining that mechanics is about studying the time evolution of single systems:

Click to enlarge

But Gibbs says he is going to do something different: he is going to look at what he’ll call an ensemble of systems, and see how the distribution of their characteristics changes over time:

Click to enlarge

He explains that these “inquiries” originally arose in connection with deriving the laws of thermodynamics:

Click to enlarge

But he argues that this area—which he’s calling statistical mechanics—is worth investigating even independent of its connection to thermodynamics:

Click to enlarge

Still, he expects this effort will be relevant to the foundations of thermodynamics:

Click to enlarge

He immediately then goes on to what he’ll claim is the way to think about the relation of “observed thermodynamics” to his exact statistical mechanics:

Click to enlarge

Soon he makes the interesting—if, in the light of history, very overly optimistic—claim that “the laws of thermodynamics may be easily obtained from the principles of statistical mechanics”:

Click to enlarge

At first the text of the book reads very much like a typical mathematical work on mechanics:

Click to enlarge

But soon it’s “going statistical”, talking about the “density” of systems in “phase” (i.e. with respect to the variables defining the configuration of the system). And a few pages in, he’s proving the fundamental result that the density of “phase fluid” satisfies a continuity equation (which we’d now call the Liouville equation):

Click to enlarge

It’s all quite elegant, and all very rooted in the calculus-based mathematics of its time. He’s thinking about a collection of instances of a system. But while with our modern computational paradigm we’d readily be able to talk about a discrete list of instances, with his calculus-based approach he has to consider a continuous collection of instances—whose treatment inevitably seems more abstract and less explicit.

He soon makes contact with the “theory of errors”, discussing in effect how probability distributions over the space of possible states evolve. But what probability distributions should one consider? By chapter 4, he’s looking at what he calls (and is still called) the “canonical distribution”:

Click to enlarge

He gives a now-classic definition for the probability as a function of energy ϵ:

Click to enlarge

He observes that this distribution combines nicely when independent parts of a system are brought together, and soon he’s noting that:

Click to enlarge

But so far he’s careful to just talk about how things are “analogous”, without committing to a true connection:

Click to enlarge

More than halfway through the book he’s defined certain properties of his probability distributions that “may … correspond to the thermodynamic notions of entropy and temperature”:

Click to enlarge

Next he’s on to the concept of a “microcanonical ensemble” that includes only states of a given energy. For him—with his continuum-based setup—this is a slightly elaborate thing to define; in our modern computational framework it actually becomes more straightforward than his “canonical ensemble”. Or, as he already says:

Click to enlarge

But what about the Second Law? Now he’s getting a little closer:

Click to enlarge

When he says “index of probability” he’s talking about the log of a probability in his ensemble, so this result is about the fact that this quantity is extremized when all the elements of the ensemble have equal probability:

Click to enlarge

Soon he’s discussing whether he can use his index as a way—like Boltzmann tried to do with his version of entropy—to measure deviations from “statistical equilibrium”:

Click to enlarge

But now Gibbs has hit one of the classic gotchas of his approach: if you look in perfect detail at the evolution of an ensemble of systems, there’ll never be a change in the value of his index—essentially because of the overall conservation of probability. Gibbs brings in what amounts to a commonsense physics argument to handle this. He says to consider putting “coloring matter” in a liquid that one stirs. And then he says that even though the liquid (like his phase fluid) is microscopically conserved, the coloring matter will still end up being “uniformly mixed” in the liquid:

Click to enlarge

He talks about how the conclusion about whether mixing happens in effect depends on what order one takes limits in. And while he doesn’t put it quite this way, he’s essentially realized that there’s a competition between the system “mixing things up more and more finely” and the observer being able to track finer and finer details. He realizes, though, that not all systems will show this kind of mixing behavior, noting for example that there are mechanical systems that’ll just keep going in simple cycles forever.

He doesn’t really resolve the question of why “practical systems” should show mixing, more or less ending with a statement that even though his underlying mechanical systems are reversible, it’s somehow “in practice” difficult to go back:

Click to enlarge

Despite things like this, Gibbs appears to have been keen to keep the majority of his book “purely mathematical”, in effect proving theorems that necessarily followed from the setup he had given. But in the penultimate chapter of the book he makes what he seems to have viewed as a less-than-satisfactory attempt to connect what he’s done with “real thermodynamics”. He doesn’t really commit to the connection, though, characterizing it more as an “analogy”:

Click to enlarge

But he soon starts to be pretty clear that he actually wants to prove the Second Law:

Click to enlarge

He quickly backs off a little, in effect bringing in the observer to soften the requirements:

Click to enlarge

But then he fires his best shot. He says that the quantities he’s defined in connection with his canonical ensemble satisfy the same equations as Clausius originally set up for temperature and entropy:

Click to enlarge

He adds that fluctuations (or “anomalies”, as he calls them) become imperceptible in the limit of a large system:

Click to enlarge

But in physical reality, why should one have a whole collection of systems as in the canonical ensemble? Gibbs suggests it would be more natural to look at the microcanonical ensemble—and in fact to look at a “time ensemble”, i.e. an averaging over time rather than an averaging over different possible states of the system:

Click to enlarge

Gibbs has proved some results (e.g. related to the virial theorem) about the relation between time and ensemble averages. But as the future of the subject amply demonstrates, they’re not nearly strong enough to establish any general equivalence. Still, Gibbs presses on.

In the end, though, as he himself recognized, things weren’t solved—and certainly the canonical ensemble wasn’t the whole story:

Click to enlarge

He discusses the tradeoff between having a canonical ensemble “heat bath” of a known temperature, and having a microcanonical ensemble with known energy. At one point he admits that it might be better to consider the time evolution of a single state, but basically decides that—at least in his continuous-probability-distribution-based formalism—he can’t really set this up:

Click to enlarge

Gibbs definitely encourages the idea that his “statistical mechanics” has successfully “derived” thermodynamics. But he’s ultimately quite careful and circumspect in what he actually says. He mentions the Second Law only once in his whole book—and then only to note that he can get the same “mathematical expression” from his canonical ensemble as Clausius’s form of the Second Law. He doesn’t mention Boltzmann’s H theorem anywhere in the book, and—apart from one footnote concerning “difficulties long recognized by physicists”—he mentions only Boltzmann’s work on theoretical mechanics.

One can view the main achievement of Gibbs’s book as having been to define a framework in which precise results about the statistical properties of collections of systems could be defined and in some cases derived. Within the mathematics and other formalism of the time, such ensemble results represented in a sense a distinctly “higher-order” description of things. Within our current computational paradigm, though, there’s much less of a distinction to be made: whether one’s looking at a single path of evolution, or a whole collection, one’s ultimately still just dealing with a computation. And that makes it clearer that—ensembles or not—one’s thrown back into the same kinds of issues about the origin of the Second Law. But even so, Gibbs provided a language in which to talk with some clarity about many of the things that come up.

Maxwell’s Demon

In late 1867 Peter Tait (1831–1901)—a childhood friend of Maxwell’s who was by then a professor of “natural philosophy” in Edinburgh—was finishing his sixth book. It was entitled Sketch of Thermodynamics and gave a brief, historically oriented and not particularly conceptual outline of what was then known about thermodynamics. He sent a draft to Maxwell, who responded with a fairly long letter:

Click to enlarge

The letter begins:

I do not know in a controversial manner the history of thermodynamics … [and] I could make no assertions about the priority of authors …

Any contributions I could make … [involve] picking holes here and there to ensure strength and stability.

Then he continues (with “ΘΔcs” being his whimsical Greekified rendering of the word “thermodynamics”):

To pick a hole—say in the 2nd law of ΘΔcs, that if two things are in contact the hotter cannot take heat from the colder without external agency.

Now let A and B be two vessels divided by a diaphragm … Now conceive a finite being who knows the paths and velocities of all the molecules by simple inspection but who can do no work except open and close a hole in the diaphragm by means of a slide without mass. Let him … observe the molecules in A and when he sees one coming … whose velocity is less than the mean [velocity] of the molecules in B let him open the hole and let it go into B [and vice versa].

Then the number of molecules in A and B are the same as at first, but the energy in A is increased and that in B diminished, that is, the hot system has got hotter and the cold colder and yet no work has been done, only the intelligence of a very observant and neat-fingered being has been employed.

Or in short [we can] … restore a uniformly hot system to unequal temperatures… Only we can’t, not being clever enough.

And so it was that the idea of “Maxwell’s demon” was launched. Tait must at some point have shown Maxwell’s letter to Kelvin, who wrote on it:

Very good. Another way is to reverse the motion of every particle of the Universe and to preside over the unstable motion thus produced.

But the first place Maxwell’s demon idea appeared in print was in Maxwell’s 1871 textbook Theory of Heat:

Click to enlarge

Much of the book is devoted to what was by then quite traditional, experimentally oriented thermodynamics. But Maxwell included one final chapter:

Click to enlarge

Even in 1871, after all his work on kinetic theory, Maxwell is quite circumspect in his discussion of molecules:

Click to enlarge

But Maxwell’s textbook goes through a series of standard kinetic theory results, much as a modern textbook would. The second-to-last section in the whole book sounds a warning, however:

Click to enlarge

Interestingly, Maxwell continues, somewhat in anticipation of what Gibbs will say 30 years later:

Click to enlarge

But then there’s a reminder that this is being written in 1871, several decades before any clear observation of molecules was made. Maxwell says:

Click to enlarge

In other words, if there are water molecules, there must be something other than a law of averages that makes them all appear the same. And, yes, it’s now treated as a fundamental fact of physics that, for example, all electrons have exactly—not just statistically—the same properties such as mass and charge. But back in 1871 it was much less clear what characteristics molecules—if they existed as real entities at all—might have.

Maxwell included one last section in his book that to us today might seem quite wild:

Click to enlarge

In other words, aware of Darwin’s (1809–1882) 1859 Origin of Species, he’s considering a kind of “speciation” of molecules, along the lines of the discrete species observed in biology. But then he notes that unlike biological organisms, molecules are “permanent”, so their “selection” must come from some kind of pure separation process:

Click to enlarge

And at the very end he suggests that if molecules really are all identical, that suggests a level of fundamental order in the world that we might even be able to flow through to “exact principles of distributive justice” (presumably for people rather than molecules):

Click to enlarge

Maxwell has described rather clearly his idea of demons. But the actual name “demon” first appears in print in a paper by Kelvin in 1874:

Click to enlarge

It’s a British paper, so—in a nod to future nanomachinery—it’s talking about (molecular) cricket bats:

Click to enlarge

Kelvin’s paper—like his note written on Maxwell’s letter—imagines that the demons don’t just “sort” molecules; they actually reverse their velocities, thus in effect anticipating Loschmidt’s 1876 “reversibility objection” to Boltzmann’s H theorem.

In an undated note, Maxwell discusses demons, attributing the name to Kelvin—and then starts considering the “physicalization” of demons, simplifying what they need to do:

Concerning Demons.
I. Who gave them this name? Thomson.
2. What were they by nature? Very small BUT lively beings incapable of doing work but able to open and shut valves which move without friction or inertia.
3. What was their chief end? To show that the 2nd Law of Thermodynamics has only a statistical certainty.
4. Is the production of an inequality of temperature their only occupation? No, for less intelligent demons can produce a difference in pressure as well as temperature by merely allowing all particles going in one direction while stopping all those going the other way. This reduces the demon to a valve. As such value him. Call him no more a demon but a valve like that of the hydraulic ram, suppose.

It didn’t take long for Maxwell’s demon to become something of a fixture in expositions of thermodynamics, even if it wasn’t clear how it connected to other things people were saying about thermodynamics. And in 1879, for example, Kelvin gave a talk all about Maxwell’s “sorting demon” (like other British people of the time he referred to Maxwell as “Clerk Maxwell”):

Click to enlarge

Kelvin describes—without much commentary, and without mentioning the Second Law—some of the feats of which the demon would be capable. But he adds:

Click to enlarge

The description of the lecture ends:

Click to enlarge

Presumably no actual Maxwell’s demon was shown—or Kelvin wouldn’t have continued for the rest of his life to treat the Second Law as an established principle.

But in any case, Maxwell’s demon has always remained something of a fixture in discussions of the foundations of the Second Law. One might think that the observability of Brownian motion would make something like a Maxwell’s demon possible. And indeed in 1912 Marian Smoluchowski (1872–1917) suggested experiments that one could imagine would “systematically harvest” Brownian motion—but showed that in fact they couldn’t. In later years, a sequence of arguments were advanced that the mechanism of a Maxwell’s demon just couldn’t work in practice—though even today microscopic versions of what amount to Maxwell’s demons are routinely being investigated.

What Happened to Those People?

We’ve finally now come to the end of the story of how the original framework for the Second Law came to be set up. And, as we’ve seen, only a fairly small number of key players were involved:

So what became of these people? Carnot lived a generation earlier than the others, never made a living as a scientist, and was all but unknown in his time. But all the others had distinguished careers as academic scientists, and were widely known in their time. Clausius, Boltzmann and Gibbs are today celebrated mainly for their contributions to thermodynamics; Kelvin and Maxwell also for other things. Clausius and Gibbs were in a sense “pure professors”; Boltzmann, Maxwell and especially Kelvin also had engagement with the more general public.

All of them spent the majority of their lives in the countries of their birth—and all (with the exception of Carnot) were able to live out the entirety of their lives without time-consuming disruptions from war or other upheavals:

Sadi Carnot (1796–1832)

Almost all of what is known about Sadi Carnot as a person comes from a single biographical note written nearly half a century after his death by his younger brother Hippolyte Carnot (who was a distinguished French politician—and sometime education minister—and father of the Sadi Carnot who would become president of France). Hippolyte Carnot began by saying that:

As the life of Sadi Carnot was not marked by any notable event, his biography would have occupied only a few lines; but a scientific work by him, after remaining long in obscurity, brought again to light many years after his death, has caused his name to be placed among those of great inventors.

The Carnots’ father was close to Napoleon, and Hippolyte explains that when Sadi was a young child he ended up being babysat by “Madame Bonaparte”—but one day wandered off, and was found inspecting the operation of a nearby mill, and quizzing the miller about it. For the most part, however, throughout his life, Sadi Carnot apparently kept very much to himself—while with quiet intensity showing a great appetite for intellectual pursuits from mathematics and science to art, music and literature, as well as practical engineering and the science of various sports.

Even his brother Hippolyte can’t explain quite how Sadi Carnot—at the age of 28—suddenly “came out” and in 1824 published his book on thermodynamics. (As we discussed above, it no doubt had something to do with the work of his father, who died two years earlier.) Sadi Carnot funded the publication of the book himself—having 600 copies printed (at least some of which remained unsold a decade later). But after the book was published, Carnot appears to have returned to just privately doing research, living alone, and never publishing again in his lifetime. And indeed he lived only another eight years, dying (apparently after some months of ill health) in the same Paris cholera outbreak that claimed General Lamarque of Les Misérables fame.

Twenty-three pages of unpublished personal notes survive from the period after the publication of Carnot’s book. Some are general aphorisms and life principles:

Speak little of what you know, and not at all of what you do not know.

Why try to be witty? I would rather be thought stupid and modest than witty and pretentious.

God cannot punish man for not believing when he could so easily have enlightened and convinced him.

The belief in an all-powerful Being, who loves us and watches over us, gives to the mind great strength to endure misfortune.

When walking, carry a book, a notebook to preserve ideas, and a piece of bread in order to prolong the walk if need be.

But others are more technical—and in fact reveal that Carnot, despite having based his book on caloric theory, had realized that it probably wasn’t correct:

When a hypothesis no longer suffices to explain phenomena, it should be abandoned. This is the case with the hypothesis which regards caloric as matter, as a subtile fluid.

The experimental facts tending to destroy this theory are as follows: The development of heat by percussion or friction of bodies … The elevation of temperature which takes place [when] air [expands into a] vacuum …

He continues:

At present, light is generally regarded as the result of a vibratory movement of the ethereal fluid. Light produces heat, or at least accompanies radiating heat, and moves with the same velocity as heat. Radiating heat is then a vibratory movement. It would be ridiculous to suppose that it is an emission of matter while the light which accompanies it could be only a movement.

Could a motion (that of radiating heat) produce matter (caloric)? No, undoubtedly; it can only produce a motion. Heat is then the result of a motion.

And then—in a rather clear enunciation of what would become the First Law of thermodynamics:

Heat is simply motive power, or rather motion which has changed form. It is a movement among the particles of bodies. Wherever there is destruction of motive power there is, at the same time, production of heat in quantity exactly proportional to the quantity of motive power destroyed. Reciprocally, wherever there is destruction of heat, there is production of motive power.

Carnot also wonders:

Liquefaction of bodies, solidification of liquids, crystallization—are they not forms of combinations of integrant molecules? Supposing heat due to a vibratory movement, how can the passage from the solid or the liquid to the gaseous state be explained?

There is no indication of how Carnot felt about this emerging rethinking of thermodynamics, or of how it might affect the results in his book. But Carnot clearly hoped to do experiments (as outlined in his notes) to test what was really going on. But as it was, he presumably didn’t get around to any of them—and his notes, ahead of their time as they were, did not resurface for many decades, by which time the ideas they contained had already been discovered by others.

Rudolf Clausius (1822–1888)

Rudolf Clausius was born in what’s now Poland (and was then Prussia), one of more than 14 children of an education administrator and pastor. He went to university in Berlin, and, after considering doing history, eventually specialized in math and physics. After graduating in 1844 he started teaching at a top high school in Berlin (which he did for 6 years), and meanwhile earned his PhD in physics. His career took off after his breakout paper on thermodynamics appeared in 1850. For a while he was a professor in Berlin, then for 12 years in Zürich, then briefly in Würzburg, then—for the remaining 19 years of his life—in Bonn.

He was a diligent—if, one suspects, somewhat stiff—professor, notable for the clarity of his lectures, and his organizational care with students. He seems to have been a competent administrator, and late in his career he spent a couple of years as the president (“rector”) of his university. But first and foremost, he was a researcher, writing about a hundred papers over the course of his career. Most physicists of the time devoted at least some of their efforts to doing actual physics experiments. But Clausius was a pioneer in the idea of being a “pure theoretical physicist”, inspired by experiments and quoting their results, but not doing them himself.

The majority of Clausius’s papers were about thermodynamics, though late in his career his emphasis shifted more to electrodynamics. Clausius’s papers were original, clear, incisive and often fairly mathematically sophisticated. But from his very first paper on thermodynamics in 1850, he very much adopted a macroscopic approach, talking about what he considered to be “bulk” quantities like energy, and later entropy. He did explore some of the potential mechanics of molecules, but he never really made the connection between molecular phenomena and entropy—or the Second Law. He had a number of run-ins about academic credit with Kelvin, Tait, Maxwell and Boltzmann, but he didn’t seem to ever pay much attention to, for example, Boltzmann’s efforts to find molecular-based probabilistic derivations of Clausius’s results.

It probably didn’t help that after two decades of highly productive work, two misfortunes befell Clausius. First, in 1870, he had volunteered to lead an ambulance corps in the Franco-Prussian war, and was wounded in the knee, leading to chronic pain (as well as to his habit of riding to class on horseback). And then, in 1875, Clausius’s wife died in the birth of their sixth child—leaving him to care for six young children (which apparently he did with great conscientiousness). Clausius nevertheless continued to pursue his research—even to the end of his life—receiving many honors along the way (like election to no less than 40 professional societies), but it never again rose to the level of significance of his early work on thermodynamics and the Second Law.

Kelvin (William Thomson) (1824–1907)

Of the people we’re discussing here, by far the most famous during their lifetime was Kelvin. In his long career he wrote more than 600 scientific papers, received dozens of patents, started several companies and served in many administrative and governmental roles. His father was a math professor, ultimately in Glasgow, who took a great interest in the education of his children. Kelvin himself got an early start, effectively going to college at the age of 10, and becoming a professor in Glasgow at the age of 22—a position in which he continued for 53 years.

Kelvin’s breakout work, done in his twenties, was on thermodynamics. But over the years he also worked on many other areas of physics, and beyond, mixing theory, experiment and engineering. Beginning in 1854 he became involved in a technical megaproject of the time: the attempt to lay a transatlantic telegraph cable. He wound up very much on the front lines, helping out as a just-in-time physicist + engineer on the cable-laying ship. The first few attempts didn’t work out, but finally in 1866—in no small part through Kelvin’s contributions—a cable was successfully laid, and Kelvin (or William Thomson, as he then was) became something of a celebrity. He was made “Sir William Thomson” and—along with two other techies—formed his first company, which had considerable success in exploiting telegraph-cable-related engineering innovations.

Kelvin’s first wife died after a long illness in 1870, and Kelvin, with no children and already enthusiastic about the sea, bought a fairly large yacht, and pursued a number of nautical-related projects. One of these—begun in 1872—was the construction of an analog computer for calculating tides (basically with 10 gears for adding up 10 harmonic tide components), a device that, with progressive refinements, continued to be used for close to a century.

Being rather charmed by Kelvin’s physicist-with-a-big-yacht persona, I once purchased a letter that Kelvin wrote in 1877 on the letterhead of “Yacht Lalla Rookh”:

Click to enlarge

The letter—in true academic style—promises that Kelvin will soon send an article he’s been asked to write on elasticity theory. And in fact he did write the article, and it was an expository one that appeared in the 9th edition of the Encyclopedia Britannica.

Kelvin was a prolific (if, to modern ears, sometimes rather pompous) writer, who took exposition seriously. And indeed—finding the textbooks available to him as a professor inadequate—he worked over the course of a dozen years (1855–1867) with his (and Maxwell’s) friend Peter Guthrie Tait to produce the influential Treatise on Natural Philosophy.

Kelvin explored many topics and theories, some more immediately successful than others. In the 1870s he suggested that perhaps atoms might be knotted vortices in the (luminiferous) aether (causing Tait to begin developing knot theory)—a hypothesis that’s in some sense a Victorian prelude to modern ideas about particles in our Physics Project.

Throughout his life, Kelvin was a devout Christian, writing that “The more thoroughly I conduct scientific research, the more I believe science excludes atheism.” And indeed this belief seems to make an appearance in his implication that humans—presumably as a result of their special relationship with God—might avoid the Second Law. But more significant at the time was Kelvin’s skepticism about Charles Darwin’s 1859 theory of natural selection, believing that there must in the end be a “continually guiding and controlling intelligence”. Despite being somewhat ridiculed for it, Kelvin talked about the possibility that life might have come to Earth from elsewhere via meteorites, believing that his estimates of the age of the Earth (which didn’t take into account radioactivity) made it too young for the things Darwin described to have occurred.

By the 1870s, Kelvin had become a distinguished man of science, receiving all sorts of honors, assignments and invitations. And in 1876, for example, he was invited to Philadelphia to chair the committee judging electrical inventions at the US Centennial International Exhibition, notably reporting, in the terms of the time:

Click to enlarge

Then in 1892 a “peerage of the realm” was conferred on him by Queen Victoria. His wife (he had remarried) and various friends (including Charles Darwin’s son George) suggested he pick the title “Kelvin”, after the River Kelvin that flowed by the university in Glasgow. And by the end of his life “Lord Kelvin” had accumulated enough honorifics that they were just summarized with “…” (the MD was an honorary degree conferred by the University of Heidelberg because “it was the only one at their disposal which he did not already possess”):

Click to enlarge

And when Kelvin died in 1907 he was given a state funeral and buried in Westminster Abbey near Newton and Darwin.

James Clerk Maxwell (1831–1879)

James Clerk Maxwell lived only 48 years but in that time managed to do a remarkable amount of important science. His early years were spent on a 1500-acre family estate (inherited by his father) in a fairly remote part of Scotland—to which he would return later. He was an only child and was homeschooled—initially by his mother, until she died, when he was 8. At 10 he went to an upscale school in Edinburgh, and by the age of 14 had written his first scientific paper. At 16 he went as an undergraduate to the University of Edinburgh, then, effectively as a graduate student, to Cambridge—coming second in the final exams (“Second Wrangler”) to a certain Edward Routh, who would spend most of his life coaching other students on those very same exams.

Within a couple of years, Maxwell was a professor, first in Aberdeen, then in London. In Aberdeen he married the daughter of the university president, who would soon be his “Observer K” (for “Katherine”) in his classic work on color vision. But after nine fairly strenuous years as a professor, Maxwell in 1865 “retired” to his family estate, supervising a house renovation, and in “rural solitude” (recreationally riding around his estate on horseback with his wife) having the most scientifically productive time of his life. In addition to his work on things like the kinetic theory of gases, he also wrote his 2-volume Treatise on Electricity and Magnetism, which ultimately took 7 years to finish, and which, with considerable clarity, described his approach to electromagnetism and what are now called “Maxwell’s Equations”. Occasionally, there were hints of his “country life”—like his 1870 “On Hills and Dales” that in his characteristic mathematicize-everything way gave a kind of “pre-topological” analysis of contour maps (perhaps conceived as he walked half a mile every day down to the mailbox at which journals and correspondence would arrive):

Click to enlarge

As a person, Maxwell was calm, reserved and unassuming, yet cheerful and charming—and given to writing (arguably sometimes sophomoric) poetry:

Click to enlarge

With a certain sense of the absurd, he would occasionally publish satirical pieces in Nature, signing them dp/dt, which in the thermodynamic notation created by his friend Tait was equal to JCM, which were his initials. Maxwell liked games and tricks, and spinning tops featured prominently in some of his work. He enjoyed children, though never had any of his own. As a lecturer, he prepared diligently, but often got too sophisticated for his audience. In writing, though, he showed both great clarity and great erudition, for example freely quoting Latin and Greek in articles he wrote for the 9th edition of the Encyclopedia Britannica (of which he was scientific co-editor) on topics such as “Atom” and “Ether”.

As we mentioned above, Maxwell was quite an enthusiast of diagrams and visual presentation (even writing an article on “Diagrams” for the Encyclopedia Britannica). He was also a capable experimentalist, making many measurements (sometimes along with his wife), and in 1861 creating the first color photograph.

In 1871 William Cavendish, 7th Duke of Devonshire, who had studied math in Cambridge, and was now chancellor of the university, agreed to put up the money to build what became the Cavendish Laboratory and to endow a new chair of experimental physics. Kelvin having turned down the job, it was offered to the still-rather-obscure Maxwell, who somewhat reluctantly accepted—with the result that for several years he spent much of his time supervising the design and building of the lab.

The lab was finished in 1874, but then William Cavendish dropped on Maxwell a large collection of papers from his great uncle Henry Cavendish, who had been a wealthy “gentleman scientist” of the late 1700s and (among other things) the discoverer of hydrogen. Maxwell liked history (as some of us do!), noticed that Cavendish had discovered Ohm’s law 50 years before Ohm, and in the end spent several years painstakingly editing and annotating the papers into a 500-page book. By 1879 Maxwell was finally ready to energetically concentrate on physics research again, but, sadly, in the fall of that year his health failed, and he died at the age of 48—having succumbed to stomach cancer, as his mother also had at almost the same age.

J. Willard Gibbs (1839–1903)

Gibbs was born near the Yale campus, and died there 64 years later, in the same house where he had lived since he was 7 years old (save for three years spent visiting European universities as a young man, and regular summer “out-in-nature” vacations). His father (who, like, “our Gibbs” was named “Josiah Willard”—making “our Gibbs” be called “Willard”) came from an old and distinguished intellectual and religious New England family, and was a professor of sacred languages at Yale. Willard Gibbs went to college and graduate school at Yale, and then spent his whole career as a professor at Yale.

He was, it seems, a quiet, modest and rather distant person, who radiated a certain serenity, regularly attended church, had a small circle of friends and lived with his two sisters (and the husband and children of one of them). He diligently discharged his teaching responsibilities, though his lectures were very sparsely attended, and he seems not to have been thought forceful enough in dealing with people to have been called on for many administrative tasks—though he became the treasurer of his former high school, and himself was careful enough with money that by the end of his life he had accumulated what would now be several million dollars.

He had begun his academic career in practical engineering, for example patenting an “improved [railway] car-brake”, but was soon drawn in more mathematical directions, favoring a certain clarity and minimalism of formulation, and a cleanliness, if not brevity, of exposition. His work on thermodynamics (initially published in the rather obscure Transactions of the Connecticut Academy) was divided into two parts: the first, in the 1870s, concentrating on macroscopic equilibrium properties, and second, in the 1890s, concentrating on microscopic “statistical mechanics” (as Gibbs called it). Even before he started on thermodynamics, he’d been interested in electromagnetism, and between his two “thermodynamic periods”, he again worked on electromagnetism. He studied Maxwell’s work, and was at first drawn to the then-popular formalism of quaternions—but soon decided to invent his own approach and notation for vector analysis, which at first he presented only in notes for his students, though it later became widely adopted.

And while Gibbs did increasingly mathematical work, he never seems to have identified as a mathematician, modestly stating that “If I have had any success in mathematical physics, it is, I think, because I have been able to dodge mathematical difficulties.” His last work was his book on statistical mechanics, which—with considerable effort and perhaps damage to his health—he finished in time for publication in connection with the Yale bicentennial in 1901 (an event which notably also brought a visit from Kelvin), only to die soon thereafter.

Gibbs had a few graduate students at Yale, a notable one being Lee de Forest, inventor of the vacuum tube (triode) electronic amplifier, and radio entrepreneur. (de Forest’s 1899 PhD thesis was entitled “Reflection of Hertzian Waves from the Ends of Parallel Wires”.) Another student of Gibbs was Lynde Wheeler, who became a government radio scientist, and who wrote a biography of Gibbs, of which I have a copy bought years ago at a used bookstore—that I was now just about to put back on a shelf when I opened its front cover and found an inscription:

Click to enlarge

And, yes, it’s a small world, and “To Willard” refers to Gibbs’s sister’s son (Willard Gibbs Van Name, who became a naturalist and wrote a 1929 book about national park deforestation).

Ludwig Boltzmann (1844–1906)

Of the people we’re discussing, Boltzmann is the one whose career was most focused on the Second Law. Boltzmann grew up in Austria, where his father was a civil servant (who died when Boltzmann was 15) and his mother was something of an heiress. Boltzmann did his PhD at the University of Vienna, where his professor notably gave him a copy of some of Maxwell’s papers, together with an English grammar book. Boltzmann started publishing his own papers near the end of his PhD, and soon landed a position as a professor of mathematical physics in Graz. Four years later he moved to Vienna as a professor of mathematics, soon moving back to Graz as a professor of “general and experimental physics”—a position he would keep for 14 years.

He’d married in 1876, and had 5 children, though a son died in 1889, leaving 3 daughters and another son. Boltzmann was apparently a clear and lively lecturer, as well as a spirited and eager debater. He seems, at least in his younger years, to have been a happy and gregarious person, with a strong taste for music—and some charming do-it-your-own-way tendencies. For example, wanting to provide fresh milk for his children, he decided to just buy a cow, which he then led from the market through the streets—though had to consult his colleague, the professor of zoology, to find out how to milk it. Boltzmann was a capable experimental physicist, as well as a creator of gadgets, and a technology enthusiast—promoting the idea of airplanes (an application for gas theory!) and noting their potential power as a means of transportation.

Boltzmann had always had mood swings, but by the early 1890s he claimed they were getting worse. It didn’t help that he was worn down by administrative work, and had worsening asthma and increasing nearsightedness (that he’d thought might be a sign of going blind). He moved positions, but then came back to Vienna, where he embarked on writing what would become a 2-volume book on Gas Theory—in effect contextualizing his life’s work. The introduction to the first volume laments that “gas theory has gone out of fashion in Germany”. The introduction to the second volume, written in 1898 when Boltzmann was 54, then says that “attacks on the theory of gases have begun to increase”, and continues:

… it would be a great tragedy for science if the theory of gases were temporarily thrown into oblivion because of a momentary hostile attitude toward it, as, for example, was the wave theory [of light] because of Newton’s authority.

I am conscious of being only an individual struggling weakly against the stream of time. But it still remains in my power to contribute in such a way that, when the theory of gases is again revived, not too much will have to be rediscovered.

But even as he was writing this, Boltzmann had pretty much already wound down his physics research, and had basically switched to exposition, and to philosophy. He moved jobs again, but in 1902 again came back to Vienna, but now also as a professor of philosophy. He gave an inaugural lecture, first quoting his predecessor Ernst Mach (1838–1916) as saying “I do not believe that atoms exist”, then discussing the philosophical relations between reality, perception and models. Elsewhere he discussed things like his view of the different philosophical character of models associated with differential equations and with atomism—and he even wrote an article on the general topic of “Models” for Encyclopedia Britannica (which curiously also talks about “in pure mathematics, especially geometry, models constructed of papier-mâché and plaster”). Sometimes Boltzmann’s philosophy could be quite polemical, like his attack on Schopenhauer, that ends by saying that “men [should] be freed from the spiritual migraine that is called metaphysics”.

Then, in 1904, Boltzmann addressed the Vienna Philosophical Society (a kind of predecessor of the Vienna Circle) on the subject of a “Reply to a Lecture on Happiness by Professor Ostwald”. Wilhelm Ostwald (1853–1932) (a chemist and social reformer, who was a personal friend of Boltzmann’s, but intellectual adversary) had proposed the concept of “energy of will” to apply mathematical physics ideas to psychology. Boltzmann mocked this, describing its faux formalism as “dangerous for science”. Meanwhile, Boltzmann gives his own Darwinian theory for the origin of happiness, based essentially on the idea that unhappiness is needed as a way to make organisms improve their circumstances in the struggle for survival.

Boltzmann himself was continuing to have problems that he attributed to then-popular but very vague “diagnosis” of “neurasthenia”, and had even briefly been in a psychiatric hospital. But he continued to do things like travel. He visited the US three times, in 1905 going to California (mainly Berkeley)—which led him to write a witty piece entitled “A German Professor’s Trip to El Dorado” that concluded:

Yes, America will achieve great things. I believe in these people, even after seeing them at work in a setting where they’re not at their best: integrating and differentiating at a theoretical physics seminar…

In 1905 Einstein published his Boltzmann-and-atomism-based results on Brownian motion and on photons. But it’s not clear Boltzmann ever knew about them. For Boltzmann was sinking further. Perhaps he’d overexerted himself in California, but by the spring of 1906 he said he was no longer able to teach. In the summer he went with his family to an Italian seaside resort in an attempt to rejuvenate. But a day before they were to return to Vienna he failed to join his family for a swim, and his youngest daughter found him hanged in his hotel room, dead at the age of 62.

Coarse-Graining and the “Modern Formulation”

After Gibbs’s 1902 book introducing the idea of ensembles, most of the language used (at least until now!) to discuss the Second Law was basically in place. But in 1912 one additional term—representing a concept already implicit in Gibbs’s work—was added: coarse-graining. Gibbs had discussed how the phase fluid representing possible states of a system could be elaborately mixed by the mechanical time evolution of the system. But realistic practical measurements could not be expected to probe all the details of the distribution of phase fluid; instead one could say that they would only sample “coarse-grained” aspects of it.

The term “coarse-graining” first appeared in a survey article entitled “The Conceptual Foundations of the Statistical Approach in Mechanics”, written for the German-language Encyclopaedia of the Mathematical Sciences by Boltzmann’s former student Paul Ehrenfest, and his wife Tatiana Ehrenfest-Afanassjewa:

Click to enlarge

The article also introduced all sorts of now-standard notation, and in many ways can be read as a final summary of what was achieved in the original development around the foundations of thermodynamics and the Second Law. (And indeed the article was sufficiently “final” that when it was republished as a book in 1959 it could still be presented as usefully summarizing the state of things.)

Looking at the article now, though, it’s notable how much it recognized was not at all settled about the Second Law and its foundations. It places Boltzmann squarely at the center, stating in its preface:

Click to enlarge

The section titles are already revealing:

Click to enlarge

And soon they’re starting to talk about “loose ends”, and lots of them. Ergodicity is something one can talk about, but there’s no known example (and with this definition it was later proved that there couldn’t be):

Click to enlarge

But, they point out, it’s something Boltzmann needed in order to justify his results:

Click to enlarge

Soon they’re talking about Boltzmann’s sloppiness in his discussion of the H curve:

Click to enlarge

And then they’re on to talking about Gibbs, and the gaps in his reasoning:

Click to enlarge

In the end they conclude:

Click to enlarge

In other words, even though people now seem to be buying all these results, there are still plenty of issues with their foundations. And despite people’s implicit assumptions, we can in no way say that the Second Law has been “proved”.

Radiant Heat, the Second Law and Quantum Mechanics

It was already realized in the 1600s that when objects get hot they emit “heat radiation”—which can be transferred to other bodies as “radiant heat”. And particularly following Maxwell’s work in the 1860s on electrodynamics it came to be accepted that radiant heat was associated with electromagnetic waves propagating in the “luminiferous aether”. But unlike the molecules from which it was increasingly assumed that one could think of matter as being made, these electromagnetic waves were always treated—particularly on the basis of their mathematical foundations in calculus—as fundamentally continuous.

But how might this relate to the Second Law? Could it be, perhaps, that the Second Law should ultimately be attributed not to some property of the large-scale mechanics of discrete molecules, but rather to a feature of continuous radiant heat?

The basic equations assumed for mechanics—originally due to Newton—are reversible. But what about the equations for electrodynamics? Maxwell’s equations are in and of themselves also reversible. But when one thinks about their solutions for actual electromagnetic radiation, there can be fundamental irreversibility. And the reason is that it’s natural to describe the emission of radiation (say from a hot body), but then to assume that, once emitted, the radiation just “escapes to infinity”—rather than ever reversing the process of emission by being absorbed by some other body.

All the various people we’ve discussed above, from Clausius to Gibbs, made occasional remarks about the possibility that the Second Law—whether or not it could be “derived mechanically”—would still ultimately work, if nothing else, because of the irreversible emission of radiant heat.

But the person who would ultimately be most intimately connected to these issues was Max Planck—though in the end the somewhat-confused connection to the Second Law would recede in importance relative to what emerged from it, which was basically the raw material that led to quantum theory.

As a student of Helmholtz’s in Berlin, Max Planck got interested in thermodynamics, and in 1879 wrote a 61-page PhD thesis entitled “On the Second Law of Mechanical Heat Theory”. It was a traditional (if slightly streamlined) discussion of the Second Law, very much based on Clausius’s approach (and even with the same title as Clausius’s 1867 paper)—and without any mention whatsoever of Boltzmann:

Click to enlarge

For most of the two decades that followed, Planck continued to use similar methods to study the Second Law in various settings (e.g. elastic materials, chemical mixtures, etc.)—and meanwhile ascended the German academic physics hierarchy, ending up as a professor of theoretical physics in Berlin. Planck was in many ways a physics traditionalist, not wanting to commit to things like “newfangled” molecular ideas—and as late as 1897 (with his assistant Zermelo having made his “recurrence objection” to Boltzmann’s work) still saying that he would “abstain completely from any definite assumption about the nature of heat”. But regardless of its foundations, Planck was a true believer in the Second Law, for example in 1891 asserting that it “must extend to all forces of nature … not only thermal and chemical, but also electrical and other”.

And in 1895 he began to investigate how the Second Law applied to electrodynamics—and in particular to the “heat radiation” that it had become clear (particularly through Heinrich Hertz’s (1857–1894) experiments) was of electromagnetic origin. In 1896 Wilhelm Wien (1864–1928) suggested that the heat radiation (or what we now call blackbody radiation) was in effect produced by tiny Hertzian oscillators with velocities following a Maxwell distribution.

Planck, however, had a different viewpoint, instead introducing the concept of “natural radiation”—a kind of intrinsic thermal equilibrium state for radiation, with an associated intrinsic entropy. He imagined “resonators” interacting through Maxwell’s equations with this radiation, and in 1899 invented a (rather arbitrary) formula for the entropy of these resonators, that implied (through the laws of electrodynamics) that overall entropy would increase—just like the Second Law said—and when the entropy was maximized it gave the same result as Wien for the spectrum of blackbody radiation. In early 1900 he sharpened his treatment and began to suggest that with his approach Wien’s form of the blackbody spectrum would emerge as a provable consequence of the universal validity of the Second Law.

But right around that time experimental results arrived that disagreed with Wien’s law. And by the end of 1900 Planck had a new hypothesis, for which he finally began to rely on ideas from Boltzmann. Planck started from the idea that he should treat the behavior of his resonators statistically. But how then could he compute their entropy? He quotes (for the first time ever) his simplification of Boltzmann’s formula for entropy:

Click to enlarge

As he explains it—claiming now, after years of criticizing Boltzmann, that this is a “theorem”:

We now set the entropy S of the system proportional to the logarithm of its probability W… In my opinion this actually serves as a definition of the probability W, since in the basic assumptions of electromagnetic theory there is no definite evidence for such a probability. The suitability of this expression is evident from the outset, in view of its simplicity and close connection with a theorem from kinetic gas theory.

But how could he figure out the probability for a resonator to have a certain energy, and thus a certain entropy? For this he turns directly to Boltzmann—who, as a matter of convenience in his 1877 paper had introduced discrete values of energy for molecules. Planck simply states that it’s “necessary” (i.e. to get the experimentally right answer) to treat the resonator energy “not as a continuous, infinitely divisible quantity, but as a discrete quantity composed of an integral number of finite equal parts”. As an example of how this works he gives a table just like the one in Boltzmann’s paper from nearly a quarter of a century earlier:

Click to enlarge

Pretty soon he’s deriving the entropy of a resonator as a function of its energy, and its discrete energy unit ϵ:

Click to enlarge

Connecting this to blackbody radiation he claims that each resonator’s energy unit is connected to its frequency according to

Click to enlarge

so that its entropy is

Click to enlarge

“[where] h and k are universal constants”.

In a similar situation Boltzmann had effectively taken the limit ϵ→0, because that’s what he believed corresponded to (“calculus-based”) physical reality. But Planck—in what he later described as an “act of desperation” to fit the experimental data—didn’t do that. So in computing things like average energies he’s evaluating Sum[x Exp[-a x], {x, 0, ∞}] rather than Integrate[x Exp [-a x], {x, 0, Infinity}]. And in doing this it takes him only a few lines to derive what’s now called the Planck spectrum for blackbody radiation (i.e. for “radiation in equilibrium”):

Click to enlarge

And then by fitting this result to the data of the time he gets “Planck’s constant” (the correct result is 6.62):

Click to enlarge

And, yes, this was essentially the birth of quantum mechanics—essentially as a spinoff from an attempt to extend the domain of the Second Law. Planck himself didn’t seem to internalize what he’d done for at least another decade. And it was really Albert Einstein’s 1905 analysis of the photoelectric effect that made the concept of the quantization of energy that Planck had assumed (more as a calculational hypothesis than anything else) seem to be something of real physical significance—that would lead to the whole development of quantum mechanics, notably in the 1920s.

Are Molecules Real? Continuous Versus Discrete

As we discussed at the very beginning above, already in antiquity there was a notion that at least things like solids and liquids might not ultimately be continuous (as they seemed), but might instead be made of large numbers of discrete “atomic” elements. By the 1600s there was also the idea that light might be “corpuscular”—and, as we discussed above, gases too. But meanwhile, there were opposing theories that espoused continuity—like the caloric theory of heat. And particularly with the success of calculus, there was a strong tendency to develop theories that showed continuity—and to which calculus could be applied.

But in the early 1800s—notably with the work of John Dalton (1766–1844)—there began to be evidence that there were discrete entities participating in chemical reactions. Meanwhile, as we discussed above, the success of the kinetic theory of gases gave increasing evidence for some kind of—at least effectively—discrete elements in gases. But even people like Boltzmann and Maxwell were reluctant to assert that gases really were made of molecules. And there were plenty of well-known scientists (like Ernst Mach) who “opposed atomism”, often effectively on the grounds that in science one should only talk about things one can actually see or experience—not things like atoms that were too small for that.

But there was something else too: with Newton’s theory of gravitation as a precursor, and then with the investigation of electromagnetic phenomena, there emerged in the 1800s the idea of a “continuous field”. The interpretation of this was fairly clear for something like an elastic solid or a fluid that exhibited continuous deformations.

Mathematically, things like gravity, magnetism—and heat—seemed to work in similar ways. And it was assumed that this meant that in all cases there had to be some fluid-like “carrier” for the field. And this is what led to ideas like the luminiferous aether as the “carrier” of electromagnetic waves. And, by the way, the idea of an aether wasn’t even obviously incompatible with the idea of atoms; Kelvin, for example, had a theory that atoms were vortices (perhaps knotted) in the aether.

But how does this all relate to the Second Law? Well, particularly through the work of Boltzmann there came to be the impression that given atomism, probability theory could essentially “prove” the Second Law. A few people tried to clarify the formal details (as we discussed above), but it seemed like any final conclusion would have to await the validation (or not) of atomism, which in the late 1800s was still a thoroughly controversial theory.

By the first decade of the 1900s, however, the fortunes of atomism began to change. In 1897 J. J. Thomson (1856–1940) discovered the electron, showing that electricity was fundamentally “corpuscular”. And in 1900 Planck had (at least calculationally) introduced discrete quanta of energy. But it was the three classic papers of Albert Einstein in 1905 that—in their different ways—began to secure the ultimate success of atomism.

First there was his paper “On a Heuristic View about the Production and Transformation of Light”, which began:

Maxwell’s theory of electromagnetic [radiation] differs in a profound, essential way from the current theoretical models of gases and other matter. We consider the state of a material body to be completely determined by the positions and velocities of a finite number of atoms and electrons, albeit a very large number. But the electromagnetic state of a region of space is described by continuous functions …

He then points out that optical experiments look only at time-averaged electromagnetic fields, and continues:

In particular, blackbody radiation, photoluminescence, [the photoelectric effect] and other phenomena associated with the generation and transformation of light seem better modeled by assuming that the energy of light is distributed discontinuously in space. According to this picture, the energy of a light wave emitted from a point source is not spread continuously over ever larger volumes, but consists of a finite number of energy quanta that are spatially localized at points of space, move without dividing and are absorbed or generated only as a whole.

In other words, he’s suggesting that light is “corpuscular”, and that energy is quantized. When he begins to get into details, he’s soon talking about the “entropy of radiation”—and, then, in three core sections of his paper, he’s basing what he’s doing on “Boltzmann’s principle”:

Click to enlarge

Two months later, Einstein produced another paper: “Investigations on the Theory of Brownian Motion”. Back in 1827 the British botanist Robert Brown (1773–1858) had seen under a microscope tiny grains (ejected by pollen) randomly jiggling around in water. Einstein began his paper:

In this paper it will be shown that according to the molecular-kinetic theory of heat, bodies of microscopically visible size suspended in a liquid will perform movements of such magnitude that they can be easily observed in a microscope, on account of the molecular motions of heat.

He doesn’t explicitly mention Boltzmann in this paper, but there’s Boltzmann’s formula again:

Click to enlarge

And by the next year it’s become clear experimentally that, yes, the jiggling Robert Brown had seen was in fact the result of impacts from discrete, real water molecules.

Einstein’s third 1905 paper, “On the Electrodynamics of Moving Bodies”—in which he introduced relativity theory—wasn’t so obviously related to atomism. But in showing that the luminiferous aether will (as Einstein put it) “prove superfluous” he was removing what was (almost!) the last remaining example of something continuous in physics.

In the years after 1905, the evidence for atomism mounted rapidly, segueing in the 1920s into the development of quantum mechanics. But what happened with the Second Law? By the time atomism was generally accepted, the generation of physicists that had included Boltzmann and Gibbs was gone. And while the Second Law was routinely invoked in expositions of thermodynamics, questions about its foundations were largely forgotten. Except perhaps for one thing: people remembered that “proofs” of the Second Law had been controversial, and had depended on the controversial hypothesis of atomism. But—they appear to have reasoned—now that atomism isn’t controversial anymore, it follows that the Second Law is indeed “satisfactorily proved”. And, after all, there were all sorts of other things to investigate in physics.

There are a couple of “footnotes” to this story. The first has to do with Einstein. Right before Einstein’s remarkable series of papers in 1905, what was he working on? The answer is: the Second Law! In 1902 he wrote a paper entitled “Kinetic Theory of Thermal Equilibrium and of the Second Law of Thermodynamics”. Then in 1903: “A Theory of the Foundations of Thermodynamics”. And in 1904: “On the General Molecular Theory of Heat”. The latter paper claims:

I derive an expression for the entropy of a system, which is completely analogous to the one found by Boltzmann for ideal gases and assumed by Planck in his theory of radiation. Then I give a simple derivation of the Second Law.

But what’s actually there is not quite what’s advertised:

Click to enlarge

It’s a short argument—about interactions between a collection of heat reservoirs. But in a sense it already assumes its answer, and certainly doesn’t provide any kind of fundamental “derivation of the Second Law”. And this was the last time Einstein ever explicitly wrote about deriving the Second Law. Yes, in those days it was just too hard, even for Einstein.

There’s another footnote to this story too. As we said, at the beginning of the twentieth century it had become clear that lots of things that had been thought to be continuous were in fact discrete. But there was an important exception: space. Ever since Euclid (~300 BC), space had almost universally been implicitly assumed to be continuous. And, yes, when quantum mechanics was being built, people did wonder about whether space might be discrete too (and even in 1917 Einstein expressed the opinion that eventually it would turn out to be). But over time the idea of continuous space (and time) got so entrenched in the fabric of physics that when I started seriously developing the ideas that became our Physics Project based on space as a discrete network (or what—in homage to the dynamical theory of heat one might call the “dynamical theory of space”) it seemed to many people quite shocking. And looking back at the controversies of the late 1800s around atomism and its application to the Second Law it’s charming how familiar many of the arguments against atomism seem. Of course it turns out they were wrong—as they seem again to be in the case of space.

The Twentieth Century

The foundations of thermodynamics were a hot topic in physics in the latter half of the nineteenth century—worked on by many of the most prominent physicists of the time. But by the early twentieth century it’d been firmly eclipsed by other areas of physics. And going forward it’d receive precious little attention—with most physicists just assuming it’d “somehow been solved”, or at least “didn’t need to be worried about”.

As a practical matter, thermodynamics in its basic equilibrium form nevertheless became very widely used in engineering and in chemistry. And in physics, there was steadily increasing interest in doing statistical mechanics—typically enumerating states of systems (quantum or otherwise), weighted as they would be in idealized thermal equilibrium. In mathematics, the field of ergodic theory developed, though for the most part it concerned itself with systems (such as ordinary differential equations) involving few variables—making it relevant to the Second Law essentially only by analogy.

There were a few attempts to “axiomatize” the Second Law, but mostly only at a macroscopic level, not asking about its microscopic origins. And there were also attempts to generalize the Second Law to make robust statements not just about equilibrium and the fact that it would be reached, but also about what would happen in systems driven to be in some manner away from equilibrium. The fluctuation-dissipation theorem about small perturbations from equilibrium—established in the mid-1900s, though anticipated in Einstein’s work on Brownian motion—was one example of a widely applicable result. And there were also related ideas of “minimum entropy production”—as well as “maximum entropy production”. But for large deviations from equilibrium there really weren’t convincing general results, and in practice most investigations basically used phenomenological models that didn’t have obvious connections to the foundations of thermodynamics, or derivations of the Second Law.

Meanwhile, through most of the twentieth century there were progressively more elaborate mathematical analyses of Boltzmann’s equation (and the H theorem) and their relation to rigorously derivable but hard-to-manage concepts like the BBGKY hierarchy. But despite occasional claims to the contrary, such approaches ultimately never seem to have been able to make much progress on the core problem of deriving the Second Law.

And then there’s the story of entropy. And in a sense this had three separate threads. The first was the notion of entropy—essentially in the original form defined by Clausius—being used to talk quantitatively about heat in equilibrium situations, usually for either engineering or chemistry. The second—that we’ll discuss a little more below—was entropy as a qualitative characterization of randomness and degradation. And the third was entropy as a general and formal way to measure the “effective number of degrees of freedom” in a system, computed from the log of the number of its achievable states.

There are definitely correspondences between these different threads. But they’re in no sense “obviously equivalent”. And much of the mystery—and confusion—that developed around entropy in the twentieth century came from conflating them.

Another piece of the story was information theory, which arose in the 1940s. And a core question in information theory is how long an “optimally compressed” message will be. And (with various assumptions) the average such length is given by a ∑p log p form that has essentially the same structure as Boltzmann’s expression for entropy. But even though it’s “mathematically like entropy” this has nothing immediately to do with heat—or even physics; it’s just an abstract consequence of needing log Ω bits (i.e. log Ω degrees of freedom) to specify one of Ω possibilities. (Still, the coincidence of definitions led to an “entropy branding” for various essentially information-theoretic methods, with claims sometimes being made that, for example, the thing called entropy must always be maximized “because we know that from physics”.)

There’d been an initial thought in the 1940s that there’d be an “inevitable Second Law” for systems that “did computation”. The argument was that logical gates (like And and Or) take 2 bits of input (with 4 overall states 11, 10, 01, 00) but give only 1 bit of output (1 or 0), and are therefore fundamentally irreversible. But in the 1970s it became clear that it’s perfectly possible to do computation reversibly (say with 2-input, 2-output gates)—and indeed this is what’s used in the typical formalism for quantum circuits.

As I’ve mentioned elsewhere, there were some computer experiments in the 1950s and beyond on model systems—like hard sphere gases and nonlinear springs—that showed some sign of Second Law behavior (though less than might have been expected). But the analysis of these systems very much concentrated on various regularities, and not on the effective randomness associated with Second Law behavior.

In another direction, the 1970s saw the application of thermodynamic ideas to black holes. At first, it was basically a pure analogy. But then quantum field theory calculations suggested that black holes should produce thermal radiation as if they had a certain effective temperature. By the late 1990s there were more direct ways to “compute entropy” for black holes, by enumerating possible (quantum) configurations consistent with the overall characteristics of the black hole. But such computations in effect assume (time-invariant) equilibrium, and so can’t be expected to shed light directly on the Second Law.

Talking about black holes brings up gravity. And in the course of the twentieth century there were scattered efforts to understand the effect of gravity on the Second Law. Would a self-gravitating gas achieve “equilibrium” in the usual sense? Does gravity violate the Second Law? It’s been difficult to get definitive answers. Many specific simulations of n-body gravitational systems were done, but without global conclusions for the Second Law. And there were cosmological arguments, particularly about the role of gravity in accounting for entropy in the early universe—but not so much about the actual evolution of the universe and the effect of the Second Law on it.

Yet another direction has involved quantum mechanics. The standard formalism of quantum mechanics—like classical mechanics—is fundamentally reversible. But the formalism for measurement introduced in the 1930s—arguably as something of a hack—is fundamentally irreversible, and there’ve been continuing arguments about whether this could perhaps “explain the Second Law”. (I think our Physics Project finally provides more clarity about what’s going on here—but also tells us this isn’t what’s “needed” for the Second Law.)

From the earliest days of the Second Law, there had always been scattered but ultimately unconvincing assertions of exceptions to the Second Law—usually based on elaborately constructed machines that were claimed to be able to achieve perpetual motion “just powered by heat”. Of course, the Second Law is a claim about large numbers of molecules, etc.—and shouldn’t be expected to apply to very small systems. But by the end of the twentieth century it was starting to be possible to make micromachines that could operate on small numbers of molecules (or electrons). And with the right control systems in place, it was argued that such machines could—at least in principle—effectively be used to set up Maxwell’s demons that would systematically violate the Second Law, albeit on a very small scale.

And then there was the question of life. Early formulations of the Second Law had tended to talk about applying only to “inanimate matter”—because somehow living systems didn’t seem to follow the same process of inexorable “dissipation to heat” as inanimate, mechanical systems. And indeed, quite to the contrary, they seemed able to take disordered input (like food) and generate ordered biological structures from it. And indeed, Erwin Schrödinger (1887–1961), in his 1944 book What Is Life? talked about “negative entropy” associated with life. But he—and many others since—argue that life doesn’t really violate the Second Law because it’s not operating in a closed environment where one should expect evolution to equilibrium. Instead, it’s constantly being driven away from equilibrium, for example by “organized energy” ultimately coming from the Sun.

Still, the concept of at least locally “antithermodynamic” behavior is often considered to be a potential general signature of life. But already by the early part of the 1900s, with the rise of things like biochemistry, and the decline of concepts like “life force” (which seemed a little like “caloric”), there developed a strong belief that the Second Law must at some level always apply, even to living systems. But, yes, even though the Second Law seemed to say that one can’t “unscramble an egg”, there was still the witty rejoinder: “unless you feed it to a chicken”.

What about biological evolution? Well, Boltzmann had been an enthusiast of Darwin’s idea of natural selection. And—although it’s not clear he made this connection—it was pointed out many times in the twentieth century that just as in the Second Law reversible underlying dynamics generate an irreversible overall effect, so also in Darwinian evolution effectively reversible individual changes aggregate to what at least Darwin thought was an “irreversible” progression to things like the formation of higher organisms.

The Second Law also found its way into the social sciences—sometimes under names like “entropy pessimism”—most often being used to justify the necessity of “Maxwell’s-demon-like” active intervention or control to prevent the collapse of economic or social systems into random or incoherent states.

But despite all these applications of the Second Law, the twentieth century largely passed without significant advances in understanding the origin and foundations of the Second Law. Though even by the early 1980s I was beginning to find results—based on computational ideas—that seemed as if they might finally give a foundational understanding of what’s really happening in the Second Law, and the extent to which the Second Law can in the end be “derived” from underlying “mechanical” rules.

What the Textbooks Said: The Evolution of Certainty

Ask a typical physicist today about the Second Law and they’re likely to be very sure that it’s “just true”. Maybe they’ll consider it “another law of nature” like the conservation of energy, or maybe they’ll think it is something that was “proved long ago” from basic principles of mathematics and mechanics. But as we’ve discussed here, there’s really nowhere in the history of the Second Law that should give us this degree of certainty. So where did all the certainty come from? I think in the end it’s a mixture of a kind of don’t-question-this-it-comes-from-sophisticated-science mystique about the Second Law, together with a century and a half of “increasingly certain” textbooks. So let’s talk about the textbooks.

While early contributions to what we now call thermodynamics (and particularly those from continental Europe) often got published as monographs, the first “actual textbooks” of thermodynamics already started to appear in the 1860s, with three examples (curiously, all in French) being:

Click to enlarge

And in these early textbooks what one repeatedly sees is that the Second Law is simply cited—without much comment—as a “principle” or “axiom” (variously attributed to Carnot, Kelvin or Clausius, and sometimes called “the Principle of Carnot”), from which theory will be developed. By the 1870s there’s a bit of confusion starting to creep in, because people are talking about the “Theorem of Carnot”. But, at least at first, by this they mean not the Second Law, but the result on the efficiency of heat engines that Carnot derived from this.

Occasionally, there are questions in textbooks about the validity of the Second Law. A notable one, that we discussed above when we talked about Maxwell’s demon, shows up under the title “Limitation of the Second Law of Thermodynamics” at the end of Maxwell’s 1871 Theory of Heat.

Tait’s largely historical 1877 Sketch of Thermodynamics notes that, yes, the Second Law hasn’t successfully been proved from the laws of mechanics:

Click to enlarge

In 1879, Eddy’s Thermodynamics at first shows even more skepticism

Click to enlarge

but soon he’s talking about how “Rankine’s theory of molecular vortices” has actually “proved the Second Law”:

Click to enlarge

He goes on to give some standard “phenomenological” statements of the Second Law, but then talks about “molecular hypotheses from which Carnot’s principle has been derived”:

Click to enlarge

Pretty soon there’s confusion like the section in Alexandre Gouilly’s (1842–1906) 1877 Mechanical Theory of Heat that’s entitled “Second Fundamental Theorem of Thermodynamics or the Theorem of Carnot”:

Click to enlarge

More textbooks on thermodynamics follow, but the majority tend to be practical expositions (that are often incredibly similar to each other) with no particular theoretical discussion of the Second Law, its origins or validity.

In 1891 there’s an “official report about the Second Law” commissioned by the British Association for the Advancement of Science (and written by a certain George Bryan (1864–1928) who would later produce a thermodynamics textbook):

Click to enlarge

There’s an enumeration of approaches so far:

Click to enlarge

Somewhat confusingly it talks about a “proof of the Second Law”—actually referring to an already-in-equilibrium result:

Click to enlarge

There’s talk of mechanical instability leading to irreversibility:

Click to enlarge

The conclusions say that, yes, the Second Law isn’t proved “yet”

Click to enlarge

but imply that if only we knew more about molecules that might be enough to nail it:

Click to enlarge

But back to textbooks. In 1895 Boltzmann published his Lectures on Gas Theory, which includes a final chapter about the H theorem and its relation to the Second Law. Boltzmann goes through his mathematical derivations for gases, then (rather over-optimistically) asserts that they’ll also work for solids and liquids:

We have looked mainly at processes in gases and have calculated the function H for this case. Yet the laws of probability that govern atomic motion in the solid and liquid states are clearly not qualitatively different … from those for gases, so that the calculation of the function H corresponding to the entropy would not be more difficult in principle, although to be sure it would involve greater mathematical difficulties.

But soon he’s discussing the more philosophical aspects of things (and by the time Boltzmann wrote this book, he was a professor of philosophy as well as physics). He says that the usual statement of the Second Law is “asserted phenomenologically as an axiom” (just as he says the infinite divisibility of matter also is at that time):

… the Second Law is formulated in such a way that the unconditional irreversibility of all natural processes is asserted as an axiom, just as general physics based on a purely phenomenological standpoint asserts the unconditional divisibility of matter without limit as an axiom.

One might then expect him to say that actually the Second Law is somehow provable from basic physical facts, such as the First Law. But actually his claims about any kind of “general derivation” of the Second Law are rather subdued:

Since however the probability calculus has been verified in so many special cases, I see no reason why it should not also be applied to natural processes of a more general kind. The applicability of the probability calculus to the molecular motion in gases cannot of course be rigorously deduced from the differential equations for the motion of the molecules. It follows rather from the great number of the gas molecules and the length of their paths, by virtue of which the properties of the position in the gas where a molecule undergoes a collision are completely independent of the place where it collided the previous time.

But he still believes in the ultimate applicability of the Second Law, and feels he needs to explain why—in the face of the Second Law—the universe as we perceive “still has interesting things going on”:

… small isolated regions of the universe will always find themselves “initially” in an improbable state. This method seems to me to be the only way in which one can understand the Second Law—the heat death of each single world—without a unidirectional change of the entire universe from a definite initial state to a final state.

Meanwhile, he talks about the idea that elsewhere in the universe things might be different—and that, for example, entropy might be systematically decreasing, making (he suggests) perceived time run backwards:

In the entire universe, the aggregate of all individual worlds, there will however in fact
occur processes going in the opposite direction. But the beings who observe such processes will simply reckon time as proceeding from the less probable to the more probable states, and it will never be discovered whether they reckon time differently from us, since they are separated from us by eons of time and spatial distances 101010 times the distance of Sirius—and moreover their language has no relation to ours.

Most other textbook discussions of thermodynamics are tamer than this, but the rather anthropic-style argument that “we live in a fluctuation” comes up over and over again as an ultimate way to explain the fact that the universe as we perceive it isn’t just a featureless maximum-entropy place.

It’s worth noting that there are roughly three general streams of textbooks that end up discussing the Second Law. There are books about rather practical thermodynamics (of the type pioneered by Clausius), that typically spend most of their time on the equilibrium case. There are books about kinetic theory (effectively pioneered by Maxwell), that typically spend most of their time talking about the dynamics of gas molecules. And then there are books about statistical mechanics (as pioneered by Gibbs) that discuss with various degrees of mathematical sophistication the statistical characteristics of ensembles.

In each of these streams, many textbooks just treat the Second Law as a starting point that can be taken for granted, then go from there. But particularly when they are written by physicists with broader experience, or when they are intended for a not-totally-specialized audience, textbooks will quite often attempt at least a little justification or explanation for the Second Law—though rather often with a distinct sleight of hand involved.

For example, when Planck in 1903 wrote his Treatise on Thermodynamics he had a chapter in his discussion of the Second Law, misleadingly entitled “Proof”. Still, he explains that:

The second fundamental principle of thermodynamics [Second Law] being, like the first, an empirical law, we can speak of its proof only in so far as its total purport may be deduced from a single self-evident proposition. We, therefore, put forward the following proposition as being given directly by experience. It is impossible to construct an engine which will work in a complete cycle, and produce no effect except the raising of a weight and the cooling of a heat-reservoir.

In other words, his “proof” of the Second Law is that nobody has ever managed to build a perpetual motion machine that violates it. (And, yes, this is more than a little reminiscent of PNP, which, through computational irreducibility, is related to the Second Law.) But after many pages, he says:

In conclusion, we shall briefly discuss the question of the possible limitations to the Second Law. If there exist any such limitations—a view still held by many scientists and philosophers—then this [implies an error] in our starting point: the impossibility of perpetual motion …

(In the 1905 edition of the book he adds a footnote that frankly seems bizarre in view of his—albeit perhaps initially unwilling—role in the initiation of quantum theory five years earlier: “The following discussion, of course, deals with the meaning of the Second Law only insofar as it can be surveyed from the points of view contained in this work avoiding all atomic hypotheses.”)

He ends by basically saying “maybe one day the Second Law will be considered necessarily true; in the meantime let’s assume it and see if anything goes wrong”:

Presumably the time will come when the principle of the increase of the entropy will be presented without any connection with experiment. Some metaphysicians may even put it forward as being a priori valid. In the meantime, no more effective weapon can be used by both champions and opponents of the Second Law than the indefatigable endeavour to follow the real purport of this law to the utmost consequences, taking the latter one by one to the highest court of appeal experience. Whatever the decision may be, lasting gain will accrue to us from such a proceeding, since thereby we serve the chief end of natural science the enlargement of our stock of knowledge.

Planck’s book came in a sense from the Clausius tradition. James Jeans’s (1877–1946) 1904 book The Dynamical Theory of Gases came instead from the Maxwell + Boltzmann tradition. He says at the beginning—reflecting the fact the existence of molecules had not yet been firmly established in 1904—that the whole notion of the molecular basis of heat “is only a hypothesis”:

Click to enlarge

Later he argues that molecular-scale processes are just too “fine-grained” to ever be directly detected:

Click to enlarge

But soon Jeans is giving a derivation of Boltzmann’s H theorem, though noting some subtleties:

Click to enlarge

His take on the “reversibility objection” is that, yes, the H function will be symmetric at every maximum, but, he argues, it’ll also be discontinuous there:

Click to enlarge

And in the time-honored tradition of saying “it is clear” right when an argument is questionable, he then claims that an “obvious averaging” will give irreversibility and the Second Law:

Click to enlarge

Later in his book Jeans simply quotes Maxwell and mentions his demon:

Click to enlarge

Then effectively just tells readers to go elsewhere:

Click to enlarge

In 1907 George Bryan (whose 1891 report we mentioned earlier) published Thermodynamics, an Introductory Treatise Dealing Mainly with First Principles and Their Direct Applications. But despite its title, Bryan has now “walked back” the hopes of his earlier report and is just treating the Second Law as an “axiom”:

Click to enlarge

And—presumably from his interactions with Boltzmann—is saying that the Second Law is basically an empirical fact of our particular experience of the universe, and thus not something fundamentally derivable:

Click to enlarge

As the years went by, many thermodynamics textbooks appeared, increasingly with an emphasis on applications, and decreasingly with a mention of foundational issues—typically treating the Second Law essentially just as an absolute empirical “law of nature” analogous to the First Law.

But in other books—including some that were widely read—there were occasional mentions of the foundations of the Second Law. A notable example was in Arthur Eddington’s (1882–1944) 1929 The Nature of the Physical World—where now the Second Law is exalted as having the “supreme position among the laws of Nature”:

Click to enlarge

Although Eddington does admit that the Second Law is probably not “mathematically derivable”:

Click to enlarge

And even though in the twentieth century questions about thermodynamics and the Second Law weren’t considered “top physics topics”, some top physicists did end up talking about them, if nothing else in general textbooks they wrote. Thus, for example, in the 1930s and 1940s people like Enrico Fermi (1901–1954) and Wolfgang Pauli (1900–1958) wrote in some detail about the Second Law—though rather strenuously avoided discussing foundational issues about it.

Lev Landau (1908–1968), however, was a different story. In 1933 he wrote a paper “On the Second Law of Thermodynamics and the Universe” which basically argues that our everyday experience is only possible because “the world as a whole does not obey the laws of thermodynamics”—and suggests that perhaps relativistic quantum mechanics (which he says, quoting Niels Bohr (1885–1962), could be crucial in the center of stars) might fundamentally violate the Second Law. (And yes, even today it’s not clear how “relativistic temperature” works.)

But this kind of outright denial of the Second Law had disappeared by the time Lev Landau and Evgeny Lifshitz (1915–1985) wrote the 1951 version of their book Statistical Mechanics—though they still showed skepticism about its origins:

There is no doubt that the foregoing simple formulations [of the Second Law] accord with reality; they are confirmed by all our everyday observations. But when we consider more closely the problem of the physical nature and origin of these laws of behaviour, substantial difficulties arise, which to some extent have not yet been overcome.

Their book continues, discussing Boltzmann’s fluctuation argument:

Firstly, if we attempt to apply statistical physics to the entire universe … we immediately encounter a glaring contradiction between theory and experiment. According to the results of statistics, the universe ought to be in a state of complete statistical equilibrium. … Everyday experience shows us, however, that the properties of Nature bear no resemblance to those of an equilibrium system; and astronomical results show that the same is true throughout the vast region of the Universe accessible to our observation.

We might try to overcome this contradiction by supposing that the part of the Universe which we observe is just some huge fluctuation in a system which is in equilibrium as a whole. The fact that we have been able to observe this huge fluctuation might be explained by supposing that the existence of such a fluctuation is a necessary condition for the existence of an observer (a condition for the occurrence of biological evolution). This argument, however, is easily disproved, since a fluctuation within, say, the volume of the solar system only would be very much more probable, and would be sufficient to allow the existence of an observer.

What do they think is the way out? The effect of gravity:

… in the general theory of relativity, the Universe as a whole must be regarded not as a closed system but as a system in a variable gravitational field. Consequently the application of the law of increase of entropy does not prove that statistical equilibrium must necessarily exist.

But they say this isn’t the end of the problem, essentially noting the reversibility objection. How should this be overcome? First, they suggest the solution might be that the observer somehow “artificially closes off the history of a system”, but then they add:

Such a dependence of the laws of physics on the nature of an observer is quite inadmissible, of course.

They continue:

At the present time it is not certain whether the law of increase of entropy thus formulated can be derived on the basis of classical mechanics. … It is more reasonable to suppose that the law of increase of entropy in the above general formulation arises from quantum effects.

They talk about the interaction of classical and quantum systems, and what amounts to the explicit irreversibility of the traditional formalism of quantum measurement, then say that if quantum mechanics is in fact the ultimate source of irreversibility:

… there must exist an inequality involving the quantum constant ℏ which ensures the validity of the law and is satisfied in the real world…

What about other textbooks? Joseph Mayer (1904–1983) and Maria Goeppert Mayer’s (1906–1972) 1940 Statistical Mechanics has the rather charming

Click to enlarge

though in the end they sidestep difficult questions about the Second Law by basically making convenient definitions of what S and Ω mean in S = k log Ω.

For a long time one of the most cited textbooks in the area was Richard Tolman’s (1881–1948) 1938 Principles of Statistical Mechanics. Tolman (basically following Gibbs) begins by explaining that statistical mechanics is about making predictions when all you know are probabilistic statements about initial conditions:

Click to enlarge

Tolman continues:

Click to enlarge

He notes that, historically, statistical mechanics was developed for studying systems like gases, where (in a vague foreshadowing of the concept of computational irreducibility) “it is evident that we should be quickly lost in the complexities of our computations” if we try to trace every molecule, but where, he claims, statistical mechanics can still accurately tell us “statistically” what will happen:

Click to enlarge

But where exactly should we get the probability distributions for initial states from? Tolman says he’s going to consider the kinds of mathematically defined ensembles that Gibbs discusses. And tucked away at the end of a chapter he admits that, well, yes, this setup is really all just a postulate—set up so as to make the results of statistical mechanics “merely a matter for computation”:

Click to enlarge

On this basis Tolman then derives Boltzmann’s H theorem, and his “coarse-grained” generalization (where, yes, the coarse-graining ultimately operates according to his postulate). For 530 pages, there’s not a single mention of the Second Law. But finally, on page 558 Tolman is at least prepared to talk about an “analog of the Second Law”:

Click to enlarge

And basically what Tolman argues is that his can reasonably be identified with thermodynamic entropy S. In the end, the argument is very similar to Boltzmann’s, though Tolman seems to feel that it has achieved more:

Click to enlarge

Very different in character from Tolman’s book, another widely cited book is Percy Bridgman’s (1882–1961) largely philosophical 1943 The Nature of Thermodynamics. His chapter on the Second Law begins:

Click to enlarge

A decade earlier Bridgman had discussed outright violations of the Second Law, saying that he’d found that the younger generation of physicists at the time seemed to often think that “it may be possible some day to construct a machine which shall violate the Second Law on a scale large enough to be commercially profitable”—perhaps, he said, by harnessing Brownian motion:

Click to enlarge

At a philosophical level, a notable treatment of the Second Law appeared in Hans Reichenbach’s (1891–1953) (unfinished-at-his-death) 1953 work The Direction of Time. Wanting to make use of the Second Law, but concerned about the reversibility objections, Reichenbach introduces the notion of “branch systems”—essentially parts of the universe that can eventually be considered isolated, but which were once connected to other parts that were responsible for determining their (“nonrandom”) effective initial conditions:

Click to enlarge

Most textbooks that cover the Second Law use one of the formulations that we’ve already discussed. But there is one more formulation that also sometimes appears, usually associated with the name “Carathéodory” or the term “axiomatic thermodynamics”.

Back in the first decade of the twentieth century—particularly in the circle around David Hilbert (1862–1943)—there was a lot of enthusiasm for axiomatizing things, including physics. And in 1908 the mathematician Constantin Carathéodory (1873–1950) suggested an axiomatization of thermodynamics. His essential idea—that he developed further in the 1920s—was to consider something like Gibbs’s phase fluid and then roughly to assert that it gets (in some measure-theoretic sense) “so mixed up” that there aren’t “experimentally doable” transformations that can unmix it. Or, in his original formulation:

In any arbitrary neighborhood of an arbitrarily given initial point there is a state that cannot be arbitrarily approximated by adiabatic changes of state.

There wasn’t much pickup of this approach—though Max Born (1882–1970) supported it, Max Planck dismissed it, and in 1939 S. Chandrasekhar (1910–1995) based his exposition of stellar structure on it. But in various forms, the approach did make it into a few textbooks. An example is Brian Pippard’s (1920–2008) otherwise rather practical 1957 The Elements of Classical Thermodynamics:

Click to enlarge

Yet another (loosely related) approach is the “postulatory formulation” on which Herbert Callen’s (1919–1993) 1959 textbook Thermodynamics is based:

Click to enlarge

In effect this is now “assuming the result” of the Second Law:

Click to enlarge

Though in an appendix he rather tautologically states:

Click to enlarge

So what about other textbooks? A famous set are Richard Feynman’s (1918–1988) 1963 Lectures on Physics. Feynman starts his discussion of the Second Law quite carefully, describing it as a “hypothesis”:

Click to enlarge

Feynman says he’s not going to go very far into thermodynamics, though quotes (and criticizes) Clausius’s statements:

Click to enlarge

But then he launches into a whole chapter on “Ratchet and pawl”:

Click to enlarge

His goal, he explains, is to analyze a device (similar to what Marian Smoluchowski had considered in 1912) that one might think by its one-way ratchet action would be able to “harvest random heat” and violate the Second Law. But after a few pages of analysis he claims that, no, if the system is in equilibrium, thermal fluctuations will prevent systematic “one-way” mechanical work from being achieved, so that the Second Law is saved.

But now he applies this to Maxwell’s demon, claiming that the same basic argument shows that the demon can’t work:

Click to enlarge

But what about reversibility? Feynman first discusses what amounts to Boltzmann’s fluctuation idea:

Click to enlarge

But then he opts instead for the argument that for some reason—then unknown—the universe started in a “low-entropy” state, and has been “running down” ever since:

Click to enlarge

By the beginning of the 1960s an immense number of books had appeared that discussed the Second Law. Some were based on macroscopic thermodynamics, some on kinetic theory and some on statistical mechanics. In all three of these cases there was elegant mathematical theory to be described, even if it never really addressed the ultimate origin of the Second Law.

But by the early 1960s there was something new on the scene: computer simulation. And in 1965 that formed the core of Fred Reif’s (1927–2019) textbook Statistical Physics:

Click to enlarge

In a sense the book is an exploration of what simulated hard sphere gases do—as analyzed using ideas from statistical mechanics. (The simulations had computational limitations, but they could go far enough to meaningfully see most of the basic phenomena of statistical mechanics.)

Even the front and back covers of the book provide a bold statement of both reversibility and the kind of randomization that’s at the heart of the Second Law:

Click to enlarge

But inside the book the formal concept of entropy doesn’t appear until page 147—where it’s defined very concretely in terms of states one can explicitly enumerate:

Click to enlarge

And finally, on page 283—after all necessary definitions have been built up—there’s a rather prosaic statement of the Second Law, almost as a technical footnote:

Click to enlarge

Looking though many textbooks of thermodynamics and statistical mechanics it’s striking how singular Reif’s “show-don’t-tell” computer-simulation approach is. And, as I describe in detail elsewhere, for me personally it has a particular significance, because this is the book that in 1972, at the age of 12, launched me on what has now been a 50-year journey to understand the Second Law and its origins.

When the first textbooks that described the Second Law were published nearly a century and a half ago they often (though even then not always) expressed uncertainty about the Second Law and just how it was supposed to work. But it wasn’t long before the vast majority of books either just “assumed the Second Law” and got on with whatever they wanted to apply it to, or tried to suggest that the Second Law had been established from underlying principles, but that it was a sophisticated story that was “out of the scope of this book” but to be found elsewhere. And so it was that a strong sense emerged that the Second Law was something whose ultimate character and origins the typical working scientist didn’t need to question—and should just believe (and protect) as part of the standard canon of science.

So Where Does This Leave the Second Law?

The Second Law is now more than 150 years old. But—at least until now—I think it’s fair to say that the fundamental ideas used to discuss it haven’t materially changed in more than a century. There’s a lot that’s been written about the Second Law. But it’s always tended to follow lines of development already defined over a century ago—and mostly those from Clausius, or Boltzmann, or Gibbs.

Looking at word clouds of titles of the thousands of publications about the Second Law over the decades we see just a few trends, like the appearance of the “generalized Second Law” in the 1990s relating to black holes:

But with all this activity why hasn’t more been worked out about the Second Law? How come after all this time we still don’t really even understand with clarity the correspondence between the Clausius, Boltzmann and Gibbs approaches—or how their respective definitions of “entropy” are ultimately related?

In the end, I think the answer is that it needs a new paradigm—that, yes, is fundamentally based on computation and on ideas like computational irreducibility. A little more than a century ago—with people still actively arguing about what Boltzmann was saying—I don’t think anyone would have been too surprised to find out that to make progress would need a new way of looking at things. (After all, just a few years earlier Boltzmann and Gibbs had needed to bring in the new idea of using probability theory.)

But as we discussed, by the beginning of the twentieth century—with other areas of physics heating up—interest in the Second Law was waning. And even with many questions unresolved people moved on. And soon several academic generations had passed. And as is typical in the history of science, by that point nobody was questioning the foundations anymore. In the particular case of the Second Law there was some sense that the uncertainties had to do with the assumption of the existence of molecules, which had by then been established. But more important, I think, was just the passage of “academic time” and the fact that what might once have been a matter of discussion had now just become a statement in the textbooks—that future academic generations should learn and didn’t need to question.

One of the unusual features of the Second Law is that at the time it passed into the “standard canon of science” it was still rife with controversy. How did those different approaches relate? What about those “mathematical objections”? What about the thought experiments that seemed to suggest exceptions? It wasn’t that these issues were resolved. It was just that after enough time had passed people came to assume that “somehow that must have all been worked out ages ago”.

And it wasn’t that there was really any pressure to investigate foundational issues. The Second Law—particularly in its implications for thermal equilibrium—seemed to work just fine in all its standard applications. And it even seemed to work in new domains like black holes. Yes, there was always a desire to extend it. But the difficulties encountered in trying to do so didn’t seem in any obvious way related to issues about its foundations.

Of course, there were always a few people who kept wondering about the Second Law. And indeed I’ve been surprised at how much of a Who’s Who of twentieth-century physics this seems to have included. But while many well-known physicists seem to have privately thought about the foundations of the Second Law they managed to make remarkably little progress—and as a result left very few visible records of their efforts.

But—as is so often the case—the issue, I believe, is that a fundamentally new paradigm was needed in order to make real progress. When the “standard canon” of the Second Law was formed in the latter part of the nineteenth century, calculus was the primary tool for physics—with probability theory a newfangled addition introduced specifically for studying the Second Law. And from that time it would be many decades before even the beginnings of the computational paradigm began to emerge, and nearly a century before phenomena like computational irreducibility were finally discovered. Had the sequence been different I have no doubt that what I have now been able to understand about the Second Law would have been worked out by the likes of Boltzmann, Maxwell and Kelvin.

But as it is, we’ve had to wait more than a century to get to this point. And having now studied the history of the Second Law—and seen the tangled manner in which it developed—I believe that we can now be confident that we have indeed successfully been able to resolve many of the core issues and mysteries that have plagued the Second Law and its foundations over the course of nearly 150 years.

Note

Almost all of what I say here is based on my reading of primary literature, assisted by modern tools and by my latest understanding of the Second Law. About some of what I discuss, there is—sometimes quite extensive—existing scholarship; some references are given in the bibliography.

Alien Intelligence and the Concept of Technology

Par : Bailey Long
16 juin 2022 à 23:58

The Nature of Alien Intelligence

“We’re going to launch lots of tiny spacecraft into interstellar space, have them discover alien intelligence, then bring back its technology to advance human technology by a million years”. I’ve heard some pretty wacky startup pitches over the years, but this might possibly be the all-time winner.

But as I thought about it, I realized that beyond the “absurdly extreme moonshot” character of this pitch, there’s some science that I’ve done that makes it clear that it’s also fundamentally philosophically confused. The nature of the confusion is interesting, however, and untangling it will give us an opportunity to illuminate some deep features of both intelligence and technology—and in the end suggest a way to think about the long-term trajectory of the very concept of technology and its relation to our universe.

Let’s start with a scenario. Let’s say one of the little spacecraft comes across a planet where it sees complicated swirling patterns:

The Jupiter Great Red Spot
&#10005

The spacecraft sends out a probe to “make contact”. The swirling pattern “responds” by changing slightly. The spacecraft analyzes the change, and sends out another probe. And pretty soon there’s a whole “conversation” going on between the spacecraft and the planet. But, you might say, that’s nothing like an “intelligence” there; there’s just a “pure physical system” that operates through physical laws.

OK, but now let’s imagine the spacecraft has returned to Earth and is checking it out. It detects complicated patterns of radio signals. It sends out a radio signal of its own. Something on Earth responds. A “conversation” ensues. Maybe the spacecraft is “talking to” a cellphone tower, doing automated handshakes with it. Maybe it reached a ham radio operator, and is exchanging Morse code with them. Or maybe—in an ultimate version of code injection—there’s a computer that’s interpreting the spacecraft’s signals as a program, and is sending back the results of running the program.

It all seems quite sophisticated—and at some level worthy of the technological civilization we’ve built up here on Earth. But let’s zoom out a bit. There’s something coming from the spacecraft, that’s causing something to happen on Earth, that’s causing something to be returned to the spacecraft.

And ultimately whatever is happening on Earth must be a physical process of some kind—operating according to the laws of physics. So what’s the difference between this and those swirling “just physics” patterns that the spacecraft found on the other planet? Everything is ultimately “just physics” after all.

OK, you might say, that’s surely true. But on Earth, even though we might have started from physics, we’ve somehow now “ascended” through chemistry and biology and technology to get to something that’s fundamentally more sophisticated. But here we run into an important—if at first surprising—piece of basic science: my Principle of Computational Equivalence.

Let’s say we represent all those “physical processes” as computations (and our Physics Project implies that all of physics is indeed ultimately computational). Now we can compare the computations that correspond to the planet with the swirling patterns to the ones that correspond to our Earth with us humans in the loop.

And what the Principle of Computational Equivalence tells us is that they’re ultimately equivalent. The computations associated with the swirling patterns are ultimately just as sophisticated as the ones we achieve with our brains and our technology here on Earth. It’s far from obvious that this would be true. But it’s something one discovers when one explores the computational universe of possible programs.

One might think that simple programs would produce only simple behavior, and that somehow the behavior would get progressively more complex with more complicated programs. But that’s not what one finds. Instead, there’s increasing evidence that almost any program that doesn’t show obviously simple behavior can in fact show behavior that is as sophisticated as anything.

It’s been known for about a century that there exist computation universal systems capable of being “programmed” to do essentially any computation. But what the Principle of Computational Equivalence says is that sophisticated computation is not only possible—even for simple programs—but is something that happens generically and ubiquitously.

So what does this mean for our spacecraft? It means that what the spacecraft sees on Earth can be computationally no more sophisticated than what it sees on the planet with the swirling patterns. Yes, we consider there to be “intelligence” here on Earth. But what the Principle of Computational Equivalence tells us is that ultimately there’s nothing abstractly different going on from what’s going on in the swirling patterns.

So if we characterize what’s going on here on Earth as an example of “intelligence” we really should say that those swirling patterns are also “examples of intelligence”. And, yes, it doesn’t seem much like human intelligence. But at an abstract computational level it’s really operating like intelligence—but to us humans it’s “alien intelligence”.

There’s a common saying: “The weather has a mind of its own”. And what the Principle of Computational Equivalence tells us is that, yes, fluid dynamics in the atmosphere—and all the swirling patterns associated with it—are examples of computation that are just as sophisticated as those associated with human minds.

But, OK, so there’s a sense in which the weather “has a mind of its own”. But it’s definitely not a “human-like mind”. Yes, the weather does sophisticated computations. But there’s no obvious way to attribute to those computations the purposes and intentions and other typical features of how we describe what goes on in a human mind. So if indeed we’re going to talk about the weather as being an intelligence, for us humans we have to consider it an “alien intelligence”.

We started off talking about spacecraft going out into the cosmos to discover alien intelligence. But what the Principle Computational Equivalence is telling us is that actually there’s what we can think of as alien intelligence all around us. Yes, we humans have managed to get to the point where we do all sorts of sophisticated computations. But computations of just the same sophistication are being done in all sorts of systems that don’t have that whole human tower of biology and technology.

For a long time it’s been a mystery why we’ve never detected alien intelligence out there in the cosmos. But actually I think there’s no lack of “alien intelligence”; indeed it’s all around us. But the point is that it really is alien. At an abstract computational level it’s like our intelligence. But in its details it’s not aligned with our intelligence. Abstractly it’s intelligence, but it’s not human-like intelligence. It’s alien intelligence.

The Role of Science and Technology

OK, so we can think of lots of systems as being examples of “alien intelligence”. But how does that alien intelligence connect to our human intelligence? Sometimes it’s close enough that we humans can immediately “anthropomorphize” the system to “understand what it’s doing in human terms”. But often we need to put effort into “making a bridge”. And in fact we can view that as being what science and technology are ultimately trying to do.

Let’s say we’re looking at swirling patterns in a fluid. The fluid is doing what it does, in effect continually running a computation that generates its behavior. But how can we “align” that with what’s going on in our brains? That’s where science comes in. Because what science is trying to do is to extract some kind of “human-relatable narrative” from the actual behavior of a system out there in the world. Or in some sense it’s trying to provide a “channel” through which we can “communicate” with the “alien intelligence” that is embodied in what’s out there in the world.

So what about technology? Fundamentally technology is about trying to take what exists out there in the world, and apply it to achieve human purposes. We have a fluid. Now we use it to create hydraulic technology that achieves some practical human purpose. We can see the history of technology as being a progressive effort to identify things out there in the world (metal ore, photoelectricity, liquid crystals, …) that can be sampled and fashioned in such a way as to achieve certain purposes we want.

And insofar as we think of what’s out there in the world as being like alien intelligence, what technology is doing is finding ways to corral that intelligence into achieving human purposes. The truth is that in most of our technology today, we’re not letting that intelligence really do anything close to what it’s capable of. We’re keeping it tightly constrained to take only steps that we can readily understand and foresee. It’s a bit like having a horse with a harness that constrains it to just walk slowly in a straight line—even though without the harness the horse could gallop around and do all sorts of elaborate things, albeit things that we might not readily be able to understand or foresee.

So let’s come back to the spacecraft. It’s reached a planet. And it’s interacting with what’s there. Perhaps there’s some weird electrical storm going on. And, yes, we can think of that as an example of alien intelligence. But if the mission of the spacecraft is to discover technology then what it needs to do is to figure out whether there’s some way to interact with the electrical storm so as to achieve some human purpose.

The storm does what the storm does. But maybe by moving some piece of metal around in just the right way it’s possible to get the storm to charge a battery. Or, more elaborately, perhaps processes in the storm could be used like an analog computer, say to compute solutions to equations. And perhaps—having seen the storm on this planet—it’s even possible to “bottle it up” and replicate it, say in a piece of consumer electronics.

One way to describe what’s going on is just to say rather prosaically that we discovered a phenomenon on the planet, that we were able to use for technology. But more colorfully we could say that we encountered an alien intelligence, we found a way to communicate with it, and then we “brought back” technology from it.

The original startup pitch was about spacecraft getting technology by discovering alien intelligence out in the cosmos. But really the whole spacecraft thing is a distraction. Because actually—as we’ve discussed—there’s plenty we can describe as “alien intelligence” all around us, even right here on Earth. And the issue is in a sense just how to “communicate with it” and find ways to “harness it” for our technological purposes.

We’ve learned in the past century that we can use electrons in semiconductors as a way to build computers. But what about other physical processes? Maybe flowing fluids, for example. Can we use that “alien intelligence” to make a new kind of computer? In the end, the point is that any technology is about finding and harnessing “alien intelligence”. That’s basically just what technology is, and always has been.

The whole “alien intelligence” part of the story, though, is much more relevant when we’re thinking of technology that makes serious use of what we can identify as sophisticated computation. If we’re just using a system from nature for its physical mass, it doesn’t really feel as if we’re using its “intelligence”. But as soon as we try, for example, to base a general computer on it, it’s a quite different story.

In getting technology from the universe we’re basically picking out certain aspects of what exists and choosing to apply these for our purposes. In doing science it seems like we have “less choice” about what aspects of the universe we deal with. After all, we might imagine that science is trying to give us a way to understand anything that’s out there in the universe. But in reality it’s much more like technology. The “scientific narratives” that we understand—at least at a given time in history—are ones that we’re in a sense “primed for”. Yes, something like fluid turbulence might give us “in-your-face” exposure to something computationally sophisticated that’s far from what we normally talk about. But what science mostly concentrates on is creating narratives that are aligned with our existing scientific understanding and discussions—much as technology is set up to be about things that are aligned with our existing human purposes.

Extracting Technology from the Ruliad

One might imagine that—wherever it ultimately comes from—technology must at least always in the end be based on the laws of physics. But what’s emerging from our Physics Project is that actually the story is considerably more complicated than that.

It all begins with the ruliad: the object that represents the entangled limit of all possible computations. The ruliad is a unique, formally necessary object, that in a sense embodies all conceivable existence. And inevitably we are embedded within the ruliad, sampling certain aspects of it to form our perception of reality.

In principle there are all sorts of kinds of observers of the ruliad, with all sorts of kinds of perceptions of reality. But the key point that has emerged as a foundation of our Physics Project is that “observers like us” have certain general characteristics—specifically that we assume that we are persistent through time, and also that we are computationally bounded—and from these characteristics alone, we can abstractly deduce from the structure of the ruliad that we must “experience” core standard laws of known physics.

The ruliad in a sense contains all possible physicses. But it’s our particular kind of sampling of the ruliad that leads us to the particular laws of physics that we currently know. An “alien intelligence” might sample the ruliad quite differently, and thus in effect “experience” quite different laws of physics.

Somewhere underneath everything we can think of there being a giant hypergraph of individual atoms of existence—but with the means of perception observers like us have, we inevitably “coarse grain” to the point where, for example, we experience this as continuous space. Another kind of observer, with different characteristics, might, for example, not do that coarse graining, might never experience continuous space, and might have a completely different perception of how the universe works.

In some sense, therefore, physics is much more like technology than we might expect. There isn’t an “absolute physics”. There’s just the physics that we as observers extract from the ruliad. Much like there’s particular technology that we choose to build from the “raw material” that exists. Put another way, both physics and technology are ultimately things we “extract” from the ruliad, in effect by making certain choices.

How we “extract” physics seems, however, much more constrained. For example, we as humans have only certain particular senses through which we are biologically set up to experience the world. Yet we have the feeling that in technology we can in effect “construct whatever we want”—although inevitably “what we want” is also still at least influenced by how we are biologically set up.

We’re very used to the idea that over time technology progresses—as we invent more, and work out new ways to use our “raw material” to achieve human purposes. But physics as a science progresses too. And in a sense what’s happening there is that we’re expanding our character as observers to be able to perceive and experience more of “what’s going on”—ultimately in the ruliad.

Part of that expansion is actually a matter of technology. We’re building telescopes and microscopes and amplifiers that allow us to extend our raw human senses to be sensitive to more things. But there’s also another part of the expansion that is in effect intellectual: we’re developing new conceptual frameworks that allow us to “corral” things we see happening in the world into forms that “fit narratives” we’ve constructed.

And the important point here is that neither our technology nor our physics is fixed. They’re in a sense co-evolving—gradually allowing more and more of the ruliad to be pulled into our narratives and our purposes. Or, put another way, what we observe is gradually expanding to encompass more and more of the ruliad, and to be able to make use of more and more of it.

Reaching Out across Rulial Space

How are “different intelligences” manifest in the ruliad? We can imagine organizing the ruliad to be laid out in some form of rulial space. And from each point in rulial space one in effect gets a “different perspective” on the ruliad. And that’s at least the beginning of the story of how “different intelligences” exist and experience the ruliad.

It’s similar to what happens with physical space: from different places in physical space one gets a different perspective on the universe. In physical space we have a concept of motion: that observers like us can move from one place in space to another while in effect maintaining our coherence and integrity.

How does this work in rulial space? We can think of different points in rulial space as corresponding to different computations, with different rules. So rulial motion in effect corresponds to making a translation between one computation and another. At the outset, it’s not obvious this would even in principle be possible. But the Principle of Computational Equivalence implies that it ultimately will be. The computations at different points in rulial space will (almost always) be equivalent in their sophistication—and as is typical with universal computation—it’ll therefore in principle be possible to have an “interpretation process” that translates between them.

But the big question is whether this can be achieved in practice. Just how far can a particular observer translate in rulial space while maintaining their coherence and integrity?

The ruliad is a complex and (if sampled across slices in time) continually changing thing. But a critical feature is that there can be structures that have a certain persistence within it. In physical space these are things like particles (as well as black holes) that behave like “stable lumps of space”—or like stable lumps in certain projections of the ruliad. In rulial space there can presumably also be structures with a certain persistence: “particles” of rulial space. And these “particles” somehow correspond to features that “survive across different computational perspectives”—or in effect represent “robust concepts”.

When we talk about “different intelligences” a very familiar example is different human minds. And in a sense we can think of different human minds as being laid out in rulial space—with each mind being at a different rulial position, and thus having a different computational rule by which it operates, and a different “experience of the ruliad”.

So how can these minds “communicate”? Ultimately it is through “rulial motion”. But potentially the most robust form of rulial motion is through rulial particles—which we’ve identified above with the abstract idea of “robust concepts”. Put in a practical way: different (human) minds operate internally in different ways. But they can still “communicate” by exchanging something that in effect “survives translation” between one mind and another: rulial particles corresponding to robust concepts (say expressed in a language).

But, OK, we can imagine rulial space with lots of human minds laid out at different places, with ones that communicate more easily closer together. So what about “alien intelligences”? Well, each one is somewhere in rulial space. But they may be far away from where our human minds are.

We can imagine our rulial particles—or “robust concepts”—being able to reach a certain distance in rulial space. The human idea of “excitement” might for example be able to reach the place in rulial space where we’d find the minds of dogs. But what about the weather, for example? Well, as an alien intelligence, it’s presumably much further away in rulial space—and, anthropomorphize it as we might—it’s not clear what its notion of “excitement” would be.

It’s an often-asked question why—with our spacecraft and radio telescopes and everything else—we haven’t ever run across what we consider to be “naturally occurring” alien intelligences. In the past we might have imagined that the answer is that there just isn’t anything like “intelligence” (outside of us humans) to be found in any part of the universe that we can probe. But the Principle of Computational Equivalence says that’s fundamentally not true, and that in fact “abstract intelligence” is thoroughly ubiquitous among systems with all but the most obviously simple behavior.

So to “find” alien intelligence it’s not that we need a more powerful radio telescope (or a better spacecraft) that can reach further in physical space. Rather, the issue is to be able to reach far enough in rulial space. Or, put another way, even if we view the weather as “having a mind of its own”, the rulial distance between “its mind” and our human minds may be too great for us to be able to “understand” and “communicate with it”.

So what will it take for us to “bridge this rulial gap”? At some level it’s just about building the right science and technology. We can think of science as being about defining a way to “translate” from the computational rules by which some particular system operates to the computational way our human minds operate. Or, in terms of rulial space, finding a way to “move” from the rulial position of the system to the rulial position of our minds—and translating from the way a system works to a “human narrative” that represents it.

Centuries ago we might have just said “the planets do what they do”; maybe their motion in space is driven by an “alien intelligence” that we don’t understand. But then along came mathematical science and we were able to “translate” from the intrinsic computation done by the planets to a mathematical description that we internalized enough to consider it a human narrative that we understand.

In some sense at any given time in intellectual history our minds “reach out a certain distance in rulial space”. We’ve developed conceptual frameworks that allow us to maintain a coherent understanding of a certain range of things—with that range growing as we invent new frameworks. At one time our “domain of understanding”—or the region of rulial space that we could reach—didn’t encompass the behavior of electricity. But our “intellectual expansion” in rulial space eventually reached this, and the result is that we can now use electricity as “raw material” from which to construct technology.

One way we “expand our reach in rulial space” is in effect conceptual: by expanding what we understand. But another way is by being able to “sense” or “measure” more. When we invent radio—or, for that matter, gravitational wave detection—there are immediately new kinds of processes that we manage to “connect to human experience”. Or, put another way, there are more distant parts of the ruliad that we’re able to reach.

More prosaically, we can say that if we want to be able to use something for technology, we’d better be able to detect that it’s there, and we’d better be able to understand it well enough that we can see how it could align with our human purposes. We can think of the ruliad as being full of alien intelligences—with plenty of capabilities to “mine”. But to be able to actually mine something for our technological purposes we have to be able to reach it across rulial space; we have to be able to connect it to us.

So what does this mean for the original startup pitch? Yes, it’s a good idea to “mine alien intelligences” for technology. In fact, that’s basically where technology always comes from. But there’s no need to send out spacecraft, “discover” alien intelligence, and so on. There are “alien intelligences” all around us; the issue is just to reach them across rulial space, and be able to “communicate” with them. But what we’ve argued is that the process of progressively reaching out in rulial space is just the general process of progressively advancing science (and the technology on which it depends).

So, yes, by all means explore more of what’s out there in the world, with more, different kinds of sensors and measurements. Then try to “understand” what you see enough to be able to tell how to align it with human purposes, and make technology out of it. But there’s no pressing need for interstellar spacecraft in this picture. It’s just a matter of doing more science to expand our domain in the ruliad, and mine more of rulial space.

The Evolution of Purpose and the Colonization of Rulial Space

We can think of technology as being about setting up things that exist in the world (or ultimately in the ruliad) to achieve human purposes. And we’ve talked about how the advance of science and technology allows us to progressively reach further in rulial space to get “raw material” for our technology. But we’ve said that technology is intended to “achieve human purposes”. So what might those purposes ultimately be?

Our purposes have certainly evolved over the course of human history. In today’s world, we might view it as purposeful to walk on a treadmill, or to trade cryptocurrencies. But it would be challenging to explain the purpose of such things to someone from even a few hundred years ago.

In a sense, purposes evolve as we build new conceptual frameworks, and as we set up technology that allows us to do new things. More abstractly, we might say that purposes are also something defined by places in rulial space. So when we talk about the evolution of purposes, what we’re really asking is where in rulial space our history and development has led us, and will lead us in the future.

And certainly in the vastness of the whole ruliad, our existing human purposes occupy just an infinitesimally tiny part. Think, for example, of the natural world even as we are currently aware of it. The vast majority of things in it do not seem in any way aligned with our purposes—and we have not been able to mine them for technology. Historically, however, there’s been progressive expansion in the domain of our purposes. There was a time when we knew about magnetic rocks, but had no purpose for magnetism. But over the course of time, from compasses to actuators to memories, more and more human purposes have emerged that connect to the phenomenon of magnetism.

And in a sense we can view the whole core trajectory of human progress as being about the expansion of the region of rulial space—and the ruliad—that represents our purposes. So how will this evolve?

As I’ve discussed extensively before, there seem to be two central features that we as entities in the ruliad have. First, that we are computationally bounded. And second, that we believe we are persistent in time. Computational boundedness is essentially the statement that the region of rulial space that we occupy is limited. In some sense our minds can coherently span a certain region of rulial space, but it’s a bounded region.

What about persistence in time? It means that even though we are always being reconstructed out of different atoms of existence (and different atoms of space), we conflate things to the point where we experience a single continuous thread of existence.

Taken together, these features suggest a picture of us being a kind of “blob” that gradually moves around in rulial space. Does it matter that there isn’t just a single human mind? Well, yes. Without some kind of “observer” there’s no real way to even define what it means to have a “blob”. And in the end it’s a story of consistency of observers observing observers. But the result is that we can think of our whole collective “flotilla” of human purposes as being something localized that moves around in rulial space, expanding the region of the ruliad that it reaches.

But just how far can this go? Imagine that at some time in the distant future we have successfully explored—and “colonized”—much of rulial space. To do this we’d certainly have to have broken out of the particular constraints of our biological construction, and made use of “additional raw material” in the ruliad.

But what would it mean to be spread across a large swath of rulial space? Our very notion of existence seems to depend on localization in rulial space. The thing that we view as “us” is something particular and coherent. To be “bigger” in rulial space is to deny that particularity and coherence, and to become something generic that does not represent any kind of “specific entity that exists”.

In a sense, it’s a pyrrhic view of the ultimate limit of our technological and other evolution. As we progress, we gradually “mine” more and more of the ruliad, pulling it into the domain of technology and of our “human” (or post-human) purposes. But in doing so, we eventually transcend the very characteristics that we identify with existence. In other words, if we go too far with our expansion in the ruliad, we simply cease to exist, at least in the sense that we currently define existence.

Put another way, if we “absorb” more and more “alien intelligence” there’s eventually no longer any coherent “us”. Of course, the very notion of coherence is something we’re basically defining from our current human view of things. And no doubt there are other definitions that could be given. But they’re certainly far away from our current place in rulial space, and it’s not even clear they can be reached without some kind of “discontinuity of motion” that would in effect fundamentally break their connection to us as we are now.

Face to Face with Alien Intelligence, Out in Rulial Space

At a fundamental level, the ruliad is a purely computational object, that we can think of as being made of pure, abstract atoms of existence (or “emes”). When observers like us sample the ruliad we can attribute to it the characteristics that correspond to our perception of physical reality. And a notable feature of that sampling is that it supports the idea of pure motion in physical space. In other words, it allows for the possibility that structures can “maintain their perceived physical integrity” while being “re-formed” out of different atoms of space, which themselves are interpretations of the pure atoms of existence in the ruliad.

But as soon as we start thinking about any kind of serious motion in rulial (rather than physical) space it no longer makes sense to talk about anything like “maintaining physical integrity”, not least because in different places in rulial space the very notion of physics changes. But wherever we are in the ruliad we can still think about what’s going on as computation. We might have some way of observing or sampling the ruliad that gives us some perception of reality—like physics, or mathematics. But if we “atomize” things down to the lowest level, we’ll always find raw computation.

As “physical observers like us” we only have limited capabilities to probe or manipulate the raw ruliad and affect what we perceive as physical reality. We can move physical objects around, maintaining what we observe of their structure. In principle we could imagine deconstructing objects into individual atoms of existence, then recreating them “transporter style” somewhere else in physical space. But as of now, we don’t know how to do this, and most likely it’s not possible for observers like us—because it would require “outcomputing” computationally irreducible features of the structure of space down at the level of individual atoms of space, which is far from what computationally bounded observers like us can expect to do.

But what about raw computation, of the kind that ultimately makes up the ruliad? There the story is different. Because we’re no longer constrained by our character as physical observers, so we’re free to in effect “make up any computation we want”. To explore the physical universe we need physical motion or something like it. And at least for observers like us the only way to achieve this seems to be to progressively move structures across physical space. But to find out what can happen in the computational universe we can effectively just write down any rule (i.e. any program) that appears “anywhere in the ruliad”, and run it.

Of course, when we run a rule on a practical computer, it’s just an emulation of what’s happening in the raw ruliad. But it’s just an abstract rule—so although it will run astronomically slower, its ultimate behavior in our emulation will inevitably be identical to what it is when implemented in terms of individual atoms of existence in the “raw ruliad”.

In principle we could take the same approach in emulating the elements that make up our physical reality. But observers like us are so big relative to the raw elements of the ruliad that we can’t expect our emulations to be at a scale where we can faithfully reproduce what we perceive. (Needless to say, in practice we can still get good approximations, and this is a particularly fertile application of our Physics Project.)

But when we’re dealing with “raw computation” down at the lowest level of the raw ruliad, we can expect to faithfully emulate it. And so it is that we can just pick a cellular automaton or a Turing machine or some other kind of computational system—that in effect comes from anywhere in the ruliad—and emulate it to find out what it does. There’s no “object we have to move” to be able to “look at that part of the ruliad”. We’re emulating things down at the level of individual atoms of existence, and seeing what happens.

We can think about our computational experiments as letting us “jump” to find out what it’s like anywhere in the ruliad. And as we “suddenly materialize” somewhere in the ruliad it’s as if we’re immediately “face to face” with whatever “alien intelligence” there is at that place in the ruliad.

But what can “observers like us” expect to make of that alien intelligence? Well, to be able “communicate” or even “relate” we somehow have to be able to “bridge the gap in rulial space”. And since we just “jumped to a place in rulial space” we don’t immediately have any “progressive path” that “incrementally” takes us from our familiar position in rulial space to wherever the alien intelligence is.

But what does that feel like in practice? The whole idea of ruliology is just to go anywhere we want in the computational universe or in the ruliad, and see what happens when we run the rules we find there. And it’s indeed routine to find that what they do seems quite “alien”. Still, they often have certain essential features that for example remind us of the natural world as we observe it. But our standard methods of science (and mathematics)—developed on the basis of being “observers like we are today”—don’t readily allow us to “understand” the behavior of these systems chosen in the course of doing ruliology. To us they usually just seem to be “showing computational irreducibility”, and behaving in ways that we can effectively get no handle on.

But still, we can in a sense view these programs “out there in the computational universe” (and in effect strewn around the ruliad) as showing us what’s possible. They’re like alien intelligences that we know exist, but that we don’t yet understand, and don’t yet know how to harness or relate to. We can see them as some kind of beacons of possible technology of the future—of things that “exist in the ruliad”, but that we haven’t yet been able to connect to human purposes.

But so how might we make this connection? Well, as it happens, I’ve devoted much of my life to what can be viewed as the construction of a systematic bridge between what’s “computationally possible” and what we humans think of as important. For that’s the story of what I call computational language—and indeed of the whole intellectual structure that is the Wolfram Language.

There’s infinite potential content in the ruliad. But one can view the goal of the Wolfram Language as being to represent—in a way that’s optimized for us humans to understand—those parts that we humans consider important. The language lets us use the concepts of computation not only to crystallize our existing thinking, but also to expand what we can think about, in effect letting us reach out further in rulial space. Computational language is the general way that we “tame the ruliad”—extend the frontier of “human colonization” in the ruliad, and in the end “mine” more and more of the ruliad for “useful technology”.

Just in terms of its practical place in the world today I’ve often said that the Wolfram Language is like an “artifact from the future”. But now we see a deep sense in which this is true. The raw ruliad is just “out there”, with “infinite potential”, but as something whose fundamental character has nothing to do with us humans. But what computational language is about is delivering what one can think of as the ultimate “meta-artifact”: something that progressively turns the raw ruliad into “human-recognizable technology”.

Much of this progress involves the specific, systematic design of the Wolfram Language. But there are also forays that in effect jump further out into rulial space. For example, we’ve often enumerated large collections of simple programs, identifying ones that satisfy a certain criterion. And sometimes that feels a lot like “leveraging alien intelligence” without “understanding” it. The rule 30 cellular automaton, for example, is a good pseudorandom generator, even though we don’t really “understand” even fairly basic things about it.

And, yes, computational language is what we need to concretely “state a criterion”, in effect expressing what we’re thinking about in computational terms—that we can use, for example, to let us explicitly search the ruliad for an “alien intelligence” that does what we want.

What does it look like out in the “raw ruliad”? It’s easy to start just looking at simple programs, say picked at random. And, yes, they have all sorts of elaborate behavior:

&#10005

But what is this behavior “achieving”? Yes, it’s following the particular underlying rules that have been given. But we don’t have any immediate way to connect it to “human purposes”. And in general we can expect that to make that connection what’s needed is for those purposes themselves to “expand”.

Maybe at some moment we call what’s produced “art”, and assign it some “aesthetic purpose”. Maybe at some point we see that it satisfies some engineering purpose that we’ve just realized we should care about. But in general, computational language is the way we can make the connection between “raw computational processes” out there in the ruliad, and our patterns of thinking about things. It’s the ultimate way for us to “communicate with alien intelligence”.

The Launch of a Rulial Space Program

We began with the far-out startup pitch of sending spacecraft to discover alien intelligence and bring its technology back to Earth. But what we’ve realized is that actually no spacecraft—of the ordinary kind—are needed. There’s “alien intelligence” to be found everywhere; you don’t have to travel to interstellar space to find it. But the challenge is to connect the “alien intelligence” to human purposes, and extract from it what we consider “useful technology”. Or, put another way, the issue is not about traversing physical space, but rather about traversing rulial space.

With our spacecraft we humans have so far reached about a 20-trillionth of the way across the physical universe. But no doubt we’ve reached a far smaller fraction of the way across the ruliad. As our science, knowledge and technology increase, we gradually reach further into rulial space. But whether it’s our failure to communicate with cetaceans or our inability to make computers out of, say, fluids, it’s clear that by many measures the distance we’ve gone so far is not so large.

In a sense the startup idea of “harnessing alien intelligence” is the meta-idea of all technology—that in our terms we can state as being to connect what’s “computationally possible” in the ruliad with purposes we humans want to achieve. And I’ve argued that the ultimate meta-technology for doing this is not spacecraft but computational language. Because computational language is what we need to make a bridge between what we care about, and “raw computation” out in the ruliad.

It’s difficult to send physical spacecraft out into interstellar space. But it’s actually a lot easier to probe the much richer possibilities of the ruliad—because in a sense it’s straightforward to put a “rulial spacecraft” anywhere. We just have to pick a rule (or program), then see what the “world” it generates is. But the challenge is then in a sense one of interpretation. What is happening in that world? Can we relate it to things we care about?

At the outset, all we’re likely to see at some “random place” in the ruliad is rampant computational irreducibility. But it’s a fundamental fact that wherever there’s computational irreducibility, there must also be slices of computational reducibility to be found. In the ordinary physical universe that we experience, those are basically our perceived laws of physics. But even in a random sample of the ruliad we can expect there’ll be computational reducibility to be found. It’ll typically be “alien stuff”, though. It might have the character of science, but it won’t be like our existing science. And most likely it won’t align with anything we currently think we care about.

But that is the great challenge and promise of mounting a “rulial space program”. To be confronted not with what we might recognize as “new life and new civilizations”, but with things for which we have no description and no current way of thinking. Perhaps we might view it merely as humbling to encounter such things, and to realize how small a part of the ruliad we yet understand. But we can also view it as a beacon of where we could go. And we can view a whole “rulial space program” as a way of systematizing the ultimate project of exploring all formally possible processes. Or we could think about it not just as defining a single “startup opportunity”—but rather as defining the “meta-opportunity” of all possible technology startups….

❌
❌