Vue normale

Il y a de nouveaux articles disponibles, cliquez pour rafraîchir la page.
À partir d’avant-hierFlux principal

85·Faut-il craindre l’intelligence artificielle (IA) ?

Par : Marc HENRY
23 juin 2024 à 18:39

Définition de l’Intelligence Artificielle

Pour ma part, je ne crains nullement ce que l’on appelle “l’Intelligence Artificielle” (IA). Depuis quelques années, on voit se développer sur internet et les réseaux sociaux, une nouvelle forme de diffusion de l’information écrit et visuelle. Cette diffusion fait appel aux nouvelles technologies fondées sur l’informatique. Rappelons pour commencer que l’IA est une branche de l’informatique. Celle qui permet aux systèmes d’apprendre et d’exécuter des tâches normalement associées à l’intelligence humaine. Comme la reconnaissance vocale, la prise de décisions ou la perception visuelle. Rappelons que dans les années 80, deux techniques importantes furent développées.

La première, le deep learning (DL), ou « apprentissage en profondeur », a permis aux ordinateurs d’apprendre par l’expérience. Le second, le « système expert » (SE), imite la capacité de l’être humain à prendre des décisions. Les ordinateurs ont commencé à utiliser un raisonnement fondé sur des « règles » en recourant principalement à une structure « si… alors » mise en œuvre pour répondre à des questions binaires. Ici, la réponse est soit oui, soit non. Depuis les années 2000, trois améliorations ont été apportées :

1. Développement des unités de traitement graphique (GPU) sans lesquelles une IA serait non crédible, provoquant juste un simple haussement d’épaules.

2. Exploitation de l’énorme quantité d’informations mémorisées (Le “Big Data”) depuis qu’internet existe (BD).

3. Développement de nouveaux algorithmes utilisant des variables cachées qui permettent de trier et d’optimiser les résultats d’un calcul.

Variables cachées

En réalité, les points #1 et #2 sont assez triviaux et de nature purement technologique. Ils n’expliquent pas le développement fulgurant observé dans le domaine de l’IA ces dernières années. Ce qui est crucial, c’est le point #3 sur lequel il convient d’être plus bavard :


Car, la description « accessible » que l’on a d’un système, est très souvent la liste de ses comportements observables. Une telle liste est très habituellement une fonction de variables qui ne sont pas immédiatement évidentes. On parle alors de variables dites “cachées” ou bien encore “latentes”. Et, ces variables “cachées” mais hautement pertinentes, sont beaucoup moins nombreuses que les variables observées qui nous sont directement accessibles. D’où le développement de nouveaux algorithmes dont le but n’est pas de traiter directement les observables accessibles. Mais, plutôt, de réduire l’espace des données disponibles à quelques variables cachées. Variables qui sont peu nombreuses, et qui décrivent au mieux les données brutes disponibles.

Analyse en Composantes Principales (ACP)

L’outil clé est ici l’analyse en composantes principales (ACP ou PCA en anglais pour principal component analysis). Cette analyse consiste à transformer des variables dites « corrélées » sur un plan statistique, en nouvelles variables décorrélées les unes des autres. Ces nouvelles variables se nomment « composantes principales » ou axes principaux. L’ACP permet donc de résumer l’information énorme disponible avec de nombreuses variables corrélées en un petit nombre de variables indépendantes les unes des autres. D’un point de vue mathématique, selon le domaine d’application, on utilise soit la transformation de Karhunen–Loève (KLT), soit la transformation de Hotelling (HT).
Les champs d’application de l’ACP sont aujourd’hui multiples, allant de la biologie à la recherche économique et sociale, et plus récemment le traitement d’images et l’apprentissage automatique. L’utilité de l’ACP réside dans les points suivants :

Description et visualisation des données complexes.

Décorrélation des données brutes, la nouvelle base étant constituée d’axes non corrélés entre eux.

Minimisation du bruit, en considérant que les axes que l’on décide d’oublier sont des axes bruités.

Effectuer une réduction de dimension des données d’entraînement en apprentissage automatique.

Il s’agit d’une approche aussi bien géométrique que statistique. Géométrique, puisque les variables se représentent dans un nouvel espace, selon des directions d’inertie maximale. Statistique, parce que la recherche porte sur des axes indépendants expliquant au mieux la variabilité — la variance — des données. En résumé, lorsqu’il s’agit de compresser un ensemble de variables aléatoires, les premiers axes de l’ACP constituent un meilleur choix, aussi bien du point de vue de l’inertie géométrique que depuis la perspective de la variance statistique.

Insuffisance du ternaire GPU/BD/ACP

Après cette démystification de ce qui se cache derrière une IA, se pose l’inévitable question de savoir si la notion “d’intelligence” est réductible à ce ternaire (GPU/BD/ACP). Pour ce qui me concerne, la réponse à cette question est un “non”, franc et massif. Et, pour bien enfoncer le clou, voici quelques éléments de réflexion.


Pour fonctionner, une IA a besoin d’accéder à une base de données de type Big Data (BD), la plus large possible. Le gros problème qu’elle rencontre ici est que pour accéder aux données, il lui faut disposer d’une source d’alimentation électrique. Débranchez l’IA du secteur, ou bien, retirez-lui sa batterie et il n’y a plus rien, de manière instantanée. Toutefois, dès que vous remettez une source d’énergie, le système redémarre et redevient totalement opérationnel.

Pour un être humain, ce débranchement instantané est impossible. Si vous le privez de nourriture, il pourra continuer à fonctionner pendant des années tant qu’il aura de l’eau à sa disposition. Évidemment, si vous lui supprimez l’eau, il finira par mourir au bout de quelques jours. Mais, en aucun cas, la perte d’activité sera instantanée comme avec une intelligence artificielle. Et, une fois l’être humain mort, impossible de redevenir opérationnel si la nourriture et l’eau redeviennent disponibles. Ces simples considérations d’alimentation, nous font toucher du doigt que si le cerveau est parfaitement capable d’utiliser le triangle GPU/BD/ACP, il ne peut pas être réduit à cela. Il y a un plus, qui dépasse le strict cadre informatique.

Et, la Conscience ?

Ce truc est, bien sûr, ce que l’on appelle la “conscience”. Et, comme il est facile de le démontrer (voir https://riviste.fupress.net/index.php/subs/article/view/161), cette conscience est aussi bien neuronale (le moi vu de l’extérieur) qu’extra-neuronale (le moi-même vu de l’intérieur). Rappelons que si le moi traite l’information transmise à l’extérieur du système, le moi-même traite pour sa part l’exformation, c’est-à-dire le contexte, qui lui reste piégé à l’intérieur du système n’étant jamais transmis (voir ici pour en savoir plus : https://marchenry.org/product/de-linformation-a-lexformation/). Pour faire plus simple, un cerveau humain traite l’information aussi bien de manière digitale via les neurones (GPU/BD/ACP) que de manière intuitive via l’eau d’hydratation de ces mêmes neurones. Bref, tant qu’une IA n’utilisera pas l’eau, elle restera largement en deçà d’une intelligence biologique (IB), bactérienne, végétale, animale ou humaine. Car, la fonction “intuition” utilise l’eau, interface de communication instantanée et non locale avec le vide quantique, appelé aussi “éther”.

Intelligence Biologique

Un dernier point concerne la notion même d’intelligence. Si l’on y réfléchit un tant soit peu, pas besoin d’être intelligent pour prendre des décisions lorsque l’on dispose d’une base de données gigantesque. Bien au contraire. On reconnaît l’intelligence quand on est capable de prendre la bonne décision avec un minimum d’information à sa disposition. L’IA, telle qu’on la connaît de nos jours, n’est donc en rien “intelligente”. Car, si l’on réduit la taille de la BD à un contenu minimal, plus rien ne sort. Avec des êtres humains, c’est totalement différent. Certains d’entre eux, appelés “génies” sont capables de faire des merveilles, voire des miracles, avec trois fois rien comme données. La raison en est simple, et l’on retrouve ici la possibilité de décider par intuition, et non par raisonnement.

Conclusion

On peut hautement être impressionné par les performances actuelles de l’intelligence artificielle. Mais, qu’il soit bien clair qu’il s’agit ici de performance technologique sous réserve d’un accès sans limites à l’électricité. Rien n’empêche de rêver à un monde dans lequel les machines remplaceraient les êtres biologiques fonctionnant à l’eau. Mais, un tel monde serait extrêmement peu efficace et d’une viabilité nulle à long terme.

Plutôt que de parler d’intelligence artificielle, on ferait peut-être mieux de parler de machines expertes dans le traitement de l’information. Car, on peut exceller dans la collecte et l’analyse d’un maximum d’information, tout en étant un parfait idiot sur le plan de la créativité. La folie actuelle autour de tous ces robots hyper-efficaces, ne pourra pas durer éternellement. Alors que la vie biologique, bactérienne ou cellulaire, existe depuis des milliards d’années. Par conséquent, s’il y a lieu d’être inquiets pour l’avenir, c’est plutôt face aux manipulations génétiques ou à la pollution toujours croissante de l’eau.

The post 85·Faut-il craindre l’intelligence artificielle (IA) ? appeared first on Marc HENRY - Natur'Eau Quant.

Silicon Valley’s unlikely ally to defend generative AI

Par : Paris Marx
14 juin 2024 à 14:00
Silicon Valley’s unlikely ally to defend generative AI

When ChatGPT was released in November 2022, many artists and publishers quickly realized another technological threat was on the horizon. Boosters were enthusiastic about how generative AI tools could churn out writing and images that would, in the process, devalue the work of writers, journalists, illustrators, and many other professions. But it all depended on the ability of those companies to take the collective knowledge of billions of people from the open web to feed into their artificial intelligence (AI) models.

Immediately, many of those professions who felt threatened saw an opportunity they could seize to present a legal speed bump for the seeming inevitability of AI adoption. Over the past year and a half, a series of groups that include visual artists, authors, and news organizations have launched copyright lawsuits directed at that vulnerability. The US Copyright Office has also refused to grant copyright protection to works created with generative AI tools — a campaign that has been supported by groups like the Writers Guild of America, SAG-AFTRA, and the Authors Guild.

If governments or the courts affirm the demands of artists and media groups, AI companies would have to remove copyrighted works from their training data or develop complex new systems to identify and compensate the owners of those copyrights. That’s something they want to avoid at all costs. In recent months, OpenAI has been proactively making deals with news publishers to gain access to their archives, but the industry isn’t stopping there.

Last week, tech lobby group Chamber of Progress launched a campaign called “Generate and Create” to ensure it remains legal for AI companies to train their models on copyrighted works. The group represents Amazon, Apple, Meta, Google, and many others. Their argument is not only that generative AI “lowers barriers for producing art,” but that training on existing works should be considered fair use. They’ve found an unlikely ally in the argument: anti-copyright activists. After taking on film and music industry lobbyists in service of the little guy in the 2000s, they’ve flipped as their libertarian politics push them defend some of the biggest and most exploitative companies in the world.

Apple hopes AI will make you buy a new iPhone

Par : Paris Marx
11 juin 2024 à 13:00
Apple hopes AI will make you buy a new iPhone

After months of obsessing by tech media, Apple’s generative AI features are (almost) here. In a prerecorded keynote at its Worldwide Developers’ Conference (WWDC) on Monday, company executives ran through a series of features like proofreading and changing the tone of text, recording and transcribing phone calls, a more capable Siri voice assistant, and an emoji generator which all fell under the label of “Apple Intelligence” — its branding term for generative AI. But none of them were anything to get too excited about.

The biggest takeaway from Apple’s keynote is that the air continues to flow out of the generative AI bubble. Last month, Google and OpenAI had their own lackluster showcases to demonstrate a series of evolutionary and even mundane features that were hard for anyone not steeped in AI boosterism to get excited about. One of the standout subpar moments of the Google event was when the company showed off how its Gemini chatbot could find a user’s license plate number in the Photos app as if it was a huge deal. Apple virtually recreated it, but with a driver’s license number instead.

AI hype is over. AI exhaustion is setting in.
Google and OpenAI’s latest showcases suggest the AI bubble’s days are numbered
Apple hopes AI will make you buy a new iPhoneDisconnectParis Marx
Apple hopes AI will make you buy a new iPhone

Unlike what OpenAI and Google were doing last year, Apple clearly isn’t trying to make people think its generative AI features are going to change the world. If we weren’t in this moment of AI hype driven by investor exuberance rather than tangible technological progress, these features would just be like any other additions to Apple’s operating systems. But they have to be singled out and given some extra fanfare because that’s usually what helps boost a company’s share price these days — or at least protects it from sliding.

After all the build up, Apple’s stock dropped in after market trading but jumped on Tuesday. The company clearly has a strategy to use its Apple Intelligence features to try to address some of the deeper problems with its business. It just remains to see if it will work, and if some of those generative AI features it’s banking on to deliver a boost might come back to bite it.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Apple’s ambitions still face challenges

Last month, Apple was in the crosshairs of artists when it released an advertisement for the new iPad Pro that showed a whole range of creative tools being destroyed to give way to the thin black slab of glass and metal. The company meant to illustrate how much creative work it believed could be done on the device, but after over a year of concern about how tech companies are deploying generative AI to take work away from artists, it sparked an uproar. As Apple rolls out its text and image generation features, the company risks more creative ire.

The image generator built into Apple Intelligence contains the strange and glossy aesthetic we’ve come to expect from AI-generated visuals. Not to mention that the use cases the feature is supposed to make possible — like being able to send personalized generated images in text messages — seem wholly unnecessary. But the worst part is that Apple isn’t being open about where it’s getting the data that powers its models.

In a news release, the company said it used licensed data and “publicly available data collected by our web-crawler,” with the option for publishers to opt out of having their content gathered as training data. It also assured users that their personal data and interactions are not being used. But the company didn’t say whether copyrighted works were absorbed into its training data without creators’ permission, suggesting it likely followed similar methods as other tech companies. Just days ago, Chamber of Progress launched a campaign to defend the legality of tech companies using copyrighted works in AI training data. Apple is a member of that lobby group.

On top of the potential anger from publishers and artists, Apple could face a public backlash once people actually get their hands on its generative AI features. After Google’s I/O event in May, it was pilloried for the terrible AI Overviews it enabled on Search that gave out an endless stream of incorrect and even dangerous answers. We’ve yet to see whether Siri will have its own controversial recommendations to make or if Apple’s built-in image generator will start spitting out offensive AI slop like so many others. That will be tested in the coming months.

Privacy is also still an open question, but the one it may be least likely to face criticism on. Apple has long presented itself as being uniquely committed to protecting user privacy, even as some groups have called those claims into question for not being as solid as the company suggests. The company’s AI features have previously done a lot of their processing directly on the device, but with generative AI it will be sending some tasks to servers it says will be highly secure, so much that it’s calling them “Private Cloud Compute.”

After looking over Apple’s breakdown of the system, cryptographer and Johns Hopkins University professor Matthew Green wrote on Twitter that “if you gave an excellent team a huge pile of money and told them to build the best ‘private’ cloud in the world, it would probably look like this.” Even with that praise, he warned there will still be a lot of vulnerabilities that could be difficult for researchers for detect. Apple also still hasn’t said whether users will be able to fully opt out of the feature if they choose to. But those protections only apply to Apple Intelligence. If a user wants to use the ChatGPT integration, almost all of that goes out the window and users must agree to share data with OpenAI on far less secure terms.

AI to drive iPhone sales

One thing you can say for Apple’s generative AI plans is that it does seem to have better curated the use cases it’s targeting instead of trying to throw everything at the wall, like some of its competitors. That doesn’t eliminate the broader problems that critics have been identifying with generative AI for more than a year, but the framing of the features does signal how Apple is hoping to use them.

The company isn’t positioning Apple Intelligence as a new subscription that users can sign up for to boost its digital services revenue. Instead, it’s just another set of features made available to people who buy its hardware — like its suite of office software. That not only further dispels the notion that generative AI is some groundbreaking development, but shows Apple is positioning the features to help it drive hardware sales — and specifically a new cycle of iPhone upgrades.

The company has been struggling with plateauing iPhone unit sales for years. It increased sales revenue by hiking prices through the introduction of its iPhone X and later iPhone Pro models, but even that stopped working last year when it had four consecutive quarters of declining revenue. Recently, iPhone sales in China saw a recovery, but that was driven by deep discounts.

Apple Intelligence will be limited to iPads and Macs with M1 chips, meaning certain models released in the last few years, but that’s even more constrained for iPhones. The only existing models that will be able to run the generative AI features when they’re released are the premium 15 Pro and Pro Max models that came out last fall, meaning the WWDC keynote was essentially a message to consumers that they’d better get ready to upgrade if they want the latest and greatest AI tools on their phone.

Roundup: Apple wants to hike iPhone prices again
Read to the end for an important statement about X… or Twitter?
Apple hopes AI will make you buy a new iPhoneDisconnectParis Marx
Apple hopes AI will make you buy a new iPhone

Apple Intelligence will be the central part of a big push to drive an iPhone upgrade supercycle in the fall to deliver good financial results for the fiscal year. Rumors suggest Apple is even planning to launch a new model alongside the iPhone 16 series that’s tentatively named the iPhone Slim and will be even more expensive than the Pro Max.

Apple Intelligence is not the company’s “next big thing.” After the Vision Pro bombed and the car project was canceled, generative AI may help drive some additional sales in its existing hardware lineup, but Apple still seems as paralyzed by its success as it’s been for some time. It’s lucky chatbots and image generators aren’t the pivotal moment Sam Altman wanted us to believe, or the company might have really been in trouble.

UPDATE (June 11, 2023): Clarified Apple’s share price movements after markets opened.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Kara Swisher's story about Sam Altman is falling apart

Par : Paris Marx
31 mai 2024 à 16:54
Kara Swisher's story about Sam Altman is falling apart

When Sam Altman was ousted as CEO of OpenAI on November 17, 2023, Kara Swisher started tweeting up a storm of “scoopage,” as she referred to her calls with high-ranking tech figures. Over the days Altman was on the outside, Swisher helped to craft a narrative that a board stacked with his internal rivals had pulled off a coup without a legitimate reason. The face of the AI boom had been betrayed and deserved to retake his position at the helm.

Swisher has been open about the fact she personally likes Altman. In her memoir Burn Book, she wrote, “I have come to know and like the young entrepreneur since I met him in 2005.” He also joined Swisher for an event in San Francisco to promote the book earlier this year. During his ouster, Swisher’s narrative was the story Altman wanted told. It cast him as the victim and helped to set the foundation for his triumphant return — one that would increase his power over the company and serve key investors like Microsoft.

In Swisher’s framing, which she repeats in her book, OpenAI was riven by divisions between AI optimists like Altman and doomers like cofounder and chief scientist Ilya Sutskever who believed artificial general intelligence (AGI) was on the horizon, presented a threat to humanity, and that the company wasn’t doing enough to mitigate that threat. The doomer camp took its opportunity to oust Altman, but misjudged its power and Altman was restored days later.

Kara Swisher’s Reality Distortion Field
In “Burn Book,” the longtime tech journalist tries to rewrite her story for the post-techlash era
Kara Swisher's story about Sam Altman is falling apartDisconnectParis Marx
Kara Swisher's story about Sam Altman is falling apart

In her commentary, Swisher often focused on a statement that Altman hadn’t been “candid” in his communications with the board, using it to suggest there was no real reason for Altman’s removal beyond factionalism, while claims he was “manipulative and headstrong” were dismissed as “sound[ing] like a typical [Silicon Valley] CEO to me.” She even went on CNN to call OpenAI’s board “incompetent” and to advocate for their resignation to open the way for Altman’s reinstatement.

Internal division was surely one factor in the board’s decision, but Swisher’s narrative and outright advocacy for Altman skillfully downplayed other factors whose details started to become clearer after his swift return as CEO. Those allegations focused more on Altman’s conduct and treatment of employees — things Swisher has a long history of ignoring as part of her “prick-to-productivity ratio.” Altman was another “prick” she cared more about being close to than seeing exposed.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Why Altman was really ousted

The day after Altman was restored as CEO, the Washington Post published a report into the side of his history that people like Swisher had little time for. Sources told the outlet the board was concerned that Altman was trying to remove checks on his power at OpenAI and had “a pattern of consistent and subtle manipulation that sows division between individuals.” A couple weeks later, journalist Nitasha Tiku built on that earlier reporting with more details suggesting Altman had been “psychologically abusive” and had allegedly been “pitting employees against each other in unhealthy ways.” Those actions had been central considerations for the board when they decided to remove him.

In 2019, Altman had also been removed as head of Y Combinator for putting “his own interests ahead of the organization,” according to the Post. At the time, he was making personal investments in companies he was selecting for inclusion in the accelerator’s fund. Even going back to his first startup Loopt, the management team had reportedly asked the board to remove him as CEO twice for “deceptive and chaotic behavior,” according to former OpenAI board member Helen Toner. Those details are often left out of the glowing profiles Altman receives today, not unlike how figures like Musk had their histories effectively rewritten once they became successful.

Elon Musk wants to relive his start-up days. He’s repeating the same mistakes.
What PayPal’s history teaches us about Elon Musk’s management of Twitter
Kara Swisher's story about Sam Altman is falling apartDisconnectParis Marx
Kara Swisher's story about Sam Altman is falling apart

During the ouster saga, OpenAI employees released a letter saying they’d resign en masse if Altman wasn’t restored as CEO. But the Washington Post found employees reported facing peer pressure to add their names and had a financial incentive to bring him back. If he wasn’t, an investment deal that would allow them to cash out would be put in jeopardy.

More recently, Vox revealed that employees were forced to sign strict non-disclosure and non-disparagement agreements once they left he company, or else they could lose their vested equity. Altman claimed he didn’t know about those clauses, but his response doesn’t line up with documents Vox obtained suggesting he would have been very aware and that company lawyers were aggressive in getting ex-employees to sign away their ability to speak about or criticize the company.

A recent interview Toner gave to the TED AI Show casts further doubt on Swisher’s version of events. Toner elaborated on the reasons Altman was removed from the company, explaining that two executives had spoken to the board about negative experiences with Altman that included lying, manipulation, and creating a toxic atmosphere at the company. Altman also didn’t disclose his ownership of the OpenAI Startup Fund and lied about the company’s safety processes. The board felt it “just couldn’t believe things that Sam was telling us,” explained Toner. She was also personally targeted by Alman when she published a research paper he didn’t like.

Ultimately, Toner explained the quick and stealthy removal of Altman as being necessary because if he knew in advance he’d “pull out all the stops, do everything in his power to undermine the board.” Those are clearly well-founded concerns, given that’s exactly what Altman did after he was ousted. He muddied the waters and ensured his narrative of events was the one that dominated the discussion about the company. Swisher served as a central conduit in that campaign.

Altman is cementing his power

After returning to OpenAI, Altman has cemented his power over the company. He brought in a new board that will be far less likely to challenge to his leadership. Earlier this month, Sutskever left the company, as did safety lead and company executive Jan Leike. OpenAI also dissolved the Superalignment team that Leike headed up. On Twitter, Leike explained that his disagreements with OpenAI leadership had been growing “for quite some time” and he was concerned “safety culture and processes have taken a backseat to shiny products.” Vox reported that five more of the company’s safety-conscious employees had left since Altman’s return as CEO as they lost trust in his leadership.

Scarlett Johansson can’t escape tech exploitation
OpenAI grabbing her voice is just the latest scandal that demands a regulatory response
Kara Swisher's story about Sam Altman is falling apartDisconnectParis Marx
Kara Swisher's story about Sam Altman is falling apart

To replace that team, OpenAI created a safety and security committee earlier this week stacked with company executives, including Altman himself. Meanwhile, the prospect of OpenAI throwing off the shackles of its non-profit status continues to be discussed and the company has reported branched out from its Microsoft dependency to sign a deal with Apple that will give it a greater sense of independence. Even in the face of the blowback from the non-disclosure scandal and the potential legal action from the drama over the use of a voice sounding very similar to Scarlett Johansson’s in GPT-4o, Altman’s power is only growing.

While Swisher might make the occasional jabs at the “man-boys” of Silicon Valley, as she refers to them in her book, she’s desperate to be in their circles and part of the power game. She’s chosen a few CEOs to more regularly criticize now that they don’t give her the access she craves, but there are many more whose narratives she will happily help turn into the official record whenever it will help them. Altman is in that group, and six months after his removal and return as CEO of OpenAI, it’s become very apparent that Swisher was echoing the Altman line and defending his interests. As she says in her book, she’s happy to give “flawed people … a little break,” as long as they have the power to justify it.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Scarlett Johansson can’t escape tech exploitation

Par : Paris Marx
22 mai 2024 à 13:16
Scarlett Johansson can’t escape tech exploitation

Less than a week after OpenAI showed off a demo of its GPT-4o with a voice that sounded so strikingly similar to Scarlett Johansson’s that Sam Altman even went to far as to tweet “her” — a reference to the 2013 Spike Jonze movie of the same name — we found out it wasn’t a coincidence at all. According to a statement released by Johansson, Altman had been trying to get the actress to be the voice of the product as far back as last September. He was seemingly desperate to make his favorite movie a reality — or at least a version of it that provided the aesthetic Altman needed to keep pushing the fantasy that his large language models will soon develop consciousness.

her

— Sam Altman (@sama) May 13, 2024

According to Johansson, Altman pitched her on the idea that putting her voice behind the chatbot would “bridge the gap between tech companies and creatives and help consumers to feel comfortable with the seismic shift concerning humans and Al.” Probably more important to Altman, hearing Johansson would be “comforting to people” — or at least the tech bros trying to recreate Her. Ultimately, Johansson turned down the offer. But Altman wasn’t done.

Two days before OpenAI’s Spring Update focused on GPT-4o, he contacted Johansson’s agent once again, asking her to reconsider. But he didn’t even wait for a response. They company went ahead with the demo, prompting questions for OpenAI about how much it sounded like Johansson, jokes on Saturday Night Live referencing it, and even the actress’ close friends to message her asking if she was involved. OpenAI has now pulled the “Sky” voice that imitates Johansson, but the whole debacle brings up a number of important issues.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

The internet’s dark side

Putting OpenAI aside for a moment, it’s hard to ignore how often Johansson gets caught up in these cycles of tech exploitation and ends up having to become the face of campaigns against it. Back in 2021, she sued Disney when it decided to put Black Widow on Disney+ the same day it was poised to hit theaters without consulting her. Part of Johansson’s compensation was tied to box office performance, so she could credibly argue she’d be paid less as a result of the decision, but it also happened to play into bigger concerns in the industry.

WarnerMedia had sent its 2021 slate of movies to streaming too, without bothering to talk to the actors and directors involved with the projects first. Those decisions heightened the growing dissatisfaction in Hollywood with the streaming model and the decisions being made by major studios. Johansson ultimately settled, but her decision to sue was seen as an important move in a growing fight between talent and studios — regardless of her A-list status. That fight came to a head last year when actors and writers went on a months-long strike demanding a better deal. But Johansson may have been primed for that legal fight due to the longer battle she’s been forced to wage against digital exploitation.

“The digital age is cannibalizing us”
Hollywood actors join writers in strike against companies using tech to degrade their profession
Scarlett Johansson can’t escape tech exploitationDisconnectParis Marx
Scarlett Johansson can’t escape tech exploitation

When explicit AI-generated photos of Taylor Swift spread across Twitter/X earlier this year, a lot more people woke up to the growing problem of non-consensual deepfake or AI-generated sexual images that generative AI tools have made much easier to create. But the problem didn’t begin with generative AI. In 2018, Johansson spoke out about deepfakes after fake videos that put her face on graphic sex scenes started spreading online and racking up millions of views. Her statement was scathing, calling the internet “a vast wormhole of darkness that eats itself.”

“The fact is that trying to protect yourself from the internet and its depravity is basically a lost cause, for the most part,” she wrote at the time. “Obviously, if a person has more resources, they may employ various forces to build a bigger wall around their digital identity. But nothing can stop someone from cutting and pasting my image or anyone else’s onto a different body and making it look as eerily realistic as desired.” For all the benefits that came of online connection, the internet continues to have a dark underbelly that often gets minimized by its biggest defenders.

By that point, Johansson was already well-acquainted with the ways the internet could be used to more easily victimize women, whether celebrities like her or just regular everyday people. Years earlier, she’d experienced having men make a life-sized robot using her face without her permission. Her email account was also been hacked and nude photos that were taken from it were later published online. (The hacker eventually got ten years in prison.) As the generative AI boom has taken off, she’s been dealing with it once again, having already sued Lisa AI for using her likeness to promote its product. Now she’s being forced to turn her attention to OpenAI.

Time for action

Since generative AI tools started taking off in November 2022, companies like OpenAI have been playing fast and loose with the rules that govern the use of copyrighted works and people’s personal data. We’ve seen countless examples of AI image generators churning out visuals that look remarkably similar to major franchise films or the graphic styles of specific artists, and voice actors have reported having their voices hijacked by AI voice generators. A proposed class action was launched last week against an AI company called LOVO which is accused of doing just that. Johansson, once again, is among those whose voices were stolen.

AI hype is over. AI exhaustion is setting in.
Google and OpenAI’s latest showcases suggest the AI bubble’s days are numbered
Scarlett Johansson can’t escape tech exploitationDisconnectParis Marx
Scarlett Johansson can’t escape tech exploitation

In the same way that Taylor Swift’s victimization at the hands of people generating sexual images of her put a spotlight on the issue, Altman’s ham-fisted attempt to build an AI assistant with Johansson’s voice will now do the same for a whole range of other abuses by companies like his own. The scandal has prompted other artists to speak out in support of Johansson and has made lawmakers take note of the problem. The reality is that the issue of Johansson’s voice is just the tip of the iceberg, and as the hype dissipates, this could be an important moment in the deflation of the AI bubble.

In her statement, Johansson connects what OpenAI did to her to those broader issues. “In a time when we are all grappling with deepfakes and the protection of our own likeness, our own work, our own identities, I believe these are questions that deserve absolute clarity,” she writes. Ultimately, she demands transparency from OpenAI and for lawmakers to move forward with regulation to protect individuals’ rights in the face of AI companies trampling all over them.

This scandal presents an opportunity for voices that have already been trying to raise the issue of theft and exploitation by AI companies to seize the spotlight Johansson has placed on generative AI and demand lawmakers take action. Protecting people’s voices and likenesses is important, but that effort can go much further to take on AI companies training on people’s work without their permission and the broader victimization their tools enable. It’s time to stop falling for Altman’s con and clean up the mess generative AI tools are making.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

AI hype is over. AI exhaustion is setting in.

Par : Paris Marx
17 mai 2024 à 15:18
AI hype is over. AI exhaustion is setting in.

At the end of Google’s snorefest of an I/O presentation on Tuesday, CEO Sundar Pichai took the stage to inform the crowd his team had used the term “AI” 120 times over the preceding two hours. He then said it one more time for good measure, so the counter would tick over to 121 and the audience could clap like a bunch of trained seals. The tired display wasn’t just an illustration of how fully Silicon Valley has embraced the latest hype cycle, but also how exhausting the whole thing has become a year and a half after the release of ChatGPT.

Throughout the keynote, Google executives showed off a ton of pretty standard features that were supposed to seem far more impressive because they had some link to Gemini, the name for its current suite of large language models (LLMs). The lemmings in the audience celebrated such groundbreaking features as getting Google Photos to find your license plate number, having a chatbot process the return for some shoes you ordered online, and getting Gemini to throw together some spreadsheets for you. The revolutionary nature of the AI future just continues to astound.

The company also brought DeepMind’s Demis Hassabis — excuse me, Sir Demis Hassabis — up to do some AI boosting of his own. That including repeating a debunked claim that Google’s AI tools had discovered a ton of “new materials” last year. Researchers who reviewed a subset of Google’s data concluded “we have yet to find any strikingly novel compounds,” with one telling 404 Media, “the Google paper falls way short in terms of it being a useful, practical contribution to the experimental materials scientists.” A little later, they got Donald Glover to lend some credibility to Google’s AI video generation efforts, making the dubious claim that it will let anyone become a director — just like we were told with home video cameras, smartphones, and other technologies that did little to break down Hollywood’s gates.

How Hollywood used the digital transition against workers
The challenge facing striking workers goes much deeper than streaming and AI
AI hype is over. AI exhaustion is setting in.DisconnectParis Marx
AI hype is over. AI exhaustion is setting in.

Gone are the days when tech executives could reasonably make us believe artificial general intelligence (AGI) was on the horizon because they’d thrown so much capital and compute behind getting their models to do slightly more advanced work than they’d done before. They certainly can’t scare us any longer with the idea that sentient AIs are on the cusp of enslaving us, if not just killing us all. Beyond the massively oversold new features that might not even need an LLM in the first place, the leaders of this AI push are getting lost in their own fantasies.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Take OpenAI CEO Sam Altman. He’s recently been making the podcast rounds sounding as though the past year of AI proselytizing has made him lose his remaining grasp on reality. Earlier this week, he appeared on the Logan Bartlett Show where he claimed AI will create a huge market for “human, in-person, fantastic experiences” — something that definitely doesn’t exist in the metaverses Altman must think we currently live in. A few days earlier, he joined the increasingly radicalized bros at All-In to explain that universal basic income is dead; instead we should strive for a universal basic compute where everyone would get a share in an LLM to use or sell as they see fit. It had big “let them eat compute” energy.

OpenAI had a showcase of its own this week that was similarly underwhelming, despite how much the company’s executives and parts of the media tried to pretend otherwise. As a few members of the OpenAI team showed off a ChatGPT voice bot doing math equations, responding to basic requests, and doing a very short voice translation demo with a flirty, female tone that got commentators obsessively comparing it to the movie Her, the tool itself kept making mistakes and unrelated comments that the presenters had to awkwardly laugh off. Over at Intelligencer, John Herrman astutely observed that AGI is no closer to being achieved, so OpenAI is getting its tools to “perform the part of an intelligent machine.”

Going back to the 1960s, we know humans are suckers for computers that pretend to be thinking machines, even when they’re doing nothing of the sort. Leading the update, OpenAI CTO Mira Murati said its updated chatbot just “feels so magical.” But as we know, magic is an illusion that always has an explanation. OpenAI just doesn’t want us to know what’s happening below the hood of its tools and keep buying the fantasy as long as possible. Murati is the same woman who became a meme back in March for the expression she pulled when asked in a Wall Street Journal interview if YouTube videos had been used to train OpenAI’s Sora video generator.

The reality is that no matter how much OpenAI, Google, and the rest of the heavy hitters in Silicon Valley might want to continue the illusion that generative AI represents a transformative moment in the history of digital technology, the truth is that their fantasy is getting increasingly difficult to maintain. The valuations of AI companies are coming down from their highs and major cloud providers are tamping down the expectations of their clients for what AI tools will actually deliver. That’s in part because the chatbots are still making a ton of mistakes in the answers they give to users, including during Google’s I/O keynote. Companies also still haven’t figured out how they’re going to make money off all this expensive tech, even as the resource demands are escalating so much their climate commitments are getting thrown out the window.

AI is fueling a data center boom. It must be stopped.
Silicon Valley believes more computation is essential for progress. But they ignore the resource burden and don’t care if the benefits materialize.
AI hype is over. AI exhaustion is setting in.DisconnectParis Marx
AI hype is over. AI exhaustion is setting in.

This whole AI cycle was fueled by fantasies, and when people stop falling for them the bubble starts to deflate. In The Guardian, John Naughton recently laid out the five stages of financial bubbles, noting AI is between stages three and four: euphoria and profit-taking. Tech companies like Microsoft and Google are still spending big to maintain the illusion, but it’s hard to deny that savvy investors see the writing on the wall and are planning their exits, if they haven’t already begun them to avoid being wiped out. The fifth stage — panic — is where we’re headed next.

That doesn’t mean generative AI will disappear. Think back to the last AI cycle in the mid-2010s when robots and AI were supposed to take all our jobs and make us destitute. The fantasy of self-driving cars is still limping along and some of those tools became entrenched, particularly the algorithmic management techniques used to carve workers out of labor protections and make it harder for them to put up a fight against bosses like Amazon and Uber.

Even though the excitement around generative AI is giving way to exhaustion, that doesn’t mean the companies behind these tools aren’t still trying to expand their power over how we used digital technology. It’s quite clear that Google is trying to further sideline the open web by ingesting it into its model then expecting people to spend even more time on its platforms than anywhere else. Things aren’t so different with OpenAI, where they’re hoping to revive the failed voice assistant push that followed the last moment of AI hype and get people used to depending on ChatGPT for virtually everything they do.

Google wants to take over the web
Its new plan for search shows how AI hype hides the real threat of increased corporate power
AI hype is over. AI exhaustion is setting in.DisconnectParis Marx
AI hype is over. AI exhaustion is setting in.

Between those visions, the Google one feels far more threatening because of the structural transformation it hopes to carry out that will further platformize our online experience, at a moment when people are feeling increasingly frustrated with the state of the internet as the services we’ve come to depend on further erode under pressure to maximize profits. But that doesn’t mean OpenAI’s efforts should be ignored. With the backing of Microsoft, it wants to sell people an illusion of intelligence to get them to take their guards down for a power play of its own.

Each of these companies present a threat in their own way, but there might be some solace in the recognition that the return to AI winter is inevitable — and the crash that’s coming could be unlike one we’ve seen in quite some time. The question is how long companies will keep spending increasingly vast sums to pump a little more hot air into the bubble to delay its total deflation. Seeing a CEO count the number of times his underlings said “AI” during a keynote suggests they’re already running on fumes.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Why Does Biological Evolution Work? A Minimal Model for Biological Evolution and Other Adaptive Processes

3 mai 2024 à 20:40

The Model

Why does biological evolution work? And, for that matter, why does machine learning work? Both are examples of adaptive processes that surprise us with what they manage to achieve. So what’s the essence of what’s going on? I’m going to concentrate here on biological evolution, though much of what I’ll discuss is also relevant to machine learning—but I’ll plan to explore that in more detail elsewhere.

OK, so what is an appropriate minimal model for biology? My core idea here is to think of biological organisms as computational systems that develop by following simple underlying rules. These underlying rules in effect correspond to the genotype of the organism; the result of running them is in effect its phenotype. Cellular automata provide a convenient example of this kind of setup. Here’s an example involving cells with 3 possible colors; the rules are shown on the left, and the behavior they generate is shown on the right:

Note: Click any diagram to get Wolfram Language code to reproduce it.

We’re starting from a single () cell, and we see that from this “seed” a structure is grown—which in this case dies out after 51 steps. And in a sense it’s already remarkable that we can generate a structure that neither goes on forever nor dies out quickly—but instead manages to live (in this case) for exactly 51 steps.

But let’s say we start from the trivial (“null”) rule that makes any pattern die out immediately. Can we end up “adaptively evolving” to the rule above? Imagine making a sequence of randomly chosen “point mutations”—each changing just one outcome in the rule, as in:

Then suppose that at each step—in a minimal analog of natural selection—we “accept” any mutation that makes the lifetime longer (though not infinite), or at least the same as before, and we reject any mutation that makes the lifetime shorter, or infinite. It turns out that with this procedure we can indeed “adaptively evolve” to the rule above (where here we’re showing only “waypoints” of progressively greater lifetime):

Different sequences of random mutations give different sequences of rules. But the remarkable fact is that in almost all cases it’s possible to “make progress”—and routinely reach rules that give long-lived patterns (here with lifetimes 107, 162 and 723) with elaborate morphological structure:

Is it “obvious” that our simple process of adaptive evolution will be able to successfully “wrangle” things to achieve this? No. But the fact that it can seems to be at the heart of why biological evolution manages to work.

Looking at the sequences of pictures above we see that there are often in effect “different mechanisms” for producing long lifetimes that emerge in different sequences of rules. Typically we first see the mechanism in simple form, then as the adaptive process continues, the mechanism gets progressively more developed, elaborated and built on—not unlike what we often appear to see in the fossil record of biological evolution.

But let’s drill down and look in a little more detail at what’s happening in the simple model we’re using. In the 3-color nearest-neighbor (k = 3, r = 1) cellular automata we’re considering, there are 26 (= 33 – 1) relevant cases in the rule (there’d be 27 if we didn’t insist that
). “Point mutations” affect a single case, changing it to one of two (= 3 – 1) possible alternative outcomes—so that there are altogether 52 (= 26 × 2) possible distinct “point mutations” that can be made to a given rule.

For example, starting from the rule

the results of possible single point mutations are:

And even with such point mutations there’s usually considerable diversity in the behavior they generate:

In quite a few cases the pattern generated is exactly the same as the one for the original rule. In other cases it dies out more quickly—or it doesn’t die out at all (either becoming periodic, or growing forever). And in this particular example, in just one case it achieves “higher fitness”, surviving longer.

If we make a sequence of random mutations, many will produce shorter-lived or infinite lifetime (“tumor”) patterns, and these we’ll reject (or, in biological terms, we can imagine they’re “selected out”):

But still there can be many “neutral mutations” that don’t change the final pattern (or at least give a pattern of the same length). And at first we might think that these don’t achieve anything. But actually they’re critical in allowing single point mutations to build up to larger mutations that can eventually give longer-lived patterns:

Tracing our whole adaptive evolution process above, the total number of point mutations involved in getting from one (increasingly long-lived) “fitness waypoint” to another is:

Here are the underlying rules associated with these fitness waypoints (where the numbers count cumulative “accepted mutations”, ignoring ones that go “back and forth”):

One way to get a sense of what’s going on is to take the whole sequence of (“accepted”) rules in the adaptive evolution process, and plot them in a dimension-reduced rendering of the (27-dimensional) rule space:

There are periods when there’s a lot of “wandering around” going on, with many mutations needed to “make progress”. And there are other periods when things go much faster, and fewer mutations are needed.

As another way to see what’s going on, we can plot the maximum lifetime achieved so far against the total number of mutation steps made:

We see plateaus (including an extremely long one) in which “no progress” is made, punctuated by sometimes-quite-large, sudden changes, often brought on by just a single mutation.

If we include “rejected mutations” we see that there’s a lot of activity going on even in the plateaus; it just doesn’t manage to make progress (one can think of each red dot that lies below a plateau as being like a mutation—or an organism—that “doesn’t make it”, and is selected out):

It’s worth noting that there can be multiple different (“phenotype”) patterns that occur across a plateau. Here’s what one sees in the particular example we’re considering:

But even between these “phenotypically different” cases, there can be many “genotypically different” rules. And in a sense this isn’t surprising, because usually only parts of the underlying rule are “coding”; other parts are “noncoding”, in the sense that they’re not sampled during the generation of the pattern from that rule.

And for example this highlights for each “fitness waypoint rule” which cells make use of a “fresh” case in the rule that hasn’t so far been sampled during the generation of the pattern:

And we see that even in the last rule shown here, only 18 of the 26 relevant cases in the rule are actually ever sampled during the generation of the pattern (from the particular, single-red-cell initial condition used). So this means that 8 cases in the rule are “undetermined” from the phenotype, implying that there are 38 = 6561 possible genotypes (i.e. rules) that will give the same result.

So far we’ve mostly been talking about one particular random sequence of mutations. But what happens if we look at many possible such sequences? Here’s how the longest lifetime (or, in effect, “fitness”) increases for 100 different sequences of random mutations:

And what’s perhaps most notable here is that it seems as if these adaptive processes indeed don’t “get stuck”. It may take a while (with the result that there are long plateaus) but these pictures suggest that eventually “adaptive evolution will find a way”, and one will get to rules that show longer lifetimes—as the progressive development of the distribution of lifetimes reflects:

The Multiway Graph of All Possible Mutation Histories

In what we’ve done so far we’ve always been discussing particular paths of adaptive evolution, determined by particular sequences of random mutations. But a powerful way to get a more global view of the process of adaptive evolution is to look—in the spirit of our Physics Project, the ruliad, etc.—not just at individual paths of adaptive evolution, but instead at the multiway graph of all possible paths. (And in making a correspondence with biology, multiway graphs give us a way to talk about adaptive evolution not just of individual sequences of organisms, but also populations.)

To start our discussion, let’s consider not the 3-color cellular automata of the previous section, but instead (nearest-neighbor) 2-color cellular automata—for which there are just 128 possible relevant rules. How are these rules related by point mutations? We can construct a graph of every possible way that one rule from this set can be transformed to another by a single point mutation:

If we imagine 5-bit rather than 7-bit rules, there are only 16 relevant ones, and we can readily see that the graph of possible mutations has the form of a Boolean hypercube:

Let’s say we start from the “null rule” . Then we enumerate the rules obtained by a single point mutation (and therefore directly connected to the null rule in the graph above)—then we see what behavior they produce, say from the initial condition ……:

Some of these rules we can view as “making progress”, in the sense that they yield patterns with longer lifetimes (not impressively longer, just 2 rather than 1). But other rules “make no progress” or generate patterns that “live forever”. Keeping only mutations that don’t lead to shorter or infinite lifetimes, we can construct a multiway graph that shows all possible mutation paths:

Although this is a very small graph (with just 15 rules appearing), we can already see hints of some important phenomena. There are “fitness-neutral” mutations that can “go both ways”. But there are also plenty of mutations that only “go one way”—because the other way they would decrease fitness. And a notable feature of the graph is that once one’s “committed” to a particular part of the graph, one often can’t reach a different one—suggesting an analogy to the existence of distinct branches in the tree of life.

Moving beyond 2-color, nearest-neighbor (k = 2, r = 1) cellular automata, we can consider k = 2, r = ones. A typical such cellular automaton is:

For k = 2, r = 1 there were a total of 128 (= 223 – 1) relevant rules. For k = 2, r = , there are a total of 32,768 (= 224 – 1). Starting with the null rule, and again using initial condition …, here are a few specific examples of adaptive evolution paths for such cellular automata:

And here is the beginning of the multiway graph for k = 2, r = rules—showing rules reached by up to two mutations starting from the null rule:

This graph contains many examples of “fitness-neutral sets”—rules that have the same fitness and that can be transformed into each other by mutations. A few examples of such fitness-neutral sets:

In the first case here, the “morphology of the phenotypic patterns” is the same for all “genotypic rules” in the fitness-neutral set. But in the other cases there are multiple morphologies within a single fitness-neutral set.

If we included all individual rules we’d get a complete k = 2, r = multiway graph with a total of 1884 nodes. But if we just include one representative from every fitness-neutral set, we get a more manageable multiway graph, with a total of 86 nodes:

Keeping only one representative from pairs of patterns that are related by left-right symmetry, we get a still-simpler graph, now with a total of 49 nodes:

There’s quite a lot of structure in this graph, with both divergence and convergence of possible paths. But overall, there’s a certain sense that different sections of the graph separate into distinct branches in which adaptive evolution in effect “pursues different ideas” about how to increase fitness (i.e. lifetime of patterns).

We can think of fitness-neutral sets as representing a certain kind of equivalence class of rules. There’s quite a range of possible structures to these sets—from ones with a single element to ones with many elements but few distinct morphologies, to ones with different morphologies for every element:

What about larger spaces of rules? For k = 2, r = 2 there are altogether about 2 billion (225 – 1) relevant rules. But if we choose to look only at left-right symmetric ones, this number is reduced to 524,288 (= 219). Here are some examples of sequences of rules produced by adaptive evolution in this case, starting from the null rule, and allowing only mutations that preserve symmetry (and now using a single black cell as the initial condition):

Once again we can identify fitness-neutral sets—though this time, in the vast majority of cases, the patterns generated by all members of a given set are the same:

Reducing out fitness-neutral sets, we can then compute the complete (transitively reduced) multiway graph for symmetric k = 2, r = 2 rules (containing a total of 60 nodes):

By reducing out fitness-neutral sets, we’re creating a multiway graph in which every edge represents a mutation that “makes progress” in increasing fitness. But actual paths of adaptive evolution based on random sequences of mutations can do any amount of “rattling around” within fitness-neutral sets—not to mention “trying” mutations that decrease fitness—before reaching mutations that “make progress”. So this means that even though the reduced multiway graph we’ve drawn suggests that the maximum number of steps (i.e. mutations) needed to adaptively evolve from the null rule to any other is 9, it can actually take any number of steps because of the “rattling around” within fitness-neutral sets.

Here’s an example of a sequence of accepted mutations in a particular adaptive evolution process—with the mutations that “make progress” highlighted, and numbers indicating rejected mutations:

We can see “rattling around” in a fitness-neutral set, with a cycle of morphologies being generated. But while this represents one way to reach the final pattern, there are also plenty of others, potentially involving many fewer mutations. And indeed one can determine from the multiway graph that an absolutely shortest path is:

This involves the sequence of rules:

We’re starting from the null rule, and at each step making a single point mutation (though because of symmetry two bits can sometimes be changed). The first few mutations don’t end up changing the “phenotypic behavior”. But after a while, enough mutations (here 6) have built up that we get morphologically different behavior. And after just 3 more mutations, we end up with our final pattern.

Our original random sequence of mutations gets to the same result, but in a much more tortuous way, doing a total of 169 mutations which often cancel each other out:

In drawing a multiway graph, we’re defining what evolutionary paths are possible. But what about probabilities? If we assume that every point mutation is equally likely, we can in effect “analyze the flow” in the multiway graph, and determine the ultimate probability that each rule will be reached (with higher probabilities here shown redder):

The Fitness Landscape

Multiway graphs give a very global view of adaptive evolution. But in understanding the process of adaptive evolution, it’s also often useful to think somewhat more locally. We can imagine that all the possible rules are laid out in a certain space, and that adaptive evolution is trying to find appropriate paths in this space. Potentially we can suppose that there’s a “fitness landscape” defined in the space, and that adaptive evolution is trying to follow a path that progressively ascends to higher peaks of fitness.

Let’s consider again the very first example we gave above—of adaptive evolution in the space of 3-color cellular automata. At each step in this adaptive evolution, there are 52 possible point mutations that can be made to the rule. And one can think of each of these mutations as corresponding to making an “elementary move” in a different direction in the (26-dimensional) space of rules.

Here’s a visual representation of what’s going on, based on the particular path of adaptive evolution from our very first example above:

What we’re showing here is in effect the sequence of “decisions” that are being made to get us from one “fitness waypoint” to another. Different possible mutations are represented by different radial directions, with the length of each line being proportional to the fitness achieved by doing that mutation. At each step the gray disk represents the previous fitness. And what we see is that many possible mutations lead to lower fitness outcomes, shown “within the disk”. But there are at least some mutations that have higher fitness, and “escape the disk”.

In the multiway graph, we’d trace every mutation that leads to higher fitness. But for a particular path of adaptive evolution as we’ve discussed it so far, we imagine we always just pick at random one mutation from this set—as indicated here by a red line. (Later we’ll discuss different strategies.)

Our radial icons can be thought of as giving a representation of the “local derivative” at each point in the space of rules, with longer lines corresponding to directions with larger slopes “up the fitness landscape”.

But what happens if we want to “knit together” these local derivatives to form a picture of the whole space? Needless to say, it’s complicated. And as a first example, consider k = 2, r = 1 cellular automaton rules.

There are a total of 128 relevant such rules, that (as we discussed above) can be thought of as connected by point mutations to form a 7-dimensional Boolean hypercube. As also discussed above, of all 128 relevant rules, only 15 appear in adaptive evolution processes (the others are in effect never selected because they represent lower fitness). But now we can ask where these rules lie on the whole hypercube:

Each node here represents a rule, with the size of the highlighted nodes indicating their corresponding fitness (computed from lifetime with initial condition ……). The node shown in green corresponds to the null rule.

Rendering this in 3D, with fitness shown as height, we get what we can consider a “fitness landscape”:

And now we can think of our adaptive evolution as proceeding along paths that never go to nodes with lower height on this landscape.

We get a more filled-in “fitness landscape” when we look at k = 2, r = rules (here with initial condition ……):

Adaptive evolution must trace out a “never-go-down” path on this landscape:

Along this path, we can make “derivative” pictures like the ones above to represent “local topography” around each point—indicating which of the possible upwards-on-the-landscape directions is taken:

The rule space over which our “fitness landscape” is defined is ultimately discrete and effectively very high-dimensional (15-dimensional for k = 2, r = rules)—and it’s quite challenging to produce an interpretable visualization of it in 3D. We’d like it if we could lay out our rendering of the rule space so that rules which differ just by one mutation are a fixed (“elementary”) 2D distance apart. In general this won’t be possible, but we’re trying to at least approximate this by finding a good layout for the underlying “mutation graph”.

Using this layout we can in principle make a “fitness landscape surface” by interpolating between discrete points. It’s not clear how meaningful this is, but it’s perhaps useful in engaging our spatial intuition:

We can try machine learning and dimension reduction, operating on the set of “rule vectors” (i.e. outcome lists) that won’t be rejected in our adaptive evolution process—and the results are perhaps slightly better:

By the way, if we use this dimension reduction for rule space, here’s how the behavior of rules lays out:

And here, for comparison, is a feature space plot based on the visual appearance of these patterns:

The Whole Space: Exhaustive Search vs. Adaptive Evolution

In adaptive evolution, we start, say, from the null rule and then make random mutations to try to reach rules with progressively larger fitness. But what about just exhaustively searching the complete space of possible rules? The number of rules rapidly becomes unmanageably big—but some cases are definitely accessible:

For example, there are just 524,288 symmetric k = 2, r = 2 rules—of which 77,624 generate patterns with finite lifetimes. Ultimately, though, there are just 77 distinct phenotypic patterns that appear—with varying lifetimes and varying multiplicity (where at least in this case the multiplicity is always associated with “unused bits” in the rule):

How do these exhaustive results compare with what’s generated in the multiway graph of adaptive evolutions? They’re almost the same, but for the addition of the two extra cases

which are generated by rules of the form (where the gray entries don’t matter):

Why don’t such rules ever appear in our adaptive evolution? The reason is that there isn’t a chain of point mutations starting from the null rule that can reach these rules without going through rules that would be rejected by our adaptive evolution process. If we draw a multiway graph that includes every possible “acceptable” rule, then we’ll see a separate part in the graph, with its own root, that contains rules that can’t be reached by our adaptive evolution from the null rule:

So now if we look at all (symmetric k = 2, r = 2) rules, here’s the distribution of lifetimes we get:

The maximum, as seen above, is 65. The overall distribution roughly follows a power law, with an exponent around –3:

As we saw above, not all rules make use of all their bits (i.e. outcomes) in generating phenotypic patterns. But what we see is that the larger the lifetime achieved, the more bits tend to be needed:

And in a sense this isn’t surprising: as we’ll discuss later, we can expect to need “more bits in the program” to specify more elaborate behavior—or, in particular, behavior that embodies a larger number for lifetime.

So what about general (i.e. not necessarily symmetric) k = 2, r = 2 rules? There are 231 ≅ 2 billion of these. If we exhaustively search through them, we find about 75 million that have finite lifetimes. The distribution of these lifetimes is:

Again it’s roughly a power law, but now with exponent around –3.5:

Here are the actual 100 patterns produced that have the longest lifetimes (in all asymmetric cases there are also rules giving left-right flipped patterns):

It’s interesting to see here a variety of “qualitatively different ideas” being used by different rules. Some (like the one with lifetime 151) we might somehow imagine could have been constructed specifically “for the purpose” of having their particular, long lifetime. But others (like the one with lifetime 308) somehow seem more “coincidental”—behaving in an apparently random way, and then just “happening to die out” after a certain number of steps.

Since we found these rules by exhaustive search, we know they’re the only possible ones with such long lifetimes (at least with k = 2, r = 2). So then we can infer that the ornate structures we see are in some sense necessary to achieve the objective of, say, having a finite lifetime of more than 100 steps. So that means that if we go through a process of adaptive evolution and achieve a lifetime above 100 steps, we see a complex pattern of behavior not because of “complicated choices in our process of adaptive evolution”, but rather because to achieve such a lifetime one has no choice but to use such a complex pattern. Or, in other words, the complexity we see is a reflection of “computational necessity”, not historical accidents of adaptive evolution.

Note also that (as we’ll discuss in more detail below) there are certain behaviors we can get, and others that we cannot. So, for example, there is a rule that gives lifetime 308, but none that gives lifetime 300. (Though, yes, if we used more complicated initial conditions or a more complicated family of rules we could find such a rule.)

Much as we saw in the symmetric k = 2, r = 2 case above, almost any long lifetimes require using all the available bits in the rule:

But, needless to say, there’s an exception—a pair of rules with lifetime 84 where the outcome for the case doesn’t matter:

But, OK, so can these long-lifetime rules be reached by single-mutation adaptive evolution from the null rule? Rather than trying to construct the whole multiway graph for the general k = 2, r = 2 case starting from the null rule, we can instead construct what’s in effect an inverse multiway graph in which we start from a given rule, then successively make all single point mutations that reach rules with progressively shorter lifetimes:

And what we see is that at least in this case such a procedure never reaches the null rule. The “furthest” it gets is to lifetime-2 rules, and among these rules the closest to the null rule are:

But it turns out that there’s no way to reach these 2-bit rules by a single point mutation from any of the 26 1-bit rules that aren’t rejected by our adaptive evolution process. And in fact this isn’t just an issue for this particular long-lifetime rule; it’s something quite general among k = 2 rules. And indeed, constructing the “forward” multiway graph starting from the null rule, we find we can only ever reach lifetime-1 rules.

Ultimately this is a particular feature of rules with just 2 colors—and it’s specific to starting with something like the null rule that has lifetime 1—but it’s an illustration of the fact that there can even be large swaths of rule space that can’t be reached by adaptive evolution with point mutations.

What about symmetric k = 2, r = 2 rules? Well, to maintain symmetry we have to deal with mutations that change not just one but two bits. And this turns out to mean that (except in the cases we discovered above) the inverse multiway system starting from long-lifetime rules always successfully reaches the null rule:

There’s something else to notice here, however. Looking at this graph, we see that there’s a way to get with just one 2-bit mutation from a lifetime-1 to a lifetime-65 rule:

We didn’t see this in our multiway graph above because we had applied transitive reduction to it. But if we don’t do that, we find that a few large lifetime jumps are possible—as we can see on this plot of possible lifetimes before and after a single point mutation:

Going beyond k = 2, r = 2 rules, we can consider symmetric k = 3, r = 1 rules, of which there are 317, or about 129 million. The distribution of lifetimes in this case is

which again roughly fits a power law, again with exponent around –3.5:

But now the maximum lifetime found is not just 308, but 2194:

One again, there are some different “ideas” on display, with a few curious examples of convergence—such as the rules we see with lifetimes 989 and 990 (as well as 1068 and 1069) which give essentially the same patterns after just exchanging colors, and adding one “prefatory” step.

What about general k = 3, r = 1 rules? There are too many to easily search exhaustively. But directed random sampling reveals plenty of long-lifetime examples, such as:

And now the tail of very long lifetimes extends further, for example with:

It’s a little easier to see what the lifetime-10863 rule does if one visualizes it in sections (and adjusts colors to get more contrast):

Sampling 100 steps out every 2000 (as well as at the very end), we see elaborate alternation between periodic and seemingly random behavior—but none of it gives any obvious clue of the remarkable fact that after 10863 steps the whole pattern will die out:

The Issue of Undecidability

As our example criterion for the “fitness” of cellular automaton rules, we’ve used the lifetimes of the patterns they generate—always assuming that if the patterns don’t terminate at all they should be considered to have fitness zero.

But how can we tell if a pattern is going to terminate? In the previous section, for example, we saw patterns that live a very long time—but do eventually terminate.

Here are some examples of the first 100 steps of patterns generated by a few k = 3, r = 1 symmetric rules:

What will happen with these patterns? We know from what we see here that none of them have lifetimes less than 100 steps. But what would allow us to say more? In a few cases we can see that the patterns are periodic, or have obvious repeating structures, which means they’ll never terminate. But in the other cases there’s no obvious way to predict what will happen. Explicitly running the rules for another 100 steps we discover some more outcomes:

Going to 500 steps there are some surprises. Rule (a) becomes periodic after 388 steps; rules (o) and (v) terminate after 265 and 377 steps, respectively:

But is there a way to systematically say what will happen “in the end” with all the remaining rules? The answer is that in general there is not; it’s something that must be considered undecidable by any finite computation.

Given how comparatively simple the cellular automaton rules we’re considering are, we might have assumed that with all our sophisticated mathematical and computational methods we’d always be able to “jump ahead of them”—and figure out their outcome without the computational effort of explicitly running each step.

But the Principle of Computational Equivalence suggests that pretty much whenever the behavior of these rules isn’t obviously simple, it will in effect be of equal computational sophistication to any other system, and in particular to any methods that we might use to predict it. And the result is the phenomenon of computational irreducibility that implies that in many systems—presumably including most of the cellular automata here—there isn’t any way to figure out their outcome much more efficiently than by explicitly tracing each of their steps. So this means that to know what will happen “in the end”—after an infinite number of steps—might take an unlimited amount of computational effort. Or, in other words, it must be considered effectively undecidable by any finite computation.

As a practical matter we might look at the observed distribution of lifetimes for a particular type of cellular automaton, and become pretty confident that there won’t be longer finite lifetimes for that type of cellular automaton. But for the k = 3, r = 1 rules from the previous section, we might have been fairly confident that a few thousand steps was the longest lifetime that would ever occur—until we discovered the 10,863-step example.

So let’s say we run a particular rule for 10,000 steps and it hasn’t died out. How can we tell if it never will? Well, we have to construct a proof of some kind. And that’s easy to do if we can see that the pattern becomes, say, completely periodic. But in general, computational irreducibility implies we won’t be able to do it. Might there, though, still be special cases where we can? In effect, those would have to correspond to “pockets of computational reducibility” where we manage to find a compressed description of the cellular automaton behavior.

There are cases like this where there isn’t strict periodicity, but where in the end there’s basically repetitive behavior (here with period 480):

And there are cases of nested behavior, which is never periodic, but is nevertheless simple enough to be predictable:

But there are always surprises. Like this example—which eventually resolves to have period 6, but only after 7129 steps:

So what does all this mean for our adaptive evolution process? It implies that in principle we could miss a very long finite lifetime for a particular rule, assuming it to be infinite. In a biological analogy we might have a genome that seems to lead to unbounded perhaps-tumor-like growth—but where actually the growth in the end “unexpectedly” stops.

Computation Theoretic Perspectives and Busy Beavers

What we’re asking about the dying out of patterns in cellular automata is directly analogous to the classic halting problem for Turing machines, or the termination problem for term rewriting, Post tag systems, etc. And in looking for cellular automata that have the longest-lived patterns, we’re studying a cellular automaton analog of the so-called busy beaver problem for Turing machines.

We can summarize the results we’ve found so far (all for single-cell initial conditions):

The profiles (i.e. widths of nonzero cells) for the patterns generated by these rules are

and the “integrals” of these curves are what give the “areas” in the table above.

For the reasons described in the previous section, we can only be certain that we’ve found lower bounds on the actual maximum lifetime—though except in the last few cases listed it seems very likely that we do in fact have the maximum lifetime.

It’s somewhat sobering, though, to compare with known results for maximum (“busy beaver”) lifetimes for Turing machines (where now s is the number of Turing machine states, the Turing machines are started from blank tapes, and they are taken to “halt” when they reach a particular halt state):

Sufficiently small Turing machines can have only modest lifetimes. But even slightly bigger Turing machines can have vastly larger lifetimes. And in fact it’s a consequence of the undecidability of the halting problem for Turing machines that the maximum lifetime grows with the size of the Turing machine faster than any computable function (i.e. any function that can be computed in finite time by a Turing machine, or whose value can be proved by a finite proof in a finite axiom system).

But, OK, the maximum lifetime increases with the “size of the rule” for a Turing machine, or a cellular automaton. But what defines the “size of a rule”? Presumably it should be roughly the number of independent bits needed to specify the rule (which we can also think of as an approximate measure of its “information content”)—or something like log2 of the number of possible rules of its type.

At the outset, we might imagine that all 232 k = 2, r = 2 rules would need 32 bits to specify them. But as we discussed above, in some cases some of the bits in the rule don’t matter when it comes to determining the patterns they produce. And what we see is that the more bits that matter (and so have to be specified), the longer the lifetimes that are possible:

So far we’ve only been discussing cellular automata with single-cell initial conditions. But if we use more complicated initial conditions what we’re effectively doing is adding more information content into the system—with the result that maximum lifetimes can potentially get larger. And as an example, here are possible lifetimes for k = 2, r = rules with a sequence of possible initial conditions:

Probabilistic Approximations?

Cellular automata are at their core deterministic systems: given a particular cellular automaton rule and a particular initial condition, every aspect of the behavior that is generated is completely determined. But is there any way that we can approximate this behavior by some probabilistic model? Or might we at least usefully be able to use such a model if we look at the aggregate properties of large numbers of different rules?

One hint along these lines comes from the power-law distributions we found above for the frequencies of different possible lifetimes for cellular automata of given types. And we might wonder whether such distributions—and perhaps even their exponents—could be found from some probabilistic model.

One possible approach is to approximate a cellular automaton by a probabilistic process—say one in which a cell becomes black with probability p if it or either of its neighbors were black on the step before. Here are some examples of what can happen with this (“directed percolation”) setup:

The behavior varies greatly with p; for small p everything dies out, while for large p it fills in:

And indeed the final density—starting from random initial conditions—has a sharp (phase) transition at around p = 0.54 as one varies p:

If instead one starts from a single initial black cell one sees a slightly different transition:

One can also plot the probabilities for different “survival times” or “lifetimes” for the pattern:

And right around the transition the distribution of lifetimes follows a power law—that’s roughly τ–1 (which happens to be what one gets from a mean field theory estimate).

So how does this relate to cellular automata? Let’s say we have a k = 2 rule, and we suppose that the colors of cells can be approximated as somehow random. Then we might suppose that the patterns we get could be like in our probabilistic model. And a potential source for the value of p to use would be the fraction of cases in the rule that give a black cell as output.

Plotting the lifetimes for k = 2, r = 2 rules against these fractions, we see that the longest lifetimes do occur when a little under half the outcomes are black (though notice this is also where the binomial distribution implies the largest number of rules are concentrated):

If we don’t try thinking about the details of cellular automaton evolution, but instead just consider the boundaries of finite-lifetime patterns we generate, we can imagine approximating these (say for symmetric rules) just by random walks—that when they collide correspond to the pattern dying out:

The standard theory of random walks then tells us that the probability to survive τ steps is proportional to τ–3/2 for large τ—a power law, though not immediately one of the same ones that we’ve observed for our cellular automaton lifetimes.

Other Adaptive Evolution Strategies

In what we’ve done so far, we’ve always taken each step of our process of adaptive evolution to pick an outcome of equal or greater fitness. But what if we adopt a “more impatient” procedure in which at each step we insist on an outcome that has strictly greater fitness?

For k = 2 it’s simply not possible with this procedure (at least with a null initial condition) to “escape” the null rule; everything that can be reached with 1 mutation still has lifetime 1. With k = 3 it’s possible to go one step, but only one, as captured by this multiway graph:

But we’re assuming here that we have to reach greater fitness with just one mutation. What if we allow two mutations at a time? Well, then we can “make progress”. And here’s the multiway graph in this case for symmetric k = 2, r = 2 rules:

We don’t reach as many phenotypic patterns as by using single mutations and allowing “fitness-neutral moves”, but where we do get, we get much quicker, without any “back and forth” in fitness-neutral spaces.

If we allow up to 3 mutations, we get still further:

And indeed we seem to get a pretty good representative sampling of “what’s out there” in this rule space, even though we reach only 37 rules, compared to the 77,624 (albeit with many duplicated phenotypic patterns) from our standard approach allowing neutral moves.

For k = 3, r = 1 symmetric rules single mutations can get 2 steps:

But now if we allow up to 2 mutations, we can go much further—and the fact that we now don’t have to deal with neutral moves means we can explicitly construct at least the first few steps of the multiway graph in this case:

We can go further if at each step we just pick a random higher-fitness rule reached with two or fewer mutations:

The adaptive evolution histories we just showed can be generated in effect by randomly trying a series of possibilities at each step, then picking the first one that exhibits increased fitness. Another approach is to use what amounts to “local exhaustive search”: at each step, look at results from all possible mutations, and pick one that gives the largest fitness. At least in smaller rule spaces, it’s common that there will be several results with the same fitness, and as an example we’ll just pick among these at random:

One might think that this approach would in effect always be an optimization of the adaptive evolution process. But in practice its systematic character can end up making it get stuck, in some sense repeatedly “trying to do the same thing” even if it “isn’t working”.

Something of an opposite approach involves loosening our criteria for which paths can be chosen—and for example allowing paths that temporarily reduce fitness, say by one step of lifetime:

In effect here we’re allowing less-than-maximally-fit organisms to survive. And we can represent the overall structure of what’s happening by a multiway graph—which now includes “backtracking” to lower fitnesses:

But although the details are different, in the end it doesn’t seem as if allowing this kind of backtracking has any dramatic effect. Somehow the basic phenomena around the process of adaptive evolution are strong enough that most of the details of how the adaptive evolution is done don’t ultimately matter much.

An Aside: Sexual Reproduction

In everything we’ve done so far, we’ve been making mutations only to individual rules. But there’s another mechanism that exists in many biological organisms: sexual reproduction, in which in effect a pair of rules (i.e. genomes) get mixed to produce a new rule. As a simple model of the crossover that typically happens with actual genomes, we can take two rules, and splice together the beginning of one with the end of the other:

In general there will be many ways to combine pairs of rules like this. In a direct analogy to our Physics Project, we can represent such “recombinations” as “events” that take two rules and produce one:

The analog of our multiway graph for all possible paths of adaptive evolution by mutations is now what we call in our Physics Project a token-event graph:

In dealing just with mutations we were able to take a single rule and progressively modify it. Now we always have to work with a “population” of rules, combining them two at a time to generate new rules. We can represent conceivable combinations among one set of rules as follows:

There are at this point many different choices we could make about how to set up our model. The particular approach we’ll use selects just n of the = n (n – 1)/2 possible combinations:

Then for each of these selected combinations we attempt a crossover, keeping those “children” (drawn here between their parents) that are not rejected as a result of having lower fitness:

Finally, to “maintain our gene pool”, we carry forward parents selected at random, so that we still end up with n rules. (And, yes, even though we’ve attempted to make this whole procedure as clean as possible, it’s still a mess—which seems to be inevitable, and which has, as we’ll discuss below, bedeviled computational studies of evolution in the past.)

OK, so what happens when we apply this procedure, say to k = 3, r = 1 rules? We’ll pick 4 rules at random as our initial population (and, yes, two happen to produce the same pattern):

Then in a sequence of steps we’ll successively pick various combinations:

And here are the distinct “phenotype patterns” produced in this process (note that even though there can be multiple copies of the same phenotype pattern, the underlying genotype rules are always distinct):

As a final form of summarization we can just plot the successive fitnesses of the patterns we generate (with the size of each dot reflecting the number of times a particular fitness occurs):

In this case we reach a steady state after 9 steps. The larger the population the longer the adaptive evolution will typically keep going. Here are a couple of examples with population 10, showing all the patterns obtained:

Showing in each case only the longest-lifetime rule found so far we get:

The results aren’t obviously different from what we were finding with mutation alone—even though now we’ve got a much more complicated model, with a whole population of rules rather than a single rule. (One obvious difference, though, is that here we can end up with overall cycles of populations of rules, whereas in the pure-mutation case that can only happen among fitness-neutral rules.)

Here are some additional examples—now obtained after 500 steps with population 25

and with population 50:

And so far as one can tell, even here there are no substantial differences from what we saw with mutation. Certainly there are detailed features introduced by sexual reproduction and crossover, but for our purposes in understanding the big picture of what’s happening in adaptive evolution it seems sufficient to do as we have done so far, and consider only mutation.

An Even More Minimal Model

By investigating adaptive evolution in cellular automata we’re already making dramatic simplifications relative, say, to actual biology. But in the effort to understand the essence of phenomena we see, it’s helpful to go even further—and instead of thinking about computational rules and their behavior, just think about vertices on a “mutation graph”, each assigned a certain fitness.

As an example, we can set up a 2D grid, assigning each point a certain random fitness:

And then, starting from a minimum-fitness point, we can follow the same kind of adaptive evolution procedure as above, at each step going to a neighboring point with an equal or greater fitness:

Typically we don’t manage to go far before we get stuck, though with the uniform distribution of fitness values used here, we still usually end on a fairly large fitness value.

We can summarize the possible paths we can take by the multiway graph:

In our cellular automaton rule space—and, for that matter, in biology—neighboring points don’t just have independent random fitnesses; instead, the fitnesses are determined by a definite computational procedure. So as a simple approximation, we can just take the fitness of each point to be a particular function of its graph coordinates. If the function forms something like a “uniform hill”, then the adaptive evolution procedure will just climb it:

But as soon as the function has “systematic bumpiness” there’s a tremendous tendency to quickly get stuck:

And if there’s some “unexpected spot of high fitness”, adaptive evolution typically won’t find it (and it certainly won’t if it’s surrounded by a lower-fitness “moat”):

So what happens if we increase the dimensionality of the “mutation space” in which we’re operating? Basically it becomes easier to find a path that increases fitness:

And we can see this, for example, if we look at Boolean hypercubes in increasing numbers of dimensions:

But ultimately this relies on the fact that in the neighborhood reachable by mutations from a given point, there’ll be a “sufficiently random” collection of fitness values that it’ll (likely) be possible to find a “direction” that’s “going up” in fitness. Yet this alone won’t in general be enough, because we also need it to be the case that there’s enough regularity in the fitness landscape that we can systematically navigate it to find its maximum—and that the maximum is not somehow “unexpected and isolated”.

What Can Adaptive Evolution Achieve?

We’ve seen that adaptive evolution can be surprisingly successful at finding cellular automata that produce patterns with long but finite lifetimes. But what about other types of “traits”? What can (and cannot) adaptive evolution ultimately manage to do?

For example, what if we’re trying to find cellular automata whose patterns don’t just live “as long as possible” but instead die after a specific number of steps? It’s clear that within any finite set of rules (say with particular k and r) there’ll only be a limited collection of possible lifetimes. For symmetric k = 2, r = 2 rules, for example, the possible lifetimes are:

But as soon as we’re dealing even with k = 3, r = 1 symmetric rules it’s already in principle possible to get every lifetime up to 100. But what about adaptive evolution? How well does it do at reaching rules with all those lifetimes? Let’s say we do single point mutation as before, but now we “accept” a mutation if it leads not specifically to a larger finite lifetime, but to a lifetime that is closer in absolute magnitude to some desired lifetime. (Strictly, and importantly, in both cases we also allow “fitness-neutral” mutations that leave the lifetime the same.)

Here are examples of what happens if we try to adaptively evolve to get lifetime exactly 50 in k = 3, r = 1 rules:

It gets close—and sometimes it overshoots—but, at least in these particular examples, it never quite makes it. Here’s what we see if we look at the lifetimes achieved with 100 different random sequences of mutations:

Basically they mostly get stuck at lifetimes close to 50, but not exactly 50. It’s not that k = 3, r = 1 rules can’t yield lifetime 50; exhaustive search shows that even many symmetric such rules can:

It’s just that our adaptive evolution process usually gets stuck before it reaches rules like these. Even though there’s usually enough “room to maneuver” in k = 3, r = 1 rule space to get to generally longer lifetimes, there’s not enough to specifically get to lifetime 50.

But what about k = 4, r = 1 rule space? There are now not 1012 but about 1038 possible rules. And in this rule space it becomes quite routine to be able to reach lifetime 50 through adaptive evolution:

It can sometimes take a while, but most of the time in this rule space it’s possible to get exactly to lifetime 50:

What happens with other “lifetime goals”? Even symmetric k = 3, r = 1 rules can achieve many lifetime values:

Indeed, the first “missing” values are 129, 132, 139, etc. And, for example, many multiples of 50 can be achieved:

But it becomes increasingly difficult for adaptive evolution to reach these specific goals. Increasing the size of the rule space always seems to help; so for example with k = 4, r = 1, if one’s aiming for lifetime 100, the actual distribution of lifetimes reached is:

In general the distribution gets broader as the lifetime sought gets larger:

We saw above that across the whole space of, say, k = 4, r = 1 rules, the frequency of progressively larger lifetimes falls roughly according to a power law. So this means that the fractional region in rule space that achieves a given lifetime gets progressively smaller—with the result that typically the paths followed by adaptive evolution are progressively more likely to get stuck before they reach it.

OK, so what about other kinds of objectives? Say ones more related to the morphologies of patterns? As a simple example, let’s consider the objective of maximizing the “widths” of finite-lifetime patterns. We can try to achieve this by adaptive evolution in which we reject any mutations that lead to decreased width (where “width” is defined as the maximum horizontal extent of the pattern). And once again this process manages to “discover” all sorts of “mechanisms” for achieving larger widths (here each pattern is labeled by its height—i.e. lifetime—and width):

There are certain structural constraints here. For example, the width can’t be too large relative to the height—because if it’s too large, patterns tend to grow forever.

But what if we specifically try to select for maximal “pattern aspect ratio” (i.e. ratio of width to height)? In essentially every case so far, adaptive evolution has in effect “invented many different mechanisms” to achieve whatever objective we’ve defined. But here it turns out we essentially see “the same idea” being used over and over again—presumably because this is the only way to achieve our objective given the overall structure of how the underlying rules we’re using work:

What if we ask something more specific? Like, say, that the aspect ratio be as close to 3 as possible. Much of the time the “solution” that adaptive evolution finds is the correct if trivial:

But sometimes it finds another solution—and often a surprisingly elaborate and complicated one:

How about if our goal is an aspect ratio of π ≈ 3.14? It turns out adaptive evolution can still do rather well here even just with the symmetric k = 3, r = 1 rules that we’re using:

We can also ask about properties of the “inside” of the pattern. For example, we can ask to maximize the lengths of uniform runs of nonwhite cells in the center column of the pattern. And, once again, adaptive evolution can successfully lead us to rules (like these random examples) where this is large:

We can go on and get still more detailed, say asking about runs of particular lengths, or the presence or number of particular subpatterns. And eventually—just like when we asked for too long a lifetime—we’ll find that the cases we’re looking for are “too sparse”, and adaptive evolution (at least in a given rule space) won’t be able to find them, even if exhaustive search could still identify at least a few examples.

But just what kinds of objectives (or fitness functions) can be handled how well by adaptive evolution, operating for example on the “raw material” of cellular automata? It’s an important question—an analog of which is also central to the investigation of machine learning. But as of now we don’t really have the tools to address it. It’s somehow reminiscent of asking what kinds of functions can be approximated how well by different methods or basis functions. But it’s more complicated. Solving it, though, would tell us a lot about the “reach” of adaptive evolution processes, not only for biology but also for machine learning.

What It Means for What’s Going On in Biology

How do biological organisms manage to be the way they are, with all their complex and seemingly clever solutions to such a wide range of challenges? Is it just natural selection that does it, or is there in effect more going on? And if “natural selection does it”, how does it actually manage it?

From the point of view of traditional engineering what we see in biology is often very surprising, and much more complex and “clever” than we’d imagine ever being able to create ourselves. But is the secret of biology in a sense just natural selection? Well, actually, there’s often an analog of natural selection going on even in engineering, as different designs get tried and only some get selected. But at least in traditional engineering a key feature is that one always tries to come up with designs where one can foresee their consequences.

But biology is different. Mutations to genomes just happen, without any notion that their consequences can be foreseen. But still one might assume that—when guided by natural selection—the results wouldn’t be too different to what we’d get in traditional engineering.

But there’s a crucial piece of intuition missing here. And it has to do with how randomly chosen programs behave. We might have assumed (based on our typical experience with programs we explicitly construct for particular purposes) that at least a simple random program wouldn’t ever do anything terribly interesting or complicated.

But the surprising discovery I made in the early 1980s is that this isn’t true. And instead, it’s a ubiquitous phenomenon that in the computational universe of possible programs, one can get immense complexity even from very simple programs. So this means that as mutation operates on a genome, it’s essentially inevitable that it’ll end up sampling programs that show highly complex behavior. At the outset, one might have imagined that such complexity could only be achieved by careful design and would inevitably be at best rare. But the surprising fact is that—because of how things fundamentally work in the computational universe—it’s instead easy to get.

But what does complexity have to do with creating “successful organisms”? To create a “successful organism” that can prosper in a particular environment there fundamentally has to be some way to get to a genome that will “solve the necessary problems”. And this is where natural selection comes in. But the fact that it can work is something that’s not at all obvious.

There are really two issues. The first is whether a program (i.e. genome) even exists that will “solve the necessary problems”. And the second is whether such a program can be found by a “thread” of adaptive evolution that goes only through intermediate states that are “fit enough” to survive. As it turns out, both these issues are related to the same fundamental features of computation—which are also responsible for the ubiquitous appearance of complexity.

Given some underlying framework—like cellular automata, or like the basic apparatus of life—is there some rule that can be implemented in that framework that will achieve some particular (computational) objective? The Principle of Computational Equivalence says that generically the answer will be yes. In effect, given almost any “underlying hardware”, it’ll ultimately be possible to come up with “software” (i.e. a rule) that achieves almost any (“physically possible”) given objective—like growing an organism of at least some kind that can survive in a particular environment. But how can we actually find a rule that achieves this?

In principle we could do exhaustive search. But that will be exponentially difficult—and in all but toy cases will be utterly infeasible in practice. So what about adaptive evolution? Well, that’s the big question. And what we’ve seen here is that—rather surprisingly—simple mutation and selection (i.e. the mechanisms of natural selection) very often provide a dramatic shortcut for finding rules that do what we want.

So why is this? In effect, adaptive evolution is finding a path through rule space that gets to where we want to go. But the surprising part is that it’s managing to do this one step at a time. It’s just trying random variations (i.e. mutations) and as soon as it finds one that’s not a “step down in fitness”, it’ll “take it”, and keep going. At the outset it’s certainly not obvious that this will work. In particular, it could be that at some point there just won’t be any “way forward”. All “directions” will lead only to lower fitness, and in effect the adaptive evolution will get stuck.

But the key observation from the experiments in our simple model here is that this typically doesn’t happen. And there seem to be basically two things going on. The first is that rule space is in effect very high-dimensional. So that means there are “many directions to choose from” in trying to find one that will allow one to “take a step forward”. But on its own this isn’t enough. Because there could be correlations between these directions that would mean that if one’s blocked in one direction one would inevitably be blocked in all others.

So why doesn’t this happen? Well, it seems to be the result of the fundamental computational phenomenon of computational irreducibility. A traditional view based on experience with mathematical science had been that if one knew the underlying rule for a system then this would immediately let one predict what the system would do. But what became clear from my explorations in the 1980s and 1990s is that in the computational universe this generically isn’t true. And instead, that the only way one can systematically find out what most computational systems will do is explicitly to run their rules, step by step, doing in effect the same irreducible amount of computational work that they do.

So if one’s just presented with behavior from the system one won’t be in a position to “decode it” and “see its simple origins”. Unless one’s capable of doing as much computational work as the system itself, one will just have to consider what it’s doing as (more or less) “random”. And indeed this seems to be at the root of many important phenomena, such as the Second Law of thermodynamics. And I also suspect it’s at the root of the effectiveness of adaptive evolution, notably in biology.

Because what computational irreducibility implies is that around every point in rule space there’ll be a certain “effective randomness” to fitnesses one sees. And if there are many dimensions to rule space that means it’s overwhelmingly likely that there’ll be “paths to success” in some directions from that point.

But will the adaptive evolution find them? We’ve assumed that there are a series of mutations to the rule, all happening “at random”. And the point is that if there are n elements in the rule, then after some fraction of n mutations we should find our “success direction”. (If we were doing exhaustive search, we’d instead have to try about kn possible rules.)

At the outset it might seem conceivable that the sequence of mutations could somehow “cleverly probe” the structure of rule space, “knowing” what directions would or would not be successful. But the whole point is that going from a rule (i.e. genotype) to its behavior (i.e. phenotype) is generically a computationally irreducible process. So assuming that mutations are generated in a computationally bounded way it’s inevitable that they can’t “break computational irreducibility” and so will “experience” the fitness landscape in rule space as “effectively random”.

OK, but what about “achieving the characteristics an organism needs”? What seems to be critical is that these characteristics are in a sense computationally simple. We want an organism to live long enough, or be tall enough, or whatever. It’s not that we need the organism to perform some specific computationally irreducible task. Yes, there are all sorts of computationally irreducible processes happening in the actual development and behavior of an organism. But as far as biological evolution is concerned all that matters is ultimately some computationally simple measure of fitness. It’s as if biological evolution is—in the sense of my recent observer theory—a computationally bounded observer of underlying computationally irreducible processes.

And to the observer what emerges is the “simple law” of biological evolution, and the idea that, yes, it is possible just by natural selection to successfully generate all sorts of characteristics.

There are all sorts of consequences of this for thinking about biology. For example, in thinking about where complexity in biology “comes from”. Is it “generated by natural selection”, perhaps reflecting the complicated sequence of historical accidents embodied in the particular collection of mutations that occurred? Or is it from somewhere else?

In the picture we’ve developed here it’s basically from somewhere else—because it’s essentially a reflection of computational irreducibility. Having said that, we should remember that the very possibility of being able to have organisms with such a wide range of different forms and functions is a consequence of the universal computational character of their underlying setup, which in turn is closely tied to computational irreducibility.

And it’s in effect because natural selection is so coarse in its operation that it does not somehow avoid the ubiquitous computational irreducibility that exists in rule space, with the result that when we “look inside” biological systems we tend to see computational irreducibility and the complexity associated with it.

Something that we’ve seen over and over again here is that, yes, adaptive evolution manages to “solve a problem”. But its solution looks very complex to us. There might be some “simple engineering solution”—involving, say, a very regular pattern of behavior. But that’s not what adaptive evolution finds; instead it finds something that to us is typically very surprising—very often an “unexpectedly clever” solution in which lots of pieces fit together just right, in a way that our usual “understand-what’s-going-on” engineering practices would never let us invent.

We might not have expected anything like this to emerge from the simple process of adaptive evolution. But—as the models we’ve studied here highlight—it seems to be an inevitable formal consequence of core features of computational systems. And as soon as we recognize that biological systems can be viewed as computational, then it also becomes something inevitable for them—and something we can view as in a sense formally derivable for them.

At the outset we might not have been able to say “what matters” in the emergence of complexity in biology. But from the models we’ve studied, and the arguments we’ve made, we seem to have quite firmly established that it’s a fundamentally computational phenomenon, that relies only on certain general computational features of biological systems, and doesn’t depend on their particular detailed components and structure.

But in the end, how “generic” is the complexity that comes out of adaptive evolution? In other words, if we were to pick programs, say completely at random, how different would the complexity they produce be from the complexity we see in programs that have been adaptively evolved “for a purpose”? The answer isn’t clear—though to know it would provide important foundational input for theoretical biology.

One has the general impression that computational irreducibility is a strong enough phenomenon that it’s the “dominant force” that determines behavior and produces complexity. But there’s still usually something a bit different about the patterns we see from rules we’ve found by adaptive evolution, compared to rules we pick at random. Often there seems to be a certain additional level of “apparent mechanism”. The details still look complicated and in some ways quite random, but there seems to be a kind of “overall orchestration” to what’s going on.

And whenever we can identify such regularities it’s a sign of some kind of computational reducibility. There’s still plenty of computational irreducibility at work. But “high fitness” rules that we find through adaptive evolution typically seem to exhibit traces of their specialness—that manifests in at least a certain amount of computational reducibility.

Whenever we manage to come up with a “narrative explanation” or a “natural law” for something, it’s a sign that we’ve found a pocket of computational reducibility. If we say that a cellular automaton manages to live long because it generates certain robust geometric patterns—or, for that matter, that an organism lives long because it proofreads its DNA—we’re giving a narrative that’s based on computational reducibility.

And indeed whenever we can successfully identify a “mechanism” in our cellular automaton behavior, we’re in effect seeing computational reducibility. But what can we say about the aggregate of a whole collection of mechanisms?

In a different context I’ve discussed the concept of a “mechanoidal phase”, distinguished, say, from solids and liquids by the presence of a “bulk orchestration” of underlying components. It’s something closely related to class 4 behavior. And it’s interesting to note that if we look, for example, at the rules we found from adaptive evolution at the end of the previous section, their evolution from random initial conditions mostly shows characteristic class 4 behavior:

In other words, adaptive evolution is potentially bringing us to “characteristically special” places in rule space—perhaps suggesting that there’s something “characteristically special” about the kind of structures that are produced in biological systems. And if we could find a way to make general statements about that “characteristic specialness” it would potentially lead us to a framework for constructing a new broad formal theory in biology.

Correspondence with Biological Phenomena

The models we’ve studied here are extremely simple in their basic construction. And at some level it’s remarkable that—without for example including any biophysics or biochemistry—they can get anywhere at all in capturing features of biological systems and biological evolution.

In a sense this is ultimately a reflection of the fundamentally computational character of biology—and the generality of computational phenomena. But it’s very striking that even the patterns of cellular automaton behavior we see look very “lifelike and organic”.

In actual biology even the shortest genomes are vastly longer than the tiny cellular automaton rules we’ve considered. But even by the time we’re looking at the length-27 “genomic sequences” in k = 3, r = 1 cellular automata, there are already 3 trillion possible sequences, which seems to be enough to see many core “combinatorially driven” biology-like phenomena.

The running of a cellular automaton rule might also at first seem far away from the actual processes that create biological organisms—involving as they do things like the construction of proteins and the formation of elaborate functional and spatial structures. But there are more analogies than one might at first imagine. For example, it’s common for only particular cases in the cellular automaton rule to be used in a given region of the pattern that’s formed, much as particular genes are typically turned on in different tissues in biological organisms.

And, for example, the “geometrical restriction” to a simple 1D array of “cells” doesn’t seem to matter much as soon as there’s sophisticated computation going on; we still get lots of structures that are actually surprisingly reminiscent of typical patterns of biological growth.

One of the defining features of biological organisms is their capability for self-reproduction. And indeed if it wasn’t for this kind of “copying” there wouldn’t be anything like adaptive evolution to discuss. Our models don’t attempt to derive self-reproduction; they just introduce it as something built into the models.

And although we’ve considered several variants, we’re basically also just building into our models the idea of mutations. And what we find is that it seems as if single point mutations made one at a time are enough to capture basic features of adaptive evolution.

We’ve also primarily considered what amounts to a single lineage—in which there’s just a single rule (or genome) at a given step. We do mutations, and we mostly “implement natural selection” just by keeping only rules that lead to patterns whose fitness is no less than what we had before.

If we had a whole population of rules it probably wouldn’t be so significant, but in the simple setup we’re using, it turns out to be important that we don’t reject “fitness-neutral mutations”. And indeed we’ve seen many examples where the system wanders around fitness-neutral regions of rule space before finally “discovering” some “innovation” that allows it to increase fitness. The way our models are set up, that “wandering” always involves changes in the “genotype”—but usually at most minor changes in the phenotype. So it’s very typical to see long periods of “apparent equilibrium” in which the phenotype changes rather little, followed by a “jump” to a new fitness level and a rather different phenotype.

And this observation seems quite aligned with the phenomenon of “punctuated equilibrium” often reported in the fossil record of actual biological evolution.

Another key feature of biological organisms and biological evolution is the formation of distinct species, as well as distinct phyla, etc. And indeed we ubiquitously see something that seems to be directly analogous in our multiway graphs of all possible paths of adaptive evolution. Typically we see distinct branches forming, based on what seem like “different mechanisms” for achieving fitness.

No doubt in actual biology there are all sorts of detailed phenomena related to reproductive or spatial isolation. But in our models the core phenomenon that seems to lead to the analog of “branching in the tree of life” is the existence of “distinctly different computational mechanisms” in different parts of rule space. It’s worth noting that at least with our finite rule spaces, branches can die out, with no “successor species” appearing the multiway graph.

And indeed looking at the actual patterns produced by rules in different parts of the multiway graph it’s easy to imagine morphologically based taxonomic classifications—that would be somewhat, though not perfectly, aligned with the phylogenetic tree defined by actual rule mutations. (At a morphological level we quite often see some level of “convergent evolution” in our multiway graphs; in small examples we sometimes also see actual “genomic convergence”—which will typically be astronomically rare in actual biological systems.)

One of the remarkable features of our models is that they allow quite global investigation of the “overall history” of adaptive evolution. In many of the simple cases we’ve discussed, the rule space we’re using is small enough that in a comparatively modest number of mutation steps we get to the “highest fitness we can reach”. But (as the examples we saw in Turing machines suggest) expanding the size of the rules we’re using even just a little can be expected to be sufficient to allow us to get astronomically further.

And the further we go, the more “mechanisms” will be “invented”. It’s an inevitable feature of systems that involve computational irreducibility that there are new and unpredictable things that will go on showing up forever—along with new pockets of computational reducibility. So even after a few billion years—and the trillion generations and 1040 or so organisms that have ever lived—there’s still infinitely further for biological evolution to go, and more and more branches to be initiated in the tree of life, involving more and more “new mechanisms”.

I suppose one might imagine that at some point biological organisms would reach a “maximum fitness”, and go no further. But even in our simple model with fitness measured in terms of pattern lifetime, there’ll be no upper limit on fitness; given any particular lifetime, it’s a feature of the fundamental theory of computation that there’ll always be a program that will yield a larger lifetime. Still, one might think, at some point enough is enough: the giraffe’s neck is long enough, etc. But if nothing else, competition between organisms will always drive things forward: yes, a particular lineage of organisms achieved a certain fitness, but then another lineage can come along and get to that fitness too, forcing the first lineage to go even further so as to not lose out.

Of course, in our simple model we’re not explicitly accounting for interactions with other organisms—or for detailed properties of the environment, as well as countless other effects. And no doubt there are many biological phenomena that depend on these effects. But the key point here is that even without explicitly accounting for any of these effects, our simple model still seems to capture many core features of biological evolution. Biological evolution—and, indeed, adaptive evolution in general—is, it seems, fundamentally a computational phenomenon that robustly emerges quite independent of the details of systems.

In the past few years our Physics Project has given strong evidence that the foundations of physics are fundamentally computational—with the core laws of physics arising as inevitable consequences of the way observers like us “parse” the ruliad of all possible computational processes. And what we’ve seen here now suggests that there’s a remarkable commonality between the foundations of physics and biology. Both are anchored in computational irreducibility. And both sample slices of computational reducibility. Physics because that’s what observers like us do to get descriptions of the world that fit in our finite minds. Biology because that’s what biological evolution does in order to achieve the “coarse objectives” set by natural selection.

The intuition of physics tends to be that there are ultimately simple models for things, whereas in biology there’s a certain sense that everything is always almost infinitely complicated, with a new effect to consider at every turn. But presumably that’s in large part because what we study in biology tends to quickly come face to face with computational irreducibility—whereas in physics we’ve been able to find things to study that avoid this. But now the commonality in foundations between physics and biology suggests that there should also be in biology the kind of structure we have in physics—complete with general laws that allow us to make useful, broad statements. And perhaps the simple model I’ve presented here can help lead us there—and in the end help build up a new paradigm for thinking about biology in a fundamentally theoretical way.

Historical Notes

There’s a long—if circuitous—history to the things I’m discussing here. Basic notions of heredity—particularly for humans—were already widely recognized in antiquity. Plant breeding was practiced from the earliest days of agriculture, but it wasn’t until the late 1700s that any kind of systematic selective breeding of animals began to be commonplace. Then in 1859 Charles Darwin described the idea of “natural selection” whereby competition of organisms in their natural environment could act like artificial selection, and, he posited, would over long periods lead to the development of new species. He ended his Origin of Species with the claim that:

… from the war of nature … the production of the higher animals, directly follows. … and whilst this planet has gone cycling on according to the fixed law of gravity, from so simple a beginning endless forms most beautiful and most wonderful have been, and are being, evolved.

What he appears to have thought is that there would somehow follow from natural selection a general law—like the law of gravity—that would lead to the evolution of progressively more complex organisms, culminating in the “higher animals”. But absent the kind of model I’m discussing here, nothing in the later development of traditional evolutionary theory really successfully supported this—or was able to give much analysis of it.

Right around the same time as Darwin’s Origin of Species, Gregor Mendel began to identify simple probabilistic laws of inheritance—and when his work was rediscovered at the beginning of the 1900s it was used to develop mathematical models of the frequencies of genetic traits in populations of organisms, with key contributions to what became the field of population genetics being made in the 1920s and 1930s by J. B. S. Haldane, R. A. Fisher and Sewall Wright, who came up with the concept of a “fitness landscape”.

On a quite separate track there had been efforts ever since antiquity to classify and understand the growth and form of biological organisms, sometimes by analogy to physical or mathematical ideas—and by the 1930s it seemed fairly clear that chemical messengers were somehow involved in the control of growth processes. But the mathematical methods used for example in population genetics basically only handled discrete traits (or simple numerical ones accessible to biometry), and didn’t really have anything to say about something like the development of complexity in the forms of biological organisms.

The 1940s saw the introduction of what amounted to electrical-engineering-inspired approaches to biology, often under the banner of cybernetics. Idealized neural networks were introduced by Warren McCulloch and Walter Pitts in 1943, and soon the idea emerged (notably in the work of Donald Hebb in 1949) that learning in such systems could occur through a kind of adaptive evolution process. And by the time practical electronic computing began to emerge in the 1950s there was widespread belief that ideas from biology—including evolution—would be useful as an inspiration for what could be done. Often what would now just be described as adaptive algorithms were couched in biological evolution terms. And even when iterative methods were used for optimization (say in industrial production or engineering design) they were sometimes presented as being grounded in biological evolution.

Meanwhile, by the 1960s, there began to be what amounted to Monte Carlo simulations of population-genetics-style evolutionary processes. A particularly elaborate example was work by Nils Barricelli on what he called “numeric evolution” in which a fairly complicated numerical-cellular-automaton-like “competition between organisms” program with “randomness” injected from details of data layout in computer memory showed what he claimed to be biological-evolution-like phenomena (such as symbiosis and parasitism).

In a different direction there was an attempt—notably by John von Neumann—to “mathematicize the foundations of biology” leading by the late 1950s to what we’d now call 2D cellular automata “engineered” in complicated ways to show phenomena like self-reproduction. The followup to this was mostly early-theoretical-computer-science work, with no particular connection to biology, and no serious mention of adaptive evolution. When the Game of Life was introduced in 1970 it was widely noted as “doing lifelike things”, but essentially no scientific work was done in this direction. By the 1970s, though, L-systems and fractals had introduced the idea of recursive tree-like or nested structures that could be generated by simple algorithms and rendered by computer graphics—and seemed to give forms close to some seen in biology. My own work on 1D cellular automata (starting in 1981) focused on systematic scientific investigation of simple programs and what they do—with the surprising conclusion that even very simple programs can produce highly complex behavior. But while I saw this as informing the generation of complexity in things like the growth of biological organisms, I didn’t at the time (as I’ll describe below) end up seriously exploring any adaptive evolution angles.

Still another thread of development concerned applying biological-like evolution not just to parameters but to operations in programs. And for example in 1958 Richard Friedberg at IBM tried making random changes to instructions in machine-code programs, but didn’t manage to get this to do much. (Twenty years later, superoptimizers in practical compilers did begin to successfully use such techniques.) Then in the 1960s, John Holland (who had at first studied learning in neural nets, and was then influenced by Arthur Burks who had worked on cellular automata with von Neumann) suggested representing what amounted to programs by simple strings of symbols that could readily be modified like genomic sequences. The typical idea was to interpret the symbols as computational operations—and to assign a “fitness” based on the outcome of those operations. A “genetic algorithm” could then be set up by having a population of strings that was adaptively evolved. Through the 1970s and 1980s occasional practical successes were reported with this approach, particularly in optimization and data classification—with much being made of the importance of sexual-reproduction-inspired crossover operations. (Something that began to be used in the 1980s was the much simpler approach of simulated annealing—that involves randomly changing values rather than programs.)

By the beginning of the 1980s the idea had also emerged of adaptively modifying the structure of mathematical expressions—and of symbolic expressions representing programs. There were notable applications in computer graphics (e.g. by Karl Sims) as well as to things like the 1984 Core War “game” involving competition between programs in a virtual machine. In the 1990s John Koza was instrumental in developing the idea of “genetic programming”, notably as a way to “automatically create inventions”, for example in areas like circuit and antenna design. And indeed to this day scattered applications of these methods continue to pop up, particularly in geometrical and mechanical design.

From the very beginning there’d been controversy around Darwin’s ideas about evolution. First, there was the issue of conflict with religious accounts of creation. But there were also—often vigorous—disagreements within the scientific community about the interpretation of the fossil record and about how large-scale evolution was really supposed to operate. A notable issue—still very active in the 1980s—was the relationship between the “freedom of evolution” and the constraints imposed by the actual dynamics of growth in organisms (and interactions between organisms). And despite much insistence that the only reasonable “scientific” (as opposed to religious) point of view was that “natural selection is all there is”, there were nagging mysteries that suggested there must be other forces at work.

Building on the possibilities of computer experimentation (as well as things like my work on cellular automata) there emerged in the mid-1980s, particularly through the efforts of Chris Langton, a focus on investigating computational models of “artificial life”. This resulted in all sorts of simulations of ecosystems, etc. that did produce a variety of evolution-related phenomena known from field biology—but typically the models were far too complex in their structure for it to be possible to extract fundamental conclusions from them. Still, there continued to be specific, simpler experiments. For example, in 1986, for his book The Blind Watchmaker, Richard Dawkins made pictures of what he called “biomorphs”, produced by adaptively adjusting parameters for a simple tree-growth algorithm based on the overall shapes generated.

In the 1980s, stimulated by my work, there were various isolated studies of “rule evolution” in cellular automata (as well as art and museum exhibits based on this), and in the 1990s there was more systematic work—notably by Jim Crutchfield and Melanie Mitchell—on using genetic algorithms to try to evolve cellular automaton rules to solve tasks like density classification. (Around this time “evolutionary computation” also began to emerge as a general term covering genetic algorithms and other usually-biology-inspired adaptive computational methods.)

Meanwhile, accelerating in the 1990s, there was great progress in understanding actual molecular mechanisms in biology, and in figuring out how genetic and developmental processes work. But even as huge amounts of data accumulated, enthusiasm for large-scale “theories of biology” (that might for example address the production of complexity in biological evolution) seemed to wane. (The discipline of systems biology attempted to develop specific, usually mathematical, models for biological systems—but there never emerged much in the way of overarching theoretical principles, except perhaps, somewhat specifically, in areas like immunology and neuroscience.)

One significant exception in terms of fundamental theory was Greg Chaitin’s concept from around 2010 of “metabiology”: an effort (see below) to use ideas from the theory of computation to understand very general features of the evolution of programs and relate them to biological evolution.

Starting in the 1950s another strand of development (sometimes viewed as a practical branch of artificial intelligence) involved the idea of “machine learning”. Genetic algorithms were one of half a dozen common approaches. Another was based on artificial neural nets. For decades machine learning languished as a somewhat esoteric field, dominated by engineering solutions that would occasionally deliver specific application results. But then in 2011 there was unexpectedly dramatic success in using neural nets for image identification, followed in subsequent years by successes in other areas, and culminating in 2022 with the arrival of large language models and ChatGPT.

What hadn’t been anticipated was that the behavior of neural nets can change a lot if they’re given sufficiently huge amounts of training. But there isn’t any good understanding of just why this is so, and just how successful neural nets can be at what kinds of tasks. Ever since the 1940s it has been recognized that there are relations between biological evolution and learning in neural nets. And having now seen the impressive things neural nets can do, it seems worthwhile to look again at what happens in biological evolution—and to try to understand why it works, not least as a prelude to understanding more about neural nets and machine learning.

Personal Notes

It’s strange to say, but most of what I’ve done here I should really have done forty years ago. And I almost did. Except that I didn’t try quite the right experiments. And I didn’t have the intuition to think that it was worth trying more.

Forty years later, I have new intuition, particularly informed by experience with modern machine learning. But even now, what made possible what I’ve done here was a chance experiment done for a somewhat different purpose.

Back in 1981 I had become very interested in the question of how complexity arises in the natural world, and I was trying to come up with models that might capture this. Meanwhile, I had just finished Version 1.0 of SMP, the forerunner to Mathematica and the Wolfram Language—and I was wondering how one might generalize its pattern-matching paradigm to “general AI”.

As it happened, right around that time, neural nets gained some (temporary) popularity. And seeing them as potentially relevant to both my topics I started simulating them and trying to see what kind of general theory I could develop about them. But I found them frustrating to work with. There seemed to be too many parameters and details to get any clear conclusions. And, at a practical level, I couldn’t get them to do anything particularly useful.

I decided that for my science question I needed to come up with something much simpler. And as a kind of minimal merger of spin systems and neural nets I ended up inventing cellular automata (only later did I discover that versions of them had been invented several times before).

As soon as I started doing experiments on them, I discovered that cellular automata were a window into an amazing new scientific world—that I have continued to explore in one way or another ever since. My key methodology, at least at first, was just to enumerate the simplest possible cellular automaton rules, and see what they did. The diversity—and complexity—of their behavior was remarkable. But the simplicity of the rules meant that the details of “successive rules” were usually fairly different—and while there were common themes in their overall behavior, there didn’t seem to be any particular structure to “rule space”. (Occasionally, though, particularly in finding examples for exposition, I would look at slightly more complicated and “multicolored” rules, and I certainly anecdotally noticed that rules with nearby rule numbers often had definite similarities in their behavior.)

It so happened that around the time I started publishing about cellular automata in 1983 there was a fair amount of ambient interest in theoretical biology. And (perhaps in part because of the “cellular” in “cellular automata”) I was often invited to theoretical biology conferences. People would sometimes ask about adaptation in cellular automata, and I would usually just emphasize what individual cellular automata could do, without any adaptation, and what significance it might have for the development of organisms.

But in 1985 I was going to a conference (at Los Alamos) on “Evolution, Games and Learning” and I decided I should take a look at the relation of these topics to cellular automata. But, too quickly, I segued away from investigating adaptation to trying to see what kind of pattern matching and other operations cellular automata might be able to be explicitly set up to do:

Click to enlarge

Many aspects of this paper still seem quite modern (and in fact should probably be investigated more now!). But—even though I absolutely had the tools to do it—I simply failed at that time to explore what I’ve now explored here.

Back in 1984 Problem 7 in my “Twenty Problems in the Theory of Cellular Automata” was “How is different behavior distributed in the space of cellular automaton rules?” And over the years I’d occasionally think about “cellular automaton rule space”, wondering, for example, what kind of geometry it might have, particularly in the continuum limit of infinitely large rules.

By the latter half of the 1980s “theoretical biology” conferences had segued to “artificial life” ones. And when I went to such conferences I was often frustrated. People would show me simulations that seemed to have far too many parameters to ever be able to conclude much. People would also often claim that natural selection was a “very simple theory”, but as soon as it was “implemented” there’d be all kinds of issues—and choices to be made—about population sizes, fitness cutoffs, interactions between organisms, and so on. And the end result was usually a descent into some kind of very specific simulation without obvious robust implications.

(In the mid-1980s I put a fair amount of effort into developing both the content and the organization of a new direction in science that I called “complex systems research”. My emphasis was on systems—like cellular automata—that had definite simple rules but highly complex behavior. Gradually, though, “complexity” started to be a popular general buzzword, and—I suspect partly to distinguish themselves from my efforts—some people started emphasizing that they weren’t just studying complex systems, they were studying complex adaptive systems. But all too often this seemed mostly to provide an excuse to dilute the clarity of what could be studied—and I was sufficiently put off that I paid very little attention.)

By the mid-1990s, I was in the middle of writing A New Kind of Science, and I wanted to use biology as an example application of my methodology and discoveries in the computational universe. In a section entitled “Fundamental Issues in Biology” I argued (as I have here) that computational irreducibility is a fundamentally stronger force than natural selection, and that when we see complexity in biology it’s most likely of “computational origin” rather than being “sculpted” by natural selection. And as part of that discussion, I included a picture of the “often-somewhat-gradual changes” in behavior that one sees with successive 1-bit changes in a k = 3, r = 1 cellular automaton rule (yes, the book was not in color):

Click to enlarge

This wasn’t done adaptively; it was basically just looking along a “random straight line” in rule space. And indeed both here and in most of the book, I was concerned with what systems like cellular automata “naturally do”, not what they can be constructed (or adaptively evolved) to do. I did give “constructions” of how cellular automata can perform particular computational tasks (like generating primes), and, somewhat obscurely, in a section on “Intelligence in the Universe” I explored finding k = 3, r = 1 rules that can successfully “double their input” (my reason for discussing these rules was to highlight the difficulty of saying whether one of these cellular automata was “constructed for a purpose” or was just “doing what it does”):

Click to enlarge

Many years went by. There’d be an occasional project at our Summer School about rule space, and occasionally about adaptation. I maintained an interest in foundational questions in biology, gradually collecting information and sometimes giving talks about the subject. Meanwhile—though I didn’t particularly internalize the connection then—by the mid-2010s, through our practical work on it in the Wolfram Language, I’d gotten quite up to speed with modern machine learning. Around the same time I also heard from my friend Greg Chaitin about his efforts (as he put it) to “prove Darwin” using the kind of computational ideas he’d applied in thinking about the foundations of mathematics.

Then in 2020 came our Physics Project, with its whole formalism around things like multiway graphs. It didn’t take long to realize that, yes, what I was calling “multicomputation” wasn’t just relevant for fundamental physics; it was something quite general that could be applied in many areas, which by 2021 I was trying to catalog:

Areas of multicomputation

I did some thinking about each of these. The one I tackled most seriously first was metamathematics, about which I finished a book in 2022. Late that year I was finishing a (50-year-in-gestation) project—informed by our Physics Project—on understanding the Second Law of thermodynamics, and as part of this I made what I thought was some progress on thinking about the fundamental character of biological systems (though not their adaptive evolution).

And then ChatGPT arrived. And in addition to being involved with it technologically, I started to think about the science of it, and particularly about how it could work. Part of it seemed to have to do with unrecognized regularities in human language, but part of it was a reflection of the emerging “meta discovery” that somehow if you “bashed” a machine learning system hard enough, it seemed like it could manage to learn almost anything.

But why did this work? At first I thought it must just be an “obvious” consequence of high dimensionality. But I soon realized there was more to it. And as part of trying to understand the boundaries of what’s possible I ended up a couple of months ago writing a piece exploring “Can AI Solve Science?”:

Can AI Solve Science?

I talked about different potential objectives for science (making predictions, generating narrative explanations, etc.) And deep inside the piece I had a section entitled “Exploring Spaces of Systems” in which I talked about science problems of the form “Can one find a system that does X?”—and asked whether systems like neural nets could somehow let one “jump ahead” in what would otherwise be huge exhaustive searches. As a sideshow to this I thought it might be interesting to compare with what a non-neural-net adaptive evolution process could do.

Remembering Greg Chaitin’s ideas about connecting the halting problem to biological evolution I wondered if perhaps one could just adaptively evolve cellular automaton rules to find ones that generated a pattern with a particular finite lifetime. I imagined it as a classic machine learning problem, with a “loss function” one needed to minimize.

And so it was that just after 1 am on February 22 I wrote three lines of Wolfram Language code—and tried the experiment:

Click to enlarge

And it worked! I managed to find cellular automaton rules that would generate patterns living exactly 50 steps:

Click to enlarge

In retrospect, I was slightly lucky. First, that this ended up being such a simple experiment to try (at least in the Wolfram Language) that I did it even though I didn’t really expect it to work. And second, that for my very first experiment I picked parameters that happened to immediately work (k = 4, lifetime 50, etc.).

But, yes, I could in principle have done the same experiment 40 years ago, though without the Wolfram Language it wouldn’t have been so easy. Still, the computers I had back then were powerful enough that I could in principle have generated the same results then as now. But without my modern experience of machine learning I don’t think I would have tried—and I would certainly have given up too easily. And, yes, it’s a little humbling to realize that I’ve gone so many years assuming adaptive evolution was out of the reach of simple, clean experiments. But it’s satisfying now to be able to check off another mystery I’ve long wondered about. And to think that much more about the foundations of biology—and machine learning—might finally be within reach.

Thanks

Thanks to Brad Klee, Nik Murzin and Richard Assar for their help.

The specific results and ideas I’ve presented here are mostly very recent, but they build on background conversations I’ve had—some recently, some more than 40 years ago—with many people, including: Sydney Brenner, Greg Chaitin, Richard Dawkins, David Goldberg, Nigel Goldenfeld, Jack Good, Jonathan Gorard, Stephen J. Gould, Hyman Hartman, John Holland, Christian Jacob, Stuart Kauffman, Mark Kotanchek, John Koza, Chris Langton, Katja Della Libera, Aristid Lindenmayer, Pattie Maes, Bill Mydlowec, John Novembre, Pedro de Oliveira, George Oster, Norman Packard, Alan Perelson, Thomas Ray, Philip Rosedale, Robert Rosen, Terry Sejnowski, Brian Silverman, Karl Sims, John Maynard Smith, Catherine Wolfram, Christopher Wolfram and Elizabeth Wolfram.

The People’s AI

Par : Doc Searls
28 mai 2024 à 17:01
Prompt: “A vast field on which the ground spells the letters A and I, with people on it, having a good time.” Via Copilot | Designer

People need their own AIs. Personally and collectively.

We won’t get them from Anthropic, Apple, Google, OpenAI, Meta, or Microsoft. Not even from Apple.

All those companies will want to provide AIaaS: AI as a Service, rather than AI that’s yours alone. Or ours, collectively.

The People’s AI can only come from people. Since it will be made of code, it will come from open-source developers working for all of us, and not just for their employers—even if those employers are companies listed above.*

That’s how we got Linux, Apache, MySQL, Python, and countless other open-source code bases on which the digital world is now built from the ground up. Our common ground is open-source code, standards, and protocols.

The sum of business that happens atop that common ground is incalculably vast. It also owes to what we first started calling because effects twenty years ago at Bloggercon. That was when people were making a lot more money because of blogging than with blogging.

Right after that it also became clear that most of the money being made in the whole tech world was because of open-source code, standards, and protocols, rather than with them. (I wrote more about it here, here, and here.)

So, thanks to because effects, the most leveraged investments anyone can make today will be in developing open source code for The People’s AI.

That’s the AI each of us will have for our own, and that we can use both by ourselves and together as communities.

Those because investments will pay off on the with side as lavishly as investments in TCP/IP, HTTP, Linux, and countless other open-source efforts have delivered across the last three decades.

Only now they’ll pay off a lot faster. For all of us.


*See what I wrote for Linux Journal in 2006 about how IBM got clueful about paying kernel developers to work for the whole world and not just one company.

Intelligence artificielle et réalité augmentée, un couple fait pour durer ! - Grégory MAUBON

Encore un article qui va parler d’IA me direz-vous ! En effet, mais vous devez constater avec moi qu’au-delà du buzzword, l’intelligence artificielle bouscule de plus en plus de domaines depuis quelques années. Les technologies immersives, et en particulier la RA ne sont pas, ne sont pas épargnées et j’en avais déjà parlé à propos des lunettes de réalité augmentée.
Permalien

The digital revolution has failed

Par : Paris Marx
8 mars 2024 à 16:22
The digital revolution has failed

The internet emerged from the defense research establishment, but as it was breaking out of those constraints in the 1990s to be unleashed onto the wider world, it had to be given a new story — and libertarian capitalists wrote it. Picking up on the melding of libertarianism, technological optimism, and neoliberal economics over the previous couple decades, they deployed a narrative that changed the way we thought about the emerging network and laid the groundwork for the commercial opportunity that followed its privatization in 1995.

The following year, Electronic Frontier Foundation cofounder John Perry Barlow released a manifesto from the World Economic Forum in Davos — a fact that should’ve immediately set off alarm bells — that became a defining narrative for the era. Cyberspace was to be a virgin frontier, divorced from the realities of the material world it depended on, where people would interact with each other as equals, free from the burdens of race, sex, or wealth, and it was those users who would construct it free from the dictates of the big, bad government. The new virtual world was to be “an act of nature and it grows itself through our collective actions,” he wrote.

Those utopian libertarian visions may have been relatively easy to believe in the web’s more anarchic years, where even though companies were having cash shoveled at them as investors and founders sought to cash in on that new frontier, it was still easy for people to spin up their own websites and stake a claim, so to speak, beyond the growing digital towns of the nascent tech capitalists. But as the boom went bust with the turn of the new millennium and the companies that remained sought to solidify their positions, the enclosure of the digital commons became a far greater priority.

While Barlow had plenty of scorn to heap on governments, he didn’t have the same disdain for corporations that saw the cyberspace he proclaimed as a libertarian paradise to be a great means to make a lot of money — something that should’ve been clear from the site where he made his so-called declaration of independence. As the investment flooded in, lawmakers prioritized the economic value (and geopolitical power) that could be wrung from the internet, while new media like Wired Magazine sprang up to advocate for the new industry. The public was kept at the center of the narrative, but in reality the wider benefits became a lesser concern as long as the money kept flowing.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

Early benefits of the internet age

While greater commercialization brought more drawbacks to virtual life, there can be no denying that there were clear benefits to be had from the global connection the internet enabled. Suddenly the information of the world was at users’ fingertips, and communities on virtually any subject could easily be sought out as the network of websites, chat rooms, and forums kept expanding. While there was some technical skill needed for certain forms of engagement, it wasn’t so hard to customize your own web page to represent your identity to fellow netizens and participate alongside everyone else.

Even as the enclosure into platforms sped up in the 2000s (and certainly annoyed the more tech-savvy early adopters), it wasn’t entirely bad: the simplicity that came with platformization made it much easier for billions more people to get online and start finding reasons to keep coming back, especially since the commercialization was still at an early stage and had done little to compromise the user experience.

For a time, everything seemed to be heading in the right direction: internet access expanded and speeds increased, our access to it moved from the desk into our pockets, whatever information we sought would be served up in milliseconds, high-quality entertainment was easily accessible without advertising at an affordable price — if it cost anything at all — and you never had to lose touch with anyone. When things seemed to be going well, it was easier to overlook the occasional stories that should’ve been a warning for what was to come.

The erosion of the web’s promise

We now look back on those times as the good old days, before the ambitions of tech companies vastly expanded and the pressure for profit accelerated the degradation of what they’d built. Our relationship with the tech industry has changed and the broad consensus on its impact on the world has been souring for years as the benefits of the digital revolution slowed while the drawbacks began to escalate through the 2010s as greed became too strong a force to contain.

These days the finger is often pointed to Facebook as the standard bearer of the turn against tech, but Google seems far more illustrative. It began as a university research project where creators Sergey Brin and Larry Page openly acknowledged how advertising would compromise the search engine’s quality. But once they spun it out as a private company and later listed on public markets, Google began its slow slide toward what it is today: embracing advertising paired with a “Don’t Be Evil” slogan that would eventually be abandoned as the pressure to keep growing the ad profits increased. Now when you turn to it for the information it claims to deliver, you’re more likely to get a bunch of listicles optimized for its search algorithms that are stuffed with affiliate links and terrible ads, if not just a bunch of AI-generated trash.

And Google’s not the only one. Facebook has never been a darling, but any commitment to promoting positive engagement among users went out the window ages ago with the need to juice ad profits, even if that meant spreading right-wing extremism and false information, abetting genocide, and even blocking the spread of real news if it meant having to share a pittance of its massive profits with struggling publishers. The greed of those two companies has sent news media spiraling, with lower ad revenue leading to successive layoffs that reduce the quality of the journalism they publish while their websites are stuffed with poor quality ads if not locked behind a paywall altogether.

Meanwhile, the promise of streaming entertainment has turned into a nightmare after corporate consolidation and the sidelining of competition. Subscription prices keep going higher, ads are increasingly part of the deal, and the old promise of unlimited access is gone as they keep yanking down programming for tax breaks and cost savings. Ecommerce hasn’t been spared either. Amazon feels like it’s taken over, but in the meantime it pulled back from being the main seller to turn itself into a poorly governed marketplace where deceptively presented low-quality goods proliferate and everyone has to pay more to pad the company’s bottom line. Don’t even get me started on the trends Shein and Temu are pioneering.

Generative AI closes off a better future
Ursula Le Guin said we must be able to imagine freedom. AI traps us in the past.
The digital revolution has failedDisconnectParis Marx
The digital revolution has failed

But it doesn’t end there. As tech companies sought to escape the confines of the web and move into the wider world, they’ve left a trail of disaster in their wake. The gig economy pretended app-based mediation was worthy of tearing hard-won workers’ rights to shreds, while data-hungry companies stuck surveillance devices anywhere they could get away with. The promise of algorithmic efficiency caused discriminatory systems to proliferate through society with little consideration for the wider repercussions. The effort to route as many interactions as possible through apps and make our smartphones as addictive as possible has spawned an epidemic of loneliness and even social disconnection.

Add to all that how the effort to stick screens, connectivity, and voice commands in as many places as possible has not only created a steady wave of devices that quickly become e-waste, but a wider problem where our appliances don’t last nearly as long in part because that tech fails so easily and our cars are less safe because key functions were shifted onto large touchscreens in anticipation of a driverless revolution that never arrived. And just as those issues were piling up, generative AI arrived on scene to make it even worse.

Enter the AI-generated swamp

The models behind the chatbots and visual generators that have taken the tech industry by storm over the past year were made by ingesting as much data as these companies could capture online as possible. That included whatever images and text they could find, including copyrighted books, works of art, news articles, and even the user-generated content and data billions of people have been spreading across the web for the past three decades. They took that collectively produced trove of information and communication as the foundation of their own businesses.

The tools they unleashed onto the world have only accelerated the trajectory of the web, filling social media feeds with AI-generated videos and text (some of it even produced by artificial users), tempting declining news media to try to pass AI-generated stories off as real, and worsening the quality of Google search results with the new wave of AI-generated garbage taking the web by storm. There’s a conspiracy theory that the web is effectively dead, and made up primarily of bots and generated content. While that may not already be true, the AI companies seem determined to make it a reality.

If you listen to CEOs like Sam Altman or venture capitalists like Marc Andreessen, they want us to believe that these tools are the beginning of a vast expansion in human potential, but that’s incredibly hard to believe for anyone who knows the history of Silicon Valley’s deception and can see through the hype to understand how these tools actually work. They’re not intelligent or prescient; they’re just churning out synthetic material that aligns with all the connections they’ve made between the training data they pulled from the open web.

Once again, the push to adopt these AI technologies isn’t about making our lives better, it’s about reducing the cost of producing ever more content to keep people engaged, to serve ads against, and keep people subscribed to struggling streaming services. The public doesn’t want the quality of news, entertainment, and human interactions to further decline because of the demands of investors for even greater profits, but that doesn’t matter. Everything must be sacrificed on the altar of tech capitalism.

AI is fueling a data center boom. It must be stopped.
Silicon Valley believes more computation is essential for progress. But they ignore the resource burden and don’t care if the benefits materialize.
The digital revolution has failedDisconnectParis Marx
The digital revolution has failed

Probably the most perverse aspect of it all is that making all of that (often quite poor quality) synthetic media has such an enormous cost. On the one hand you have the labor impact, where the people doing some of the jobs we’d most want done by fellow humans and maybe even to have more people engaging in — the creative work of writing and art that enriches human society — are the first to be targeted by people who seem most divorced from the human condition. But then layer on top of that the energy, water, and mineral cost of running all the data centers behind those tools, along with Altman’s comments about the potential need for geoengineering to make his dystopian future a reality, and it shows how little the proliferation of AI tools has to offer us.

Another internet is possible

There can only be one conclusion from all of this: the digital revolution has failed. The initial promise was a deception to lay the foundation for another corporate value-creation scheme, but the benefits that emerged from it have been so deeply eroded by commercial imperatives that the drawbacks far outweigh the remaining redeeming qualities — and that only gets worse with every day generative AI tools are allowed to keep flooding the web with synthetic material.

The time for tinkering around the edges has passed, and like a phoenix rising from the ashes, the only hope to be found today is in seeking to tear down the edifice the tech industry has erected and to build new foundations for a different kind of internet that isn’t poisoned by the requirement to produce obscene and ever-increasing profits to fill the overflowing coffers of a narrow segment of the population.

There were many networks before the internet, and there can be new networks that follow it. We don’t have to be locked into the digital dystopia Silicon Valley has created in a network where there was once so much hope for something else entirely. The ongoing erosion already seems to be sending people fleeing by ditching smartphones (or at least trying to reduce how much they use them), pulling back from the mess that social media has become, and ditching the algorithmic soup of streaming services.

Personal rejection is a welcome development, but as the web declines, we need to consider what a better alternative could look like and the political project it would fit within. We also can’t fall for any attempt to cast a libertarian “declaration of independence” as a truly liberatory future for everyone.

Sign up for Disconnect

Tech critic Paris Marx provides the critical take on Silicon Valley you’ve been missing

Email sent! Check your inbox to complete your signup.

No spam. Unsubscribe anytime.

2023 : une année de consolidation, de diversification et d’innovations

30 janvier 2024 à 11:00

2023 : une année de consolidation, de diversification et d’innovations

Numéricité est un studio des transformations, pour un numérique au service de l’impact et de l’intérêt général. Choisir Numéricité, c’est opter pour une méthode originale qui transforme la transformation (et les organisations).

L’année 2023 a été marquée par une consolidation et une diversification de nos activités.

Nous avons affirmé notre expertise dans la formation et le coaching des équipes projets, l’élaboration de stratégies data, la conception d’ateliers innovants et stimulants, ainsi que le développement de services numériques performants, toujours en conformité avec le RGPD.

Retour sur les moments phares qui ont rythmé cette année.

Numéricité a poursuivi sa croissance et s’est affirmé comme un acteur de la transformation numérique. Pour continuer d’accompagner au mieux les projets de nos clients, huit nouvelles personnes ont rejoint la famille Numéricité. Par ordre d’arrivée :

  • Romain Tales, Digital Transformation Specialist
  • Alexandre Le Borgne, Juriste en droit du numérique public
  • Daniel Manéa, Apprenti développeur full-stack
  • Mathilde Bras, Strategy & Innovation Manager
  • Maxime Lubrano, Responsable communication multimédia et stratégie éditoriale
  • Sarah Pilhan, Chargée de communication multimédia et stratégie éditoriale
  • Antoine Lelong, Développeur full-stack, apprenti recruté en CDI
  • Sam Khouma, Chef de projet & Product owner

Notre réseau d’experts s’est renforcé. Mathilde Hoang, Experte en transformation numérique, nous a rejoint en février. Cette collaboration nous permet d’avoir une deuxième implantation internationale, à Hanoï au Vietnam. Nous avons également renouvelé notre collaboration avec Vincent Motte, Chef de projet, et maintenons ainsi notre présence à Abidjan en Côte-d’Ivoire.

Afin d’œuvrer à la réussite des projets de nos clients, nous mettons nos expertises en commun. Celles-ci sont de plus en plus reconnues et demandées. Certains de nos experts sont d’ailleurs référencés par de grandes institutions internationales telles que le FMI ou la Banque Mondiale, et sélectionnés par ailleurs dans le cadre du Collège Numérique France 2030 du Secrétariat général pour l’investissement.

Un pôle communication s’est structuré au sein de Numéricité. Le défi ? Traduire et transmettre les expertises de l’équipe et mieux faire connaitre l’accompagnement 360° que l’on peut proposer qu’il s’agisse de conformité juridique, de développement produit en mode agile ou d’accompagnement en transformation numérique.

Cette année a également été marquée par la poursuite de nos partenariats avec d’autres acteurs de la transformation numérique tels qu’Octo, Nextmap, La Javaness, Hubvisory, Bearing Point, Malt et Opteamis.

Au sein de cet écosystème, nous faisons de notre connaissance fine du numérique public français une force. Plusieurs membres de l’équipe en ont été les forgerons dans une précédente vie professionnelle. En conciliant cela aux méthodes et techniques issues de la culture numérique, nous proposons une approche originale en mesure de donner de nouvelles capacités aux organisations que nous accompagnons.

Ensemble, transformons la transformation numérique pour en faire un levier d’amélioration de vos organisations : 

Les entreprises et administrations publiques sont convaincues de la valeur de la donnée qui est un actif stratégique pour l’ensemble de l’organisation et pas seulement pour un seul service ou une direction métier. C’est pourquoi les entreprises et les administrations publiques se transforment afin de mieux valoriser ces données et investissent dans la construction et la mise en place d’une stratégie data au service de leur stratégie globale.

Nous appuyons cette transformation à travers la co-construction de stratégies data portant sur l’ouverture des données, la modernisation des SI afin de les orienter vers une meilleure circulation de la donnée, et construisons des produits data à destination des métiers. Nous intervenons ainsi dans des secteurs à l’impact quotidien : les mobilités, la santé, le logement, le travail, l’inclusion, l’environnement, la fiscalité, les douanes, etc.

Nous réinterrogeons et transformons les organisations à travers une meilleure prise en compte des données au service de l’impact : c’est une des traductions concrètes de notre offre Transfo-soft.

Afin de faciliter la vie des usagers, citoyens comme entreprises, notre pôle Intelligence juridique et conformité a appuyé l’ouverture des données, comme celles produites par France Chaleur Urbaine ou MaCantine.

Moderniser les SI est un prérequis pour faciliter la circulation de la donnée. Pour cela, nous proposons une expertise d’audit nous permettant de formuler des recommandations d’actions concrètes. L’objectif est d’engager la transformation par les métiers, en améliorant leur process internes ou leur organisation produit par exemple.

Nous avons accompagné de nombreux acteurs, territoires et gouvernements dans la concertation, la construction et la priorisation de leurs stratégies data. Un seul impératif : mettre l’humain, que ce soit les métiers, les acteurs économiques et/ou les usagers au cœur de la réflexion.

Pour embarquer ces différents collectifs, qu’ils soient spécialistes ou non de la data, nous disposons d’un large éventail de formats : open-lab, ateliers, entretiens, expérimentations, cas d’usage, datadays, hackathons. Il s’agit là d’une première brique de notre approche produit des stratégies data, visant ensuite à s’interroger sur les données à mobiliser, la gouvernance à constituer, les compétences à renforcer, avant de réaliser des expérimentations à petite échelle pour confirmer les orientations.

Ces missions ont largement ponctué notre activité en 2023. Voici quelques exemples emblématiques :

  • Réaliser le bilan de la feuille de route du numérique 2021-2023 du Ministère de l’Éducation nationale et de la Jeunesse et des Académies et identifier les actions du futur plan relatif aux données et codes sources pour la période 2024-2027 à travers un open-lab.
  • Mobiliser le potentiel des datas pour améliorer les mobilités à l’échelle régionale lors d’un hackathon au cours duquel les participants pouvaient répondre aux quatre défis pensés en lien avec la stratégie open data d’Île-de-France Mobilités.
  • Aligner les différentes parties prenantes du Groupe Vyv sur une position ambitieuse autour de l’ouverture des données assurantielles, qui permette de mettre l’usager au centre de l’utilisation de ces données, en organisant une série d’open-lab.
  • Accompagner la rédaction de la feuille de route de la Direction du Numérique et de la Modernisation de Nouvelle-Calédonie, en organisant en partenariat avec Atlas Management une série d’ateliers, en proposant un parangonnage des exemples internationaux et en coachant des prototypes (POC) data appuyés sur la donnée.

Nous suivons la même méthode produit dans la conception de services numériques appuyés sur des données. 2023 a été marquée par la conception de 10 produits numériques par notre Pôle design et technologies.

Par exemple, avec La Fabrique des ministères sociaux et l’ARS Île-de-France, nous avons conçu et développé un outil de datavisualisation des causes médicales de décès, s’appuyant sur quelques 315K certificats de décès franciliens émis depuis 2020 et l’industrialisation d’un flux de données de l’INSERM.

Nous avons également accompagné l’Agence Nationale d’Information sur le Logement (ANIL) dans l’élaboration d’une base de données partagée avec ses relais départementaux, afin de faciliter le partage d’informations au sein du réseau. AdilStatWeb2 a été consulté 815 000 fois et une interface de croisement des données a été créée.

 

La démocratisation des intelligences artificielle génératives (IAG) est en train de bousculer nos pratiques à une vitesse fulgurante. Avec ces avancées, viennent des défis complexes en matière de sécurité, de responsabilité et de protection des droits fondamentaux.

Pour en savoir plus sur l’encadrement législatif de l’intelligence artificielle, (re)découvrez nos derniers articles de décryptage :

Écoutez notre podcast sur la régulation du numérique par le législateur avec Éric Bothorel, député de la 5e circonscription des Côtes-d’Armor : 

En interne, nous mobilisons plusieurs outils d’IAG (ChatGPT, MidJourney et Copilot) pour gagner en maîtrise et en efficacité tout en explorant les fonctionnalités pour de futures missions. Nous sommes d’ailleurs en train d’entraîner un NuméricitéGPT, connecté à OpenAI et basé sur le scraping des contenus de notre site, pour répondre à des besoins de communication et de partage des savoirs en interne.

Dans le cadre de nos missions, nous nous sommes particulièrement intéressés au domaine des agents conversationnels appuyés sur de grands modèles de langage (Large Language Model ou LLM). Ces innovations nous ont permis de développer des solutions adaptables, qui garantissent la protection des données exploitées.

Nous avons entraîné un modèle basé sur le LLM Mistral pour expérimenter l’usage d’un agent conversationnel sur des questions de santé sexuelle. Pour répondre, l’agent s’appuie sur des contenus web scrapés sur différents sites officiels.

Nous avons prototypé deux agents conversationnels connectés à OpenAI grâce à une clé d’API. L’hébergement de ces bots sur nos serveurs garanti la souveraineté en assurant qu’aucune données fiscales sensibles ne soient transmises à OpenAI. Grâce à ces bots, les agents de la Direction Générale des Impôts (DGI) de Mauritanie acquièrent de l’autonomie dans leurs tâches quotidiennes en pouvant générer de manière automatisée des rapports d’analyse des indicateurs issus de leur tableau de bord (bot Reporter) et interroger directement leurs jeux de données (FileBot).

Notre pôle Intelligence juridique et conformité participe à déterminer et maîtriser les impacts de l’IA dans l’administration. Cela l’a conduit à élaborer différents guides de bonnes pratiques. Avec La Javaness, nous avons accompagné Pôle Emploi afin de confronter l’offre de service de la plateforme IA aux différents usages, afin d’en adapter le contenu et inscrire ce produit et les services associés dans un processus d’amélioration. L’enjeu était de créer un socle IA du service public de l’emploi qui soit un gage de souveraineté et d’efficacité pour la puissance publique.

Numéricité prône une approche 360° de la conformité : nos missions s’articulent entre la protection de la vie privée et la circulation des données. Nous intervenons tant à l’échelle d’un projet qu’à l’échelle d’une organisation et proposons pour cela une conception agile de la conformité. Pour cela, nous travaillons au plus près des équipes, en leur proposant de nombreux formats adaptés à leurs besoins, que ce soit des cliniques, des formations, des stands ou du question/réponse via messagerie instantanée.

Nous avons accompagné la conformité d’acteurs de toutes tailles, publics comme privés, dont plus d’une centaine de Startups d’État.

Nous avons pu également administrer un premier audit “blanc” à l’Unédic, inspiré des audits de certification appliqués par la CNIL, afin de formuler des recommandations et des corrections.

Chez Numéricité, nous avons l’open source chevillé au corps (et au cœur). Dans l’ensemble de nos projets, nous nous efforçons de mettre en place ou orienter nos clients vers des solutions open source.

Par exemple, nous avons déployé Onyxia afin de permettre aux agents de quatre pays d’Afrique de l’Ouest partenaires du projet DATAFID de se former et s’entraîner à la datascience avec pour objectif d’améliorer les recettes fiscales et douanières. Pour en savoir plus sur Onyxia, et plus largement sur le projet DATAFID. 

Nos convictions dans les communs se sont aussi traduites par notre participation à la sixième édition de Numérique en Commun[s]. À cette occasion, nous avons pu proposer deux ateliers sur la conception de produits data avec et pour des utilisateurs non spécialistes de la donnée, et sur le RGPD comme accélérateur des projets. Nous avons également accompagné l’Institut national de l’information géographique et forestière (IGN), partenaire de l’évènement, dans la conception et l’éditorialisation de ses formats.

Vivez notre aventure à Numérique en Commun[s] et découvrez les cinq principes d’action pour développer des produits data avec et pour des utilisateurs non spécialistes de la donnée issus de l’atelier :

Cette rétrospective de nos activités en 2023 serait incomplète sans parler nos interventions à l’international et notre expertise e-gov. Nous intervenons actuellement dans 15 pays.

Carte des pays d'intervention de Numéricité en 2023

Nos interventions e-gov sont au croisement de l’ensemble de nos expertises.

  • Accompagnement dans l’élaboration de stratégies data
    • Au Rwanda, nous avons poursuivi notre mission afin d’évaluer la maturité numérique des administrations, mettre en place et animer la gouvernance du réseau interministériel des “chiefs digital officers”.
    • À Madagascar, nous sommes intervenus sur les enjeux du partage des données entre administration et la nécessaire rationalisation des architectures du SI en vue de la numérisation de l’État-civil, afin d’Insuffler la dynamique et la culture de l’interopérabilité au sein d’administrations fonctionnant en silo.
    • Au Botswana, nous avons formulé un cadrage stratégique des activités de l’opérationnalisation de la stratégie SmartBots, visant à stimuler la transformation numérique dans l’économie, le gouvernement et la société.

Nos experts référencés par le Fonds Monétaire International (FMI) ont par ailleurs appuyé la réalisation d’études préparatoires au Cameroun et en Tunisie.

    • En juillet 2023, nous avons été sollicités afin de construire et faciliter un séminaire de partage d’expérience organisé par le FMI pour 8 administrations fiscales d’Afrique centrale à Yaoundé : Burundi, Cameroun, Guinée équatoriale, République centrafricaine, République du Congo, République Démocratique du Congo, Sao Tomé-et-Principe, et Tchad.
    • En octobre 2023, nous avons évaluer la gestion de la conformité et l’analyse des données de la DGI tunisienne afin de formuler des recommandations techniques, et sensibiliser les agents et les cadres au lien avec le Compliance Risk Management.

  • Ingénierie juridique et conformité :
    • En Libye, au Bénin, en Mauritanie, au Cameroun et au Vietnam, nous avons réalisé une cartographie du droit applicable, et formulé des guides des bonnes pratiques afin d’orienter les édifices juridiques de ces pays.
    • Pour Ho Chi Minh Ville, nous avons analysé les textes prévus dans le cadre du déploiement d’une plateforme de services publics en lignes afin d’identifier les manques pour la mise en application opérationnelle (gouvernance des données, régulation, responsabilité, interopérabilité, etc.). Nous avons également formulé une série de recommandations en nous appuyant sur les meilleurs pratiques issues des standards internationaux en la matière.

Pour en savoir plus sur notre accompagnement de la Libye dans l’appréhension des enjeux de la transformation numérique

  • Ruptures et expérimentations technologiques :
    • Mobiliser des modèles d’intelligence artificielle pour améliorer l’efficacité du contrôle douanier ? Nous l’avons rendu possible en expérimentant RandomForestClassifier, un algorithme d’apprentissage supervisé, avec les douanes nigériennes.
    • Automatiser la génération de rapports de données grâce à l’intelligence artificielle ? Nous avons prototypé deux bots pour la Direction Générale des Impôts mauritanienne. Le bot “Reporter” et le “FileBot” sont des Proof of concept servant de base à une étude préalable avant de développer un modèle d’intelligence artificielle générative.

  • Formations et transmission d’expertise :
    • Nous avons rédigé un guide à destination des cadres de l’administration vietnamienne, sur les bonnes pratiques françaises en matière de dématérialisation de l’administration.
    • En Côte-d’Ivoire, au Togo et en Mauritanie, 180 agents des administrations fiscales et douanières ont été formés aux techniques de datascience, à travers plus de 80 sessions de formations, donnant lieu à la création d’une quarantaine de tableaux de bord.
    • Les agents des impôts du Sénégal ont suivi nos formations aux techniques de webscraping et datamining.

2023 a été une année de succès, d’innovation et d’impact pour Numéricité. Nous nous imposons comme un acteur de l’écosystème de la transformation numérique, tout en y apportant notre touche d’originalité.

Nous continuons de transformer la transformation numérique avec une équipe talentueuse d’experts et des projets ambitieux, en portant une attention particulière à l’impact et à l’éthique. 

Pour en découvrir davantage sur notre méthode et envisager de nouveaux projets ensemble, contactez-nous : contact@numericite.eu

Feed Time

Par : Doc Searls
4 avril 2024 à 21:25
I asked ChatGPT to give me “people eating blogs” and got this after it suggested some details.

Two things worth blogging about that happened this morning.

One was getting down and dirty trying to make DALL-E 3 work. That turned into giving up trying to find DALL-E (in any version) on the open Web and biting the $20/month bullet for a Pro account with ChatGPT, which for some reason maintains its DALL-E 3 Web page while having “Try in ChatGPT↗︎” on that page link to the ChatGPT home page rather than a DALL-E one. I gather that the free version of DALL-E is now the one you get at Microsoft’s Copilot | Designer, while the direct form of DALL-E is what you get when you prompt ChatGPT (now 4.0 for Pro customers… or so I gather) to give you an image that credits nothing to DALL-E.

The other thing was getting some great help from Dave Winer in putting the new Feedroll category of my Feedland feeds placed on this blog, in a way similar stylistically to old-fashioned blogrolls (such as the one here). You’ll find it in the right column of this blog now. One cool difference from blogrolls is that the feedroll is live. Very cool. I’m gradually expanding it.

Meanwhile, after failing to get ChatGPT or Copilot | Designer to give me the image I needed on another topic (which I’ll visit here later) I prompted them to give me an image that might speak to a feedroll of blogs. ChatGPT gave me the one above, not in response to “people eating blogs” (my first attempt), but instead to “People eating phone, mobile and computer screens of type.” Microsoft | Designer gave me these:

Redraw your own inconclusions.

Looking for DALL-E 3 Help

Par : Doc Searls
4 avril 2024 à 15:48

dall-e 3 header
I just returned to DALL-E 3 after using its Microsoft version (currently called Copilot | Designer) for a while. But I can’t get in.

See how it says “Try in ChatGPT↗︎?” When I do that, it goes to https://chat.openai.com/. After I log in there, it offers no clue about where DALL-E 3 is. So I asked, “How do I get to the DALL-E 3 page after I’m logged in here?”

ChatGPT answered, “As of my last update in January 2022, DALL-E 3 is a hypothetical iteration of the DALL-E model, and there’s no information about its existence or how to access it. If it were to be released, you might typically find access through OpenAI’s website or platform if they make it available to the public. However, without further updates or announcements, it’s challenging to provide specific instructions. If OpenAI releases DALL-E 3 or any similar model in the future, you can check their official website or announcements for information on how to access it.”

Additional useless dialog followed. So now I’ll ask the Lazyweb. Does anyone know how to get into DALL-E 3?

Personal AI at VRM Day and IIW

Par : Doc Searls
20 mars 2024 à 21:07

Prompt: A woman uses personal AI to know, get control of, and put to better use all available data about her property, health, finances, contacts, calendar, subscriptions, shopping, travel, and work. Via Microsoft Copilot Designer, with spelling corrections by the author.

Most AI news is about what the giants (OpenAI/Microsoft, Meta, Google/Apple, Amazon, Adobe, Nvidia) are doing (seven $trillion, anyone?), or what AI is doing for business (all of Forbes’ AI 50). Against all that, personal AI appears to be about where personal computing was in 1974: no longer an oxymoron but discussed more than delivered.

For evidence, look up “personal AI.” All the results will be about business (see here and here) or “assistants” that are just suction cups on the tentacles of giants (Siri, Google Assistant, Alexa, Bixby), or wannabes that do the same kind of thing (Lindy, Hound, DataBot).

There may be others, but three exceptions I know are Kin, Personal AI and Pi.

Personal AI is finding its most promoted early uses on the side of business more than the side of customers. Zapier, for example, explains that Personal AI “can be used as a productivity or business tool.”

Kin and Pi are personal assistants that help you with your life by surveilling your activities for your own benefit. I’ve signed up for both, but have only experienced Pit,” or “just vent,” when I ask it to help me with the stuff outlined in (and under) the AI-generated image above, it wants to hook me up with a bunch of siloed platforms that cost money, or to do geeky things (PostgreSQL, MongoDB, Python on my own computer. Provisional conclusion: Pi means well, but the tools aren’t there yet. [Later… Looks like it’s going to morph into some kind of B2B thing, or be abandoned outright, now that Inflection AI’s CEO, Mustafa Suleyman is gone to Microsoft. Hmm… will Microsoft do what we’d like in this space?]

Open source approaches are out there: OpenDAN, Khoj, Kwaai , and Llama are four, and I know at least one will be at VRM Day and IIW.

So, since personal AI may finally be what pushes VRM into becoming a Real Thing, we’ll make it the focus of our next VRM Day.

As always, VRM Day will precede IIW in the same location: the Boole Room of the Computer History Museum in Mountain View, just off Highway 101 in the heart of Silicon Valley. It’ll be on Monday, 15 April, and start at 9am. There’s a Starbucks across the street and ample parking because the museum is officially closed on Mondays, but the door is open. We lunch outdoors (it’s always clear) at the sports bar on the other corner.

Registration is open now at this Eventbrite link:

https://vrmday2024a.eventbrite.com

You can also just show up, but registering gives us a rough headcount, which is helpful for bringing in the right number of chairs and stuff like that.

See you there!

 

DAOFest Marseille ❤️

Avec 50 participants et plus de 10 intervenants, nous avons réinventé le format des Barcamp. J’ai vraiment apprécié cette expérience. Je suis extrêmement satisfait du format et des interactions. Il était audacieux d’aborder une innovation radicale telle que les DAO à travers mon expérience sur la cartographie, remettant en question la souveraineté. Mais aussi déconstruire …

Continuer la lecture « DAOFest Marseille ❤️ »

How to Think Computationally about AI, the Universe and Everything

27 octobre 2023 à 21:47

Transcript of a talk at TED AI on October 17, 2023, in San Francisco

Human language. Mathematics. Logic. These are all ways to formalize the world. And in our century there’s a new and yet more powerful one: computation.

And for nearly 50 years I’ve had the great privilege of building an ever taller tower of science and technology based on that idea of computation. And today I want to tell you some of what that’s led to.

There’s a lot to talk about—so I’m going to go quickly… sometimes with just a sentence summarizing what I’ve written a whole book about.

You know, I last gave a TED talk thirteen years ago—in February 2010—soon after Wolfram|Alpha launched.

TED Talk 2010

And I ended that talk with a question: is computation ultimately what’s underneath everything in our universe?

I gave myself a decade to find out. And actually it could have needed a century. But in April 2020—just after the decade mark—we were thrilled to be able to announce what seems to be the ultimate “machine code” of the universe.

Wolfram Physics Project

And, yes, it’s computational. So computation isn’t just a possible formalization; it’s the ultimate one for our universe.

It all starts from the idea that space—like matter—is made of discrete elements. And that the structure of space and everything in it is just defined by the network of relations between these elements—that we might call atoms of space. It’s very elegant—but deeply abstract.

But here’s a humanized representation:

A version of the very beginning of the universe. And what we’re seeing here is the emergence of space and everything in it by the successive application of very simple computational rules. And, remember, those dots are not atoms in any existing space. They’re atoms of space—that are getting put together to make space. And, yes, if we kept going long enough, we could build our whole universe this way.

Eons later here’s a chunk of space with two little black holes, that eventually merge, radiating ripples of gravitational radiation:

And remember—all this is built from pure computation. But like fluid mechanics emerging from molecules, what emerges here is spacetime—and Einstein’s equations for gravity. Though there are deviations that we just might be able to detect. Like that the dimensionality of space won’t always be precisely 3.

And there’s something else. Our computational rules can inevitably be applied in many ways, each defining a different thread of time—a different path of history—that can branch and merge:

But as observers embedded in this universe, we’re branching and merging too. And it turns out that quantum mechanics emerges as the story of how branching minds perceive a branching universe.

The little pink lines here show the structure of what we call branchial space—the space of quantum branches. And one of the stunningly beautiful things—at least for a physicist like me—is that the same phenomenon that in physical space gives us gravity, in branchial space gives us quantum mechanics.

In the history of science so far, I think we can identify four broad paradigms for making models of the world—that can be distinguished by how they deal with time.

4 paradigms

In antiquity—and in plenty of areas of science even today—it’s all about “what things are made of”, and time doesn’t really enter. But in the 1600s came the idea of modeling things with mathematical formulas—in which time enters, but basically just as a coordinate value.

Then in the 1980s—and this is something in which I was deeply involved—came the idea of making models by starting with simple computational rules and then just letting them run:

Can one predict what will happen? No, there’s what I call computational irreducibility: in effect the passage of time corresponds to an irreducible computation that we have to run to know how it will turn out.

But now there’s something even more: in our Physics Project things become multicomputational, with many threads of time, that can only be knitted together by an observer.

It’s a new paradigm—that actually seems to unlock things not only in fundamental physics, but also in the foundations of mathematics and computer science, and possibly in areas like biology and economics too.

You know, I talked about building up the universe by repeatedly applying a computational rule. But how is that rule picked? Well, actually, it isn’t. Because all possible rules are used. And we’re building up what I call the ruliad: the deeply abstract but unique object that is the entangled limit of all possible computational processes. Here’s a tiny fragment of it shown in terms of Turing machines:

OK, so the ruliad is everything. And we as observers are necessarily part of it. In the ruliad as a whole, everything computationally possible can happen. But observers like us can just sample specific slices of the ruliad.

And there are two crucial facts about us. First, we’re computationally bounded—our minds are limited. And second, we believe we’re persistent in time—even though we’re made of different atoms of space at every moment.

So then here’s the big result. What observers with those characteristics perceive in the ruliad necessarily follows certain laws. And those laws turn out to be precisely the three key theories of 20th-century physics: general relativity, quantum mechanics, and statistical mechanics and the Second Law.

It’s because we’re observers like us that we perceive the laws of physics we do.

We can think of different minds as being at different places in rulial space. Human minds who think alike are nearby. Animals further away. And further out we get to alien minds where it’s hard to make a translation.

How can we get intuition for all this? We can use generative AI to take what amounts to an incredibly tiny slice of the ruliad—aligned with images we humans have produced.

We can think of this as a place in the ruliad described using the concept of a cat in a party hat:

Zooming out, we see what we might call “cat island”. But pretty soon we’re in interconcept space. Occasionally things will look familiar, but mostly we’ll see things we humans don’t have words for.

In physical space we explore more of the universe by sending out spacecraft. In rulial space we explore more by expanding our concepts and our paradigms.

We can get a sense of what’s out there by sampling possible rules—doing what I call ruliology:

Even with incredibly simple rules there’s incredible richness. But the issue is that most of it doesn’t yet connect with things we humans understand or care about. It’s like when we look at the natural world and only gradually realize we can use features of it for technology. Even after everything our civilization has achieved, we’re just at the very, very beginning of exploring rulial space.

But what about AIs? Just like we can do ruliology, AIs can in principle go out and explore rulial space. But left to their own devices, they’ll mostly be doing things we humans don’t connect with, or care about.

The big achievements of AI in recent times have been about making systems that are closely aligned with us humans. We train LLMs on billions of webpages so they can produce text that’s typical of what we humans write. And, yes, the fact that this works is undoubtedly telling us some deep scientific things about the semantic grammar of language—and generalizations of things like logic—that perhaps we should have known centuries ago.

You know, for much of human history we were kind of like LLMs, figuring things out by matching patterns in our minds. But then came more systematic formalization—and eventually computation. And with that we got a whole other level of power—to create truly new things, and in effect to go wherever we want in the ruliad.

But the challenge is to do that in a way that connects with what we humans—and our AIs—understand.

And in fact I’ve devoted a large part of my life to building that bridge. It’s all been about creating a language for expressing ourselves computationally: a language for computational thinking.

The goal is to formalize what we know about the world—in computational terms. To have computational ways to represent cities and chemicals and movies and formulas—and our knowledge about them.

It’s been a vast undertaking—that’s spanned more than four decades of my life. It’s something very unique and different. But I’m happy to report that in what has been Mathematica and is now the Wolfram Language I think we have now firmly succeeded in creating a truly full-scale computational language.

In effect, every one of the functions here can be thought of as formalizing—and encapsulating in computational terms—some facet of the intellectual achievements of our civilization:

It’s the most concentrated form of intellectual expression I know: finding the essence of everything and coherently expressing it in the design of our computational language. For me personally it’s been an amazing journey, year after year building the tower of ideas and technology that’s needed—and nowadays sharing that process with the world on open livestreams.

A few centuries ago the development of mathematical notation, and what amounts to the “language of mathematics”, gave a systematic way to express math—and made possible algebra, and calculus, and ultimately all of modern mathematical science. And computational language now provides a similar path—letting us ultimately create a “computational X” for all imaginable fields X.

We’ve seen the growth of computer science—CS. But computational language opens up something ultimately much bigger and broader: CX. For 70 years we’ve had programming languages—which are about telling computers in their terms what to do. But computational language is about something intellectually much bigger: it’s about taking everything we can think about and operationalizing it in computational terms.

You know, I built the Wolfram Language first and foremost because I wanted to use it myself. And now when I use it, I feel like it’s giving me a superpower:

I just have to imagine something in computational terms and then the language almost magically lets me bring it into reality, see its consequences and then build on them. And, yes, that’s the superpower that’s let me do things like our Physics Project.

And over the past 35 years it’s been my great privilege to share this superpower with many other people—and by doing so to have enabled such an incredible number of advances across so many fields. It’s a wonderful thing to see people—researchers, CEOs, kids—using our language to fluently think in computational terms, crispening up their own thinking and then in effect automatically calling in computational superpowers.

And now it’s not just people who can do that. AIs can use our computational language as a tool too. Yes, to get their facts straight, but even more importantly, to compute new facts. There are already some integrations of our technology into LLMs—and there’s a lot more you’ll be seeing soon. And, you know, when it comes to building new things, a very powerful emerging workflow is basically to start by telling the LLM roughly what you want, then have it try to express that in precise Wolfram Language. Then—and this is a critical feature of our computational language compared to a programming language—you as a human can “read the code”. And if it does what you want, you can use it as a dependable component to build on.

OK, but let’s say we use more and more AI—and more and more computation. What’s the world going to be like? From the Industrial Revolution on, we’ve been used to doing engineering where we can in effect “see how the gears mesh” to “understand” how things work. But computational irreducibility now shows that won’t always be possible. We won’t always be able to make a simple human—or, say, mathematical—narrative to explain or predict what a system will do.

And, yes, this is science in effect eating itself from the inside. From all the successes of mathematical science we’ve come to believe that somehow—if only we could find them—there’d be formulas to predict everything. But now computational irreducibility shows that isn’t true. And that in effect to find out what a system will do, we have to go through the same irreducible computational steps as the system itself.

Yes, it’s a weakness of science. But it’s also why the passage of time is significant—and meaningful. We can’t just jump ahead and get the answer; we have to “live the steps”.

It’s going to be a great societal dilemma of the future. If we let our AIs achieve their full computational potential, they’ll have lots of computational irreducibility, and we won’t be able to predict what they’ll do. But if we put constraints on them to make them predictable, we’ll limit what they can do for us.

So what will it feel like if our world is full of computational irreducibility? Well, it’s really nothing new—because that’s the story with much of nature. And what’s happened there is that we’ve found ways to operate within nature—even though nature can still surprise us.

And so it will be with the AIs. We might give them a constitution, but there will always be consequences we can’t predict. Of course, even figuring out societally what we want from the AIs is hard. Maybe we need a promptocracy where people write prompts instead of just voting. But basically every control-the-outcome scheme seems full of both political philosophy and computational irreducibility gotchas.

You know, if we look at the whole arc of human history, the one thing that’s systematically changed is that more and more gets automated. And LLMs just gave us a dramatic and unexpected example of that. So does that mean that in the end we humans will have nothing to do? Well, if you look at history, what seems to happen is that when one thing gets automated away, it opens up lots of new things to do. And as economies develop, the pie chart of occupations seems to get more and more fragmented.

And now we’re back to the ruliad. Because at a foundational level what’s happening is that automation is opening up more directions to go in the ruliad. And there’s no abstract way to choose between them. It’s just a question of what we humans want—and it requires humans “doing work” to define that.

A society of AIs untethered by human input would effectively go off and explore the whole ruliad. But most of what they’d do would seem to us random and pointless. Much like now most of nature doesn’t seem like it’s “achieving a purpose”.

One used to imagine that to build things that are useful to us, we’d have to do it step by step. But AI and the whole phenomenon of computation tell us that really what we need is more just to define what we want. Then computation, AI, automation can make it happen.

And, yes, I think the key to defining in a clear way what we want is computational language. You know—even after 35 years—for many people the Wolfram Language is still an artifact from the future. If your job is to program it seems like a cheat: how come you can do in an hour what would usually take a week? But it can also be daunting, because having dashed off that one thing, you now have to conceptualize the next thing. Of course, it’s great for CEOs and CTOs and intellectual leaders who are ready to race onto the next thing. And indeed it’s impressively popular in that set.

In a sense, what’s happening is that Wolfram Language shifts from concentrating on mechanics to concentrating on conceptualization. And the key to that conceptualization is broad computational thinking. So how can one learn to do that? It’s not really a story of CS. It’s really a story of CX. And as a kind of education, it’s more like liberal arts than STEM. It’s part of a trend that when you automate technical execution, what becomes important is not figuring out how to do things—but what to do. And that’s more a story of broad knowledge and general thinking than any kind of narrow specialization.

You know, there’s an unexpected human-centeredness to all of this. We might have thought that with the advance of science and technology, the particulars of us humans would become ever less relevant. But we’ve discovered that that’s not true. And that in fact everything—even our physics—depends on how we humans happen to have sampled the ruliad.

Before our Physics Project we didn’t know if our universe really was computational. But now it’s pretty clear that it is. And from that we’re inexorably led to the ruliad—with all its vastness, so hugely greater than all the physical space in our universe.

So where will we go in the ruliad? Computational language is what lets us chart our path. It lets us humans define our goals and our journeys. And what’s amazing is that all the power and depth of what’s out there in the ruliad is accessible to everyone. One just has to learn to harness those computational superpowers. Which starts here. Our portal to the ruliad:

New AI Usage Data Shows Who’s Using AI — and Uncovers a Population of ‘Super-Users’ - Salesforce News

Dans cette étude flash de Salesforce sur les usages de l'IA on retrouve des conclusions connues : les plus jeunes sont les plus importants utilisateurs de l'IA générative et ils ont des usages clairs et précis. Globalement, les plus âgés comprennent moins bien la technologie, ce qui les bride dans leurs utilisations
Permalien

Remembering Doug Lenat (1950–2023) and His Quest to Capture the World with Logic

6 septembre 2023 à 00:23

Logic, Math and AI

In many ways the great quest of Doug Lenat’s life was an attempt to follow on directly from the work of Aristotle and Leibniz. For what Doug was fundamentally trying to do over the forty years he spent developing his CYC system was to use the framework of logic—in more or less the same form that Aristotle and Leibniz had it—to capture what happens in the world. It was a noble effort and an impressive example of long-term intellectual tenacity. And while I never managed to actually use CYC myself, I consider it a magnificent experiment—that if nothing else ultimately served to demonstrate the importance of building frameworks beyond logic alone in usefully representing and reasoning about the world.

Doug Lenat started working on artificial intelligence at a time when nobody really knew what might be possible—or even easy—to do. Was AI (whatever that might mean) just a clever algorithm—or a new type of computer—away? Or was it all just an “engineering problem” that simply required pulling together a bigger and better “expert system”? There was all sorts of mystery—and quite a lot of hocus pocus—around AI. Did the demo one was seeing actually prove something, or was it really just a trivial (if perhaps unwitting) cheat?

I first met Doug Lenat at the beginning of the 1980s. I had just developed my SMP (“Symbolic Manipulation Program”) system, that was the forerunner of Mathematica and the modern Wolfram Language. And I had been quite exposed to commercial efforts to “do AI” (and indeed our VCs had even pushed my first company to take on the dubious name “Inference Corporation”, complete with a “=>” logo). And I have to say that when I first met Doug I was quite dismissive. He told me he had a program (that he called “AM” for “Automated Mathematician”, and that had been the subject of his Stanford CS PhD thesis) that could discover—and in fact had discovered—nontrivial mathematical theorems.

“What theorems?” I asked. “What did you put in? What did you get out?” I suppose to many people the concept of searching for theorems would have seemed like something remarkable, and immediately exciting. But not only had I myself just built a system for systematically representing mathematics in computational form, I had also been enumerating large collections of simple programs like cellular automata. I poked at what Doug said he’d done, and came away unconvinced. Right around the same time I happened to be visiting a leading university AI group, who told me they had a system for translating stories from Spanish into English. “Can I try it?” I asked, suspending for a moment my feeling that this sounded like science fiction. “I don’t really know Spanish”, I said, “Can I start with just a few words?” “No”, they said, “the system works only with stories.” “How long does a story have to be?” I asked. “Actually it has to be a particular kind of story”, they said. “What kind?” I asked. There were a few more iterations, but eventually it came out: the “system” translated one particular story from Spanish into English! I’m not sure if my response included an expletive, but I wondered what kind of science, technology, or anything else this was supposed to be. And when Doug told me about his “Automated Mathematician”, this was the kind of thing I was afraid I was going to find.

Years later, I might say, I think there’s something AM could have been trying to do that’s valid, and interesting, if not obviously possible. Given a particular axiom system it’s easy to mechanically generate infinite collections of “true theorems”—that in effect fill metamathematical space. But now the question is: which of these theorems will human mathematicians find “interesting”? It’s not clear how much of the answer has to do with the “social history of mathematics”, and how much is more about “abstract principles”. I’ve been studying this quite a bit in recent years (not least because I think it could be useful in practice)—and have some rather deep conclusions about its relation to the nature of mathematics. But I now do wonder to what extent Doug’s work from all those years ago might (or might not) contain heuristics that would be worth trying to pursue even now.

CYC

I ran into Doug quite a few times in the early to mid-1980s, both around a company called Thinking Machines (to which I was a consultant) and at various events that somehow touched on AI. There was a fairly small and somewhat fragmented AI community in those days, with the academic part in the US concentrated around MIT, Stanford and CMU. I had the impression that Doug was never quite at the center of that community, but was somehow nevertheless a “notable member”, who—particularly with his work being connected to math—was seen as “doing upscale things” around AI.

In 1984 I wrote an article for a special issue of Scientific American on “computer software” (yes, software was trendy then). My article was entitled “Computer Software in Science and Mathematics”, and the very next article was by Doug, entitled “Computer Software for Intelligent Systems”. The summary at the top of my article read: “Computation offers a new means of describing and investigating scientific and mathematical systems. Simulation by computer may be the only way to predict how certain complicated systems evolve.” And the summary for Doug’s article read: “The key to intelligent problem solving lies in reducing the random search for solutions. To do so intelligent computer programs must tap the same underlying ‘sources of power’ as human beings”. And I suppose in many ways both of us spent most of our next four decades essentially trying to fill out the promise of these summaries.

A key point in Doug’s article—with which I wholeheartedly agree—is that to create something one can usefully identify as “AI”, it’s essential to somehow have lots of knowledge of the world built in. But how should that be done? How should the knowledge be encoded? And how should it be used?

Doug’s article in Scientific American illustrated his basic idea:

Click to enlarge

Encode knowledge about the world in the form of statements of logic. Then find ways to piece together these statements to derive conclusions. It was, in a sense, a very classic approach to formalizing the world—and one that would at least in concept be familiar to Aristotle and Leibniz. Of course it was now using computers—both as a way to store the logical statements, and as a way to find inferences from them.

At first, I think Doug felt the main problem was how to “search for correct inferences”. Given a whole collection of logical statements, he was asking how these could be knitted together to answer some particular question. In essence it was just like mathematical theorem proving: how could one knit together axioms to make a proof of a particular theorem? And especially with the computers and algorithms of the time, this seemed like a daunting problem in almost any realistic case.

But then how did humans ever manage to do it? What Doug imagined was that the critical element was heuristics: strategies for guessing how one might “jump ahead” and not have to do the kind of painstaking searches that systematic methods seemed to imply would be needed. Doug developed a system he called EURISKO that implemented a range of heuristics—that Doug expected could be used not only for math, but basically for anything, or at least anything where human-like thinking was effective. And, yes, EURISKO included not only heuristics, but also at least some kinds of heuristics for making new heuristics, etc.

But OK, so Doug imagined that EURISKO could be used to “reason about” anything. So if it had the kind of knowledge humans do, then—Doug believed—it should be able to reason just like humans. In other words, it should be able to deliver some kind of “genuine artificial intelligence” capable of matching human thinking.

There were all sorts of specific domains of knowledge to consider. But Doug particularly wanted to push in what seemed like the most broadly impactful direction—and tackle the problem of commonsense knowledge and commonsense reasoning. And so it was that Doug began what would become a lifelong project to encode as much knowledge as possible in the form of statements of logic.

In 1984 Doug’s project—now named CYC—became a flagship part of MCC (Microelectronics and Computer Technology Corporation) in Austin, TX—an industry-government consortium that had just been created to counter the perceived threat from the Japanese “Fifth Generation Computer Project”, that had shocked the US research establishment by putting immense resources into “solving AI” (and was actually emphasizing many of the same underlying rule-based techniques as Doug). And at MCC Doug had the resources to hire scores of people to embark on what was expected to be a few thousand person-years of effort.

I didn’t hear much about CYC for quite a while, though shortly after Mathematica was released in 1988 Marvin Minsky mused to me about how it seemed like we were doing for math-like knowledge what CYC was hoping to do for commonsense knowledge. I think Marvin wasn’t convinced that Doug had the technical parts of CYC right (and, yes, they weren’t using Marvin’s theories as much as they might). But in those years Marvin seemed to feel that CYC was one of the few AI projects going on that actually made any sense. And indeed in my archives I find a rather charming email from Marvin in 1992, attaching a draft of a science fiction novel (entitled The Turing Option) that he was writing with Harry Harrison, which contained mention of CYC:

June 19, 2024

When Brian and Ben reached the lab, the computer was running
but the tree-robot was folded and motionless. “Robin,
activate.”

“Robin will have to use different concepts of progress for
different kinds of problems. And different kinds of subgoals
for reducing those different kinds of differences.”

“Won’t that require enormous amounts of knowledge?”

“It will indeed—and that’s one reason human education takes
so long. But Robin should already contain a massive amount of
just that kind of information—as part of his CYC-9 knowledge-
base.”

“There now exists a procedural model for the behavior of a
human individual, based on the prototype human described in
section 6.001 of the CYC-9 knowledge base. Now customizing
parameters on the basis of the example person Brian Delaney
described in the employment, health, and security records of
Megalobe Corporation.”

A brief silence ensued. Then the voice continued.

“The Delaney model is judged as incomplete as compared to those
of other persons such as President Abraham Lincoln, who has
3596.6 megabytes of descriptive text, or Commander James
Bond, who has 16.9 megabytes.”

Later, one of the novel’s characters observes: “Even if we started with nothing but the
old Lenat–Haase representation-languages, we’d still be far ahead of what any animal ever evolved.” (Ken Haase was a student of Marvin’s who critiqued and extended Doug’s work on heuristics.)

I was exposed to CYC again in 1996 in connection with a book called HAL’s Legacy—to which both Doug and I contributed—published in honor of the fictional birthday of the AI in the movie 2001. But mostly AI as a whole was in the doldrums, and almost nobody seemed to be taking it seriously. Sometimes I would hear murmurs about CYC, mostly from government and military contacts. Among academics, Doug would occasionally come up, but rather cruelly he was most notable for his name being used for a unit of “bogosity”—the lenat—of which it was said that “Like the farad it is considered far too large a unit for practical use, so bogosity is usually expressed in microlenats”.

Doug Meets Wolfram|Alpha

Many years passed. I certainly hadn’t forgotten Doug, or CYC. And a few times people suggested connecting CYC in some way to our technology. But nothing ever happened. Then in the spring of 2009 we were nearing the first release of Wolfram|Alpha, and it seemed like I finally had something that I might meaningfully be able to talk to Doug about.

I sent a rather tentative email:

Subject: something you might find interesting…
Date: Thu, 05 Mar 2009 11:15:04 -0500
From: Stephen Wolfram
To: Doug Lenat

We’re in the final stages of a rather large project that I think relates to
some of your interests.

I just made a small blog post about it:

http://blog.wolfram.com/2009/03/05/wolframalpha-is-coming/

I’d be pleased to give you a webconference demo if you’re interested.

I hope you’ve been well all these years.

— Stephen

Doug quickly responded:

Subject: Re: something you might find interesting…
Date: Thu, 5 Mar 2009 13:23:31 -0600
From: Doug Lenat
To: Stephen Wolfram

Hi, Stephen.

You have become a master of understatement! This certainly
does relate to the 1000 person-years we’ve spent building Cyc’s ontology,
knowledge base, and inference engines, over the last 25 years. I’d very
much like to see a webconference demo, so we identify the opportunities for
synergy.

Regards
Doug

It was definitely a “you’re on my turf” kind of response. And I wasn’t sure what to expect from Doug. But a few days later we had a long call with Doug and some of the senior members of what was now the Cycorp team. And Doug did something that deeply impressed me. Rather than for example nitpicking that Wolfram|Alpha was “not AI” he basically just said “We’ve been trying to do something like this for years, and now you’ve succeeded”. It was a great—and even inspirational—show of intellectual integrity. And whatever I might think of CYC and Doug’s other work (and I’d never formed a terribly clear opinion), this for me put Doug firmly in the category of people to respect.

Doug wrote a blog post entitled “I was positively impressed with Wolfram Alpha”, and immediately started inviting us to various AI and industry-pooh-bah events to which he was connected.

Doug seemed genuinely pleased that we had made such progress in something so close to his longtime objectives. I talked to him about the comparison between our approaches. He was just working with “pure human-like reasoning”, I said, like one would have had to do in the Middle Ages. But, I said, “In a sense we cheated”. Because we used all the things that got invented in modern times in science and math and so on. If he wanted to work out how some mechanical system would behave, he would have to reason through it: “If you push this down, that pulls up, then this rolls”, etc. But with what we’re doing, we just have to turn everything into math (or something like it), then systematically solve it using equations and so on.

And there was something else too: we weren’t trying to use just logic to represent the world, we were using the full power and richness of computation. In talking about the Solar System, we didn’t just say that “Mars is a planet contained in the Solar System”; we had an algorithm for computing its detailed motion, and so on.

Doug and CYC had also emphasized the scraps of knowledge that seem to appear in our “common sense”. But we were interested in systematic, computable knowledge. We didn’t just want a few scattered “common facts” about animals. We wanted systematic tables of properties of millions of species. And we had very general computational ways to represent things: not just words or tags for things, but systematic ways to capture computational structures, whether they were entities, graphs, formulas, images, time series, or geometrical forms, or whatever.

I think Doug viewed CYC as some kind of formalized idealization of how he imagined human minds work: providing a framework into which a large collection of (fairly undifferentiated) knowledge about the world could be “poured”. At some level it was a very “pure AI” concept: set up a generic brain-like thing, then “it’ll just do the rest”. But Doug still felt that the thing had to operate according to logic, and that what was fed into it also had to consist of knowledge packaged up in the form of logic.

But while Doug’s starting points were AI and logic, mine were something different—in effect computation writ large. I always viewed logic as something not terribly special: a particular formal system that described certain kinds of things, but didn’t have any great generality. To me the truly general concept was computation. And that’s what I’ve always used as my foundation. And it’s what’s now led to the modern Wolfram Language, with its character as a full-scale computational language.

There is a principled foundation. But it’s not logic. It’s something much more general, and structural: arbitrary symbolic expressions and transformations of them. And I’ve spent much of the past forty years building up coherent computational representations of the whole range of concepts and constructs that we encounter in the world and in our thinking about it. The goal is to have a language—in effect, a notation—that can represent things in a precise, computational way. But then to actually have the built-in capability to compute with that representation. Not to figure out how to string together logical statements, but rather to do whatever computation might need to be done to get an answer.

But beyond their technical visions and architectures, there is a certain parallelism between CYC and the Wolfram Language. Both have been huge projects. Both have been in development for more than forty years. And both have been led by a single person all that time. Yes, the Wolfram Language is certainly the larger of the two. But in the spectrum of technical projects, CYC is still a highly exceptional example of longevity and persistence of vision—and a truly impressive achievement.

Later Years

After Wolfram|Alpha came on the scene I started interacting more with Doug, not least because I often came to the SXSW conference in Austin, and would usually make a point of reaching out to Doug when I did. Could CYC use Wolfram|Alpha and the Wolfram Language? Could we somehow usefully connect our technology to CYC?

When I talked to Doug he tended to downplay the commonsense aspects of CYC, instead talking about defense, intelligence analysis, healthcare, etc. applications. He’d enthusiastically tell me about particular kinds of knowledge that had been put into CYC. But time and time again I’d have to tell him that actually we already had systematic data and algorithms in those areas. Often I felt a bit bad about it. It was as if he’d been painstakingly planting crops one by one, and we’d come through with a giant industrial machine.

In 2010 we made a big “Timeline of Systematic Data and the Development of Computable Knowledge” poster—and CYC was on it as one of the six entries that began in the 1980s (alongside, for example, the web). Doug and I continued to talk about somehow working together, but nothing ever happened. One problem was the asymmetry: Doug could play with Wolfram|Alpha and Wolfram Language any time. But I’d never once actually been able to try CYC. Several times Doug had promised API keys, but none had ever materialized.

Eventually Doug said to me: “Look, I’m worried you’re going to think it’s bogus”. And particularly knowing Doug’s history with alleged “bogosity” I tried to assure him my goal wasn’t to judge. Or, as I put it in a 2014 email: “Please don’t worry that we’ll think it’s ‘bogus’. I’m interested in finding the good stuff in what you’ve done, not criticizing its flaws.”

But when I was at SXSW the next year Doug had something else he wanted to show me. It was a math education game. And Doug seemed incredibly excited about its videogame setup, complete with 3D spacecraft scenery. My son Christopher was there and politely asked if this was the default Unity scenery. I kept on saying, “Doug, I’ve seen videogames before; show me the AI!” But Doug didn’t seem interested in that anymore, eventually saying that the game wasn’t using CYC—though did still (somewhat) use “rule-based AI”.

I’d already been talking to Doug, though, about what I saw as being an obvious, powerful application of CYC in the context of Wolfram|Alpha: solving math word problems. Given a problem, say, in the form of equations, we could solve pretty much anything thrown at us. But with a word problem like “If Mary has 7 marbles and 3 fall down a drain, how many does she now have?” we didn’t stand a chance. Because to solve this requires commonsense knowledge of the world, which isn’t what Wolfram|Alpha is about. But it is what CYC is supposed to be about. Sadly, though, despite many reminders, we never got to try this out. (And, yes, we built various simple linguistic templates for this kind of thing into Wolfram|Alpha, and now there are LLMs.)

Independent of anything else, it was impressive that Doug had kept CYC and Cycorp running all those years. But when I saw him in 2015 he was enthusiastically telling me about what I told him seemed to me to be a too-good-to-be-true deal he was making around CYC. A little later there was a strange attempt to sell us the technology of CYC, and I don’t think our teams interacted again after that.

I personally continued to interact with Doug, though. I sent him things I wrote about the formalization of math. He responded pointing me to things he’d done on AM. On the tenth anniversary of Wolfram|Alpha Doug sent me a nice note, offering that “If you want to team up on, e.g., knocking the Winograd sentence pairs out of the park, let me know.” I have to say I wondered what a “Winograd sentence pair” was. It felt like some kind of challenge from an age of AI long past (apparently it has to do with identifying pronoun reference, which of course has become even more difficult in modern English usage).

And as I write this today, I realize a mistake I made back in 2016. I had for years been thinking about what I’ve come to call “symbolic discourse language”—an extension of computational language that can represent “everyday discourse”. And—stimulated by blockchain and the idea of computational contracts—I finally wrote something about this in 2016, and I now realize that I overlooked sending Doug a link to it. Which is a shame, because maybe it would have finally been the thing that got us to connect our systems.

And Now There Are LLMs

Doug was a person who believed in formalism, particularly logic. And I have the impression that he always considered approaches like neural nets not really to have a chance of “solving the problem of AI”. But now we have LLMs. So how do they fit in with things like the ideas of CYC?

One of the surprises of LLMs is that they often seem, in effect, to use logic, even though there’s nothing in their setup that explicitly involves logic. But (as I’ve described elsewhere) I’m pretty sure what’s happened is that LLMs have “discovered” logic much as Aristotle did—by looking at lots of examples of statements people make and identifying patterns in them. And in a similar way LLMs have “discovered” lots of commonsense knowledge, and reasoning. They’re just following patterns they’ve seen, but—probably in effect organized into what I’ve called a “semantic grammar” that determines “laws of semantic motion”—that’s enough to often achieve some fairly impressive commonsense-like results.

I suspect that a great many of the statements that were fed into CYC could now be generated fairly successfully with LLMs. And perhaps one day there’ll be good enough “LLM science” to be able to identify mechanisms behind what LLMs can do in the commonsense arena—and maybe they’ll even look a bit like what’s in CYC, and how it uses logic. But in a sense the very success of LLMs in the commonsense arena strongly suggests that you don’t fundamentally need deep “structured logic” for that. Though, yes, the LLM may be immensely less efficient—and perhaps less reliable—than a direct symbolic approach.

It’s a very different story, by the way, with computational language and computation. LLMs are through and through based on language and patterns to be found through it. But computation—as it can be accessed through structured computational language—is something very different. It’s about processes that are in a sense thoroughly non-human, and that involve much deeper following of general formal rules, as well as much more structured kinds of data, etc. An LLM might be able to do basic logic, as humans have. But it doesn’t stand a chance on things where humans have had to systematically use formal tools that do serious computation. Insofar as LLMs represent “statistical AI”, CYC represents a certain level of “symbolic AI”. But computational language and computation go much further—to a place where LLMs can’t and shouldn’t follow, and should just call them as tools.

Doug always seemed to have a very optimistic view of the promise of AI. In 2013 he wrote to me:

Of course you are coming at this from the opposite end of the Chunnel than
we are, but you’re proceeding, frankly, much more rapidly toward us than we
are toward you. I probably appreciate the significance of what you’ve
accomplished more than almost anyone else: when your and our approaches do
meet up, the combination will be the existence of real AI on Earth. I
think that’s the main motivation in your life, as it is in mine: to live to
see real AI, with the obvious sweeping change in all aspects of life when
there is (i) cradle-to-grave 24×7 Aristotle mentoring and advising for
every human being and, in effect, (ii) a Land of Faerie intelligence
effectively present [e.g., that one can converse with] in every door, floor
tile,…every tangible object above a certain microscopic size.) And to
live to see and be users ourselves in an era of massively amplified human
intelligence …

The last mail I received from Doug was on January 10, 2023—telling me that he thought it was great that I was talking about connecting our tech to ChatGPT. He said, though, that he found it “increasingly worrisome that these models train on CONVINCINGNESS rather than CORRECTNESS”, then gave an example of ChatGPT getting a math word problem wrong.
His email ended:

Yes, let’s chat again at your convenience… it bothers both of us, I
believe, that our systems aren’t leveraging each other! That just bothers
me more and more as I get old (not just older).

Sadly we never did chat again. We now have a team actively working on symbolic discourse language, and just last week I mentioned CYC to them—and lamented that I’d never been able to try it. And then on Friday I heard that Doug had died. A remarkable pioneer of AI who steadfastly pursued his vision over the whole course of his career, and was taken far too soon.

❌
❌