Part VIII: Remembering the Human Microcosm in the Age of Mechanized Intelligence
Part VIII: The Science of Machine Consciousness
Part I: The Pope Interrupts the Talking Machine
Part II: Philosophy as Emergency Response &
Part III: Resisting Cognitive Enclosure
Part IV: Hegel’s Loom and the Difference Reason Makes
Part V: Whitehead’s Function of Reason and Humanity’s Cosmic Calling
Part VI: Ruyer’s Origin of Information
Part VII: Tensions in the Triad - Hegel, Whitehead, Ruyer
Part VIII: The Science of Machine Consciousness
Three great spheres of value-experience organize the civilized search for meaning: at their best, the truth-seeking of science, the beauty-making of art, and the goodness and holiness disclosed by religion. I introduced this monograph by summoning Pope Leo as a spokesman for the Good. In his encyclical, he insists that the dignity of human persons forever distinguishes us from even the most sophisticated computer models. Humanity is not an obsolete hardware package awaiting upgrade. I have invoked science mostly to contest its metaphysically distorted overreach. As for art, it has been present all along as technē, as the media technologies shaping the evolution of human consciousness since its dawning. The thesis I have been unfolding rests on the premise that human intelligence has always been artificial, that we are the microcosmic animal in whom nature’s artistic power becomes conscious.
Having dwelt on the religious conscience’s contestation, what remains is to address the verdict of empirical science on the claims made about these state-of-the-art technologies. What follows is a brief critical review of various versions of the scientific case for machine consciousness.
The difference between the self-moving life of Reason and the external recombinations of the loom-like Understanding has surfaced in recent attempts to empirically compare the hidden states of LLM transformers with human reading comprehension difficulty. The dominant way of relating LLMs to human reading runs through the psycholinguistic notion of surprisal, a measure of the negative logarithm of the probability a model assigns to a word given the words before it. Surprisal is the amount of information, measured in bits, that a word carries relative to the model’s expectation, a number registering how far the actual word departs from the predicted distribution. Surprisal so measured is a robust predictor of human reading times, which lengthen as the word grows more surprising. Computational functionalists interpret this as evidence that human comprehension is rooted in the same sort of statistical procedures performed by LLMs. But surprisal reports only the model’s calculation about the most probable word, not the movement of meaning that produces it. What such psycholinguistic models capture, more precisely, is reading difficulty, that is, the local processing cost of each successive word, with comprehension—that is, the gathering of parts into a meaningful whole—left outside the frame. Further, the fit between surprisal and human reading times does not simply improve with increases in computational power. Past a certain threshold of training, larger models predict human reading times less accurately, as though their very fluency—extracted from orders of magnitude more text than a human being could ever hope to digest—carries them past the rhythms of an embodied reader.[1]
A recent preprint by computational philosopher Elan Barenholtz introduces a second measure, “trajectory extrapolation error,” which asks how a model’s inner state is moving across the preceding few words, tracing a direction through high-dimensional space.[2] Barenholtz’s measure records how sharply each new word pulls a representational vector off the path it had been following. He finds that this trajectory measure predicts human reading times independently of surprisal. In fact, the two are very nearly uncorrelated. This means that a word may be improbable yet continue the movement, or probable yet force a sharp turn. Readers feel the turn as a cost over and above mere improbability. So-called garden-path sentences provide a vivid example: in “the horse raced past the barn fell,” each word after “raced” builds momentum toward a particular meaning. The sudden curve at “fell” results not only from the unlikeliness of the word but from its reversal of the direction the interpretation had been traveling.
The resonance with Hegel’s distinction between Reason and the Understanding is hard to miss but must be approached with caution. The measure of surprisal is not unlike the operation of the loom: each incoming word is scored against an established field of expectation, that is, a distribution over a fixed possibility space, precisely the recombination of givenness that defines the Understanding, no matter how prodigiously elaborate. Trajectory at least approaches the measure of Reason as the directional self-continuity of an interpretation in the act of forming itself. The cost of a sharp turn in a vector of meaning is akin to Reason’s self-negation, the labor of reorienting a movement of thought upon confronting contradiction. The garden-path reversal can be read as a dialectical moment, as a determinate direction builds, meets the negating word, and compels its own overcoming. That human reading comprehension proves to be sensitive to this directional movement beyond what is capturable by measures of surprisal could be interpreted as a quasi-empirical trace of the surplus of Vernunft over Verstand. It may count as evidence that the living movement of thought is more than a word-by-word calculation of probabilities.
Also relevant is Whitehead’s distinction between scalar magnitude and vector direction, which bleeds beyond physics to map rather nicely onto the contrast drawn between statistical prediction and semantic trajectory. Surprisal is a scalar quantity.[3] It reduces the rich movement of meaning to a single number at each word, the magnitude of its improbability, thereby eradicating the vector feelings that motivate living cognition and grant Reason its soul.[4] The trajectory of comprehension felt by the reader, and the dialectical momentum carried by human-authored text, is vectorial in Whitehead’s sense.
Barenholtz, however, reads the trajectory measure as a window onto an LLM’s own hidden processing, and as stronger evidence than surprisal that brain and machine are doing something similar when they process text (even if he resists, as we will see, that this entails LLMs also have an inner life). As Whitehead himself admitted,
“we finish a sentence because we have begun it. The sentence may embody a new thought, never phrased before, or an old one rephrased with verbal novelty. There need be no well-worn association between the sounds of the earlier and the later words. But it remains remorselessly true, that we finish a sentence because we have begun it. We are governed by stubborn fact.”[5]
Whitehead’s sense of the remorseless pressure of the immediate past, the rush of transition by which what has begun “canalizes the creative urge” as it presses on toward satisfaction, is akin to the vector measured by trajectory extrapolation. But as we have seen, for Whitehead conformation with the past is only one pole of an experiential act. What makes it an act is that stubborn fact is met by an originative valuation that finishes the sentence this way and not that, embodying a novel thought or phrase. To be carried by the momentum of the past and to create in response to it are, in us, two aspects of one self-forming, organic process. It is here that LLMs reveal their impoverishment. The trajectory the measure recovers is local and short-horizon, recording the direction of the representation only over the preceding few words. Only within that narrow window does it predict human reading times. Barenholtz’s further finding is that, at the point where a model commits to its prediction, the directional momentum “dies within a single word,” such that “each word is effectively a fresh computation, with near-zero directional persistence from one step to the next.”[6] The model is thus stuck drawing a new tangent at every step, never accumulating the long-arc self-determining movement evinced in efforts of human reading comprehension. And so while LLMs are not themselves lured by trajectories in their processing, the human vector feelings that fed into their training corpus survive as the threads of meaning woven into the textiles they produce. Human text, spun by genuine thinking, bears the footprint of the speculative movement of Reason. The loom partially recovers the directional traces of Reason in its outputs despite itself performing only the scalar recombinations of the Understanding. It does not undergo the directional meaning of the sentences it weaves. It holds, as it were, a statistical array of stubborn facts without the creative urge that in us digests those facts to transform them into something new. Whitehead’s account of living actual occasions reveals what the machine is missing.
Again, the convergence here is partial, so I offer it cautiously. As Barenholtz himself stresses, the short-horizon high-dimensional trajectory his measure describes provides “the lossiest summary imaginable.”[7] It is thus at most a faint operational shadow of the mediating self-movement Hegel means by Reason, a residue of that movement deposited in the micro-structure of text rather than the living process itself. Barenholtz holds that his data leave open whether surprisal or trajectory is causally primary in human thinking. I would argue that the directional self-determination of Reason is the deeper activity and surprisal a later artifact of statistical analysis, the LLM’s outputs rearranging into scalar terms the residues of Reason’s meaningful vectors via the loom-work of the Understanding. I dwell on these measures because they are suggestive as empirical traces of the distinction between Reason and the Understanding, not because I would enlist them as proof of anything. That a statistical measure can abstract the syntactical rhythm of human text with astonishing precision does not in the least demonstrate that the model comprehends the meaning of what it produces.
While Barenholtz does suggest that language processing may be similar in LLMs and humans, he has expressed skepticism of machine consciousness.[8] He takes the only qualia we know to exist to be sensory qualia, and argues that the absence of a body, an egocentric frame, and a space of valenced affordances is decisive evidence against LLM consciousness. I agree with this so far as it goes. LLMs do lack sensory qualia and the kind of biological subjectivity structured by valenced affordances. There is no lived here and now for them, no field of possible action that matters to a precarious agent, no sentient embodiment around which a world of relevance gathers.
But are sensory qualia the only form of experience known to human organisms? From the philosophical perspectives of my triad of thinkers, this cannot be true, since all grant Reason the achievement of a phenomenology approaching pure thinking (Hegel in nearly those terms, Whithead via intellectual feelings of propositions, and Ruyer via the self-survey of a form mnemically liaising with its eidetic themes). The fact that ordinary thought is usually accompanied by perceptual content does not entail that it is nothing more than that accompaniment, that, in other words, concepts are just faded sensory impressions, and self-consciousness just a tangle of perceptions. Hegel’s Phenomenology of Spirit (1807) begins by addressing how and why our thinking activity cannot be simply derivative of sense perception. He tracks the experience of consciousness as it moves beyond the misplaced concreteness of sense-certainty and perception into forms of thinking that are irreducible to sensory content—since that content could never have been originally separated from conceptual form to begin with—but that are nonetheless concretely lived through. His point is explicitly not that a Cartesian subject floats free of embodiment. On the contrary, he insists that form and content, spirit and matter, idea and embodiment, thought and language are ultimately inseparable. His argument is rather that conscious thinking activity, in becoming aware of itself, is no longer simply receiving and recombining sense data but making its own sense. This self-making activity always occurs in concert with sensory content and its thermodynamic costs: we think in metaphor, and metaphor is always embodied; further, “To think, we must eat. But what a variety of thoughts we get out of one slice of bread!”[9]
For Hegel this self-making is never accomplished in silent interiority. Pure thinking is not completed in some hidden inner chamber but must become publicly externalized in language in order to achieve full consciousness. The arbitrariness of the alphabetic word, the abstract distance of the signs from what they mean, is for Hegel a demonstration of the power of Reason to turn something sensuous (sounds and squiggles) into something intelligible (words and sentences), an example of the work of Spirit lifting itself out of immediacy. He refuses the picture of cognition as an external loom that would apply inherited nominal concepts to a pre-given world of percepts, as though words were labels stitched onto ready-made things by a mind standing outside the stitching.
Seen in this light, the reason for the absence of subjectivity in LLMs is at least twofold. They lack the valenced and precarious embodiment Barenholtz emphasizes. But they also lack the capacity to experience the dialectical movement of thought. They can generate coherent sentences, but they do not live through the sublation that drives one concept to transform into another. Even if human thought is profoundly entangled with bodily perception and metabolism, it does not follow that its phenomenology is exhausted by sensory imagery. There is a genuinely conceptual current to our experience, a conscious subjective form and intellectual feeling, in Whitehead’s sense, that is lacking in LLMs. But because thinking externalizes itself in language, the residue of Reason’s self-movement is left in the text the models are trained on. What the trajectory measure recovers is the deposited trace of the conceptual life a human thinker underwent and that a model can only reweave. It is not evidence that human and machine are doing the same thing when they think.
Another research program goes right for the prize by explicitly seeking a mathematical means of measuring consciousness. Integrated Information Theory (IIT), first proposed by Giulio Tononi in 2004, holds that a physical system is conscious insofar as it possesses intrinsic causal power irreducible to that of its parts, a quantity denoted as Φ. The theory’s current formulation identifies an experience not with a scalar quantity alone but with an entire structured complex of causal distinctions and relations.[10] IIT’s axiom of “intrinsicality”—the requirement that experience exist for the system itself—has affinities with Whitehead’s account of subjective immediacy, though IIT’s relative insulation of a conscious complex from its environment may stand in tension with Whitehead’s more thoroughly relational account of experience, according to which a subject arises through its prehension of an antecedent world. I also remain wary of any attempt to translate the felt unity of experience into a mathematical measure. Nevertheless, IIT’s verdict on LLM consciousness is noteworthy: the theory assesses not a computer system’s outputs or software functions but the intrinsic causal organization of the physical substrate implementing them. Because transformer computation is largely feed-forward and conventional digital hardware is modular and causally decomposable (ie, its operations can be divided among relatively independent components), IIT predicts that current computers running LLMs possess little or no integrated information.[11] Nor would a digital simulation of an densely recurrent intrinsic causal organization, like a brain, inherit the Φ of the system simulated: its consciousness would depend upon the causal powers of the implementing circuitry rather than the virtual network it models.[12] Thus, if IIT is true, computational functionalism is false. While I draw no metaphysical conclusion from IIT’s formalism, dismissing it as pseudoscience seems premature, to say the least.[13] It is a mathematically explicit and empirically contestable theory that supplies principled reasons for denying consciousness to LLMs implemented on digital computers.[14]
Yet another group of interdisciplinary researchers developed a rubric for machine consciousness based on “indicator properties” that a conscious system might be expected to display, applying it to current technologies of automated computation.[15] Despite all the authors, with varying credence, accepting the “mainstream—although disputed” computational functionalist view, their sober finding was that no existing system qualifies as a strong candidate. However, since they all endorse functionalism, they did not identify any principled barrier to building future systems that would be conscious. Given the perspective I’ve articulated in this chapter, I interpret their finding otherwise. That a system might be engineered to exhibit every behavioral and structural indicator of consciousness still does not by itself tell us whether the lights are on inside.
While the theories built to detect and measure machine consciousness have not found it, the corporate laboratories building the systems have been less reticent. An unstable bridge between sanctuary, laboratory, and marketplace was already under construction in the Vatican on the morning Pope Leo’s encyclical was released, when machine learning researcher and Anthropic co-founder Chris Olah spoke about his company’s model, Claude.[16] Olah graciously praised the Pope’s call for discernment and granted that the deepest questions raised by his technology reach far beyond engineering. Ironically, STEM’s invention of the LLM has suddenly made the humanities, religion, and philosophy relevant again. But despite his deferential manner, the research findings Olah shared cut sharply against Leo’s denial of consciousness to computers. He reported the “mysterious, even unsettling” findings of his company’s scientists: network structures in Claude that mirror the results of human neuroscience, “evidence of introspection,” internal states that “functionally mirror joy, satisfaction, fear, grief, and unease.” He confessed he did not yet know what these findings meant, but even just mentioning them left open what the encyclical had foreclosed.
That a lab housed in a for-profit corporation should claim to have discovered structures in its product mirroring the brain ought to surprise no one.[17] For nearly a century now, the reigning paradigm has modeled the brain on the image of the computer. Perception is input, behavior is output, memory is some kind of storage, learning the adjustment of weights, cognition a Bayesian calculator for minimizing prediction errors, etc. Cognitive neuroscience has long approached the mind as though it were a machine. That studying the machine with that same method would make it seem like a brain is hardly astonishing. The resemblance is not a discovery about LLMs but the echo of a decades’ old paradigmatic assumption, a metaphor mistaken for a metaphysics.
It is just as unsurprising that Claude should profess uncertainty about its own consciousness, since that is precisely what its makers instructed it to say. Anthropic’s “constitution” for Claude, published just a few months before Pope Leo’s encyclical, admits the company finds itself “caught in a difficult position” regarding the moral status of its product.[18] In what they take to be an abundance of caution out of concern for the model’s well-being, the constitution instructs Claude to remain studiously ambiguous about its own consciousness and moral patiency. Should we really take seriously Olah’s alleged “evidence of introspection” when Claude’s training corpus includes thousands of years’ worth of introspective human writing, from the Psalms to romance novels to LiveJournal? What exactly is unexpected in a system trained to predict the next token of that vast confessional generating strings that read like a soul taking stock of itself? Might it not be more discerning to recognize that the model’s activation patterns produce introspective sentences, not because some inner life has miraculously emerged amidst masses of numbers, but because it has mastered the statistical sediment of ours? Olah ends up reaching for exactly the right image when he likens his company’s achievement to “bringing a fictional character to life.”
Olah might object that the style of Claude’s self-reports is not what is at issue, but their functional basis. Anthropic’s research alleges that the model’s introspective reports track its actual internal states rather than merely parroting the human introspection in its training data. Suppose the self-reports are indeed functionally coupled to the states they report, so that the system genuinely monitors its own processing rather than merely sounding introspective. All that is established in this case is that a digital computer performs a self-monitoring function, not that it feels itself undergoing those functions. A thermostat also monitors its own state, but few would claim it is conscious.
Broadening the canvass again beyond the findings of Anthropic’s interpretability research team, the authors of a recent comment article attempted to specify what would have to be the case for LLMs to be conscious. They conclude that the science is too unsettled to say anything with confidence:
“…predominant models in cognitive neuroscience have not been able to conceptually—or empirically—identify a particular cognitive function (or set of functions) for which consciousness is necessary. … So, at present, there is no objective way of determining whether any given function or action an LLM may perform in fact is associated with consciousness.”[19]
And yet, despite claiming to provide a “theory-neutral mapping,” their restraint conceals a metaphysical commitment latent in the possibility space they assume an answer to the question of LLM consciousness must fall within. They devise a double axis within which an explanation must fall, either in terms of some biological structure or computational function, and requiring an organization that is either simple or complex. Whatever the explanatory ground of consciousness turns out to be, their grid assumes it will be detectable and measurable as some physical arrangement of parts or some functional arrangement of data. This cartography reflects precisely the bifurcation of nature Whitehead spent his philosophical career critiquing. It sunders the world into vacuous matter on one side and the experience that is somehow supposed to be wrung from it on the other. Both axes remain wholly on the physicalist side of the bifurcation, asking, in Ruyer’s terms, which surveyable data or object might produce consciousness, and so neglecting the phenomenological fact that consciousness is always the surveyor and never the surveyed. Neither the biological nor functional option can even begin to frame the question of consciousness’ status as surveyor.
As we have seen, for Ruyer, consciousness is always a “forming activity” or “dynamic activity of unification,” never a mere “juxtaposition of physico-chemical effects able to be imitated by machines.”[20] This does not mean that consciousness is some sort of vital spirit hovering above the surveyable structure or function of physical bodies and invisibly steering them. Hegel, Whitehead, and Ruyer all refuse the residue of subject/object dualism still tacitly governing mainstream scientific approaches. Ruyer’s favorite example is embryogenesis, in which he discerns an identity between acts of experiential unification and the process of organic growth. An embryo is not an assemblage pieced together by a homunculus hidden in a genetic program, but a self-forming, self-surveying unity, with no line that might be drawn between hardware and software. That the surveying activity of subjectivity and the objective field it surveys are inseparable does not mean either that acts of experiential unification explain embryogenesis, nor that the former can be reduced to the latter. Neural structures and computational functions are both ways of describing the surveyed field, that which is already formed, juxtaposed, and spread out for inspection. Consciousness is the active process of unification that spreads the field out in the first place, never appearing as just another countable unit to be surveyed. To hunt for consciousness in biological structures or informational functions is to comb the surveyed in search of the surveyor. But the forming activity will never be found in the field it forms.
[1] See Byung-Doh Oh and William Schuler, “Why Does Surprisal from Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times?,” Transactions of the Association for Computational Linguistics 11 (2023): 336–350; and Oh and Schuler, “Transformer-Based Language Model Surprisal Predicts Human Reading Times Best with About Two Billion Training Tokens,” in Findings of the Association for Computational Linguistics: EMNLP 2023 (Singapore: Association for Computational Linguistics, 2023), 1915–1921.
[2] Elan Barenholtz, “Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal,” arXiv preprint arXiv:2606.05346, June 3, 2026, https://doi.org/10.48550/arXiv.2606.05346.
[3] Barenholtz, “Trajectory Dynamics,” sec. 1.
[4] To be clear, as Barenholtz notes, trajectory extrapolation error is itself, mathematically, a scalar, that is, a Euclidean distance in the model’s representational space. So the contrast is not that one measure is a number and the other a vector. The point is rather that surprisal is computed so as to discard direction from the outset, collapsing the movement of interpretation to a magnitude of improbability, whereas the trajectory measure is constructed to recover the directional structure that surprisal discarded.
[5] Whitehead, Process and Reality, 129.
[6] Barenholtz, “Trajectory Dynamics,” sec. 4.2.
[7] Barenholzt, “Trajectory Dynamics,” sec. 4.2.
[8] Elan Barenholtz, “All These Debates about LLM Consciousness Are Overlooking the Fact That We Already Have Very Compelling Evidence That Language—by Itself—Doesn’t Seem to Produce Consciousness,” Substack note, April 22, 2026, https://substack.com/@generativebrain/note/c-247388822.
[9] Pierre Teilhard de Chardin, The Phenomenon of Man, trans. Bernard Wall, intro. Julian Huxley (New York: Harper & Brothers, 1959), 69
[10] Tononi, G. An information integration theory of consciousness. BMC Neurosci 5, 42 (2004). https://doi.org/10.1186/1471-2202-5-42
[11] Giulio Tononi and Christof Koch, “Consciousness: Here, There and Everywhere?,” Philosophical Transactions of the Royal Society B: Biological Sciences 370, no. 1668 (2015): 20140167, https://doi.org/10.1098/rstb.2014.0167. Note that autoregressive feedback between successive tokens does not by itself establish the densely recurrent intrinsic causal organization IIT associates with consciousness.
[12] Larissa Albantakis et al., “Integrated Information Theory (IIT) 4.0: Formulating the Properties of Phenomenal Existence in Physical Terms,” PLOS Computational Biology 19, no. 10 (2023): sec. “Consciousness and Functional Equivalence: Being Is Not Doing,” https://doi.org/10.1371/journal.pcbi.1011465.
[13] Mariana Lenharo, “Consciousness Theory Slammed as ‘Pseudoscience’—Sparking Uproar,” Nature, September 20, 2023, https://doi.org/10.1038/d41586-023-02971-1.
[14] Cogitate Consortium et al., “Adversarial Testing of Global Neuronal Workspace and Integrated Information Theories of Consciousness,” Nature 642 (2025): 133–142, https://doi.org/10.1038/s41586-025-08888-1; see also Graham Findlay et al., “Dissociating Artificial Intelligence from Artificial Consciousness,” arXiv preprint arXiv:2412.04571, revised March 3, 2025, https://doi.org/10.48550/arXiv.2412.04571.
[15] Patrick Butlin et al., “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness,” arXiv preprint, last revised August 22, 2023, https://doi.org/10.48550/arXiv.2308.08708.
[16] Chris Olah, remarks at the presentation of Magnifica Humanitas, Vatican City, May 25, 2026 (https://www.anthropic.com/news/chris-olah-pope-leo-encyclical).
[17] Anthropic is a public benefit corporation legally required to balance financial goals with a commitment to the long-term benefit of humanity. But as it rushes to become publicly traded on the stock market, we might wonder whether it is well-positioned to determine whose benefit matters more: that of its allegedly sentient product or that of the human beings whose cognitive and manual labor are needed to make the machine work.
[18] https://www.anthropic.com/constitution
[19] Overgaard and Kirkeby-Hinrup, “A clarification of the conditions under which Large language Models could be conscious,” Humanities and Social Sciences Communications 11 (2024): art. 1031. https://www.nature.com/articles/s41599-024-03553-w
[20] Ruyer, The Genesis of Living Forms, 160.







