06The Brain’s Bitter Lesson
“The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.” — Rich Sutton, The Bitter Lesson, 2019
Let’s strip away the biology for a second. If you look at the brain purely as a computer scientist, what is it?
It is a massive graph of densely connected nodes (neurons) firing off noisy electrical signals across multiple time scales. It has extreme redundancy and almost no labels. To decode the brain, you have to extract meaning from a giant, chaotic graph of raw data.
But we have seen this exact architecture before; we even already solved it.
We just did not call it a brain. We called it the internet.
The internet is a vast ocean of unlabeled text, images, connections and noise, far too big for human engineers to manually sort. But we made it tractable, and true to Sutton’s rule, we didn’t do it through clever, hand-crafted tricks.
We did it by scaling an embarrassingly simple technique relentlessly: pointing very large models at very large amounts of data, and letting them figure it out.
The brain is the exact same shape of problem. It is a graph. Messier than the internet, denser than the internet, and profoundly harder to instrument. But the fundamental math is the same. Which means the playbook that conquered the silicon graph should work for the biological one.
This is no longer just a compelling handwavy theory, it is a published, measurable reality.
In early 2025, researchers at Meta’s AI lab released a paper titled Scaling laws for decoding images from brain activityScaling laws for decoding images from brain activity8 datasets, 84 volunteers, 498 hours of brain recording, 2.3 million trials across EEG, MEG, 3T fMRI, 7T fMRI. The first published scaling law for non-invasive neural decoding: log-linear in data, no plateau observed.⁠▸13. They fed an AI decoder nearly 500 hours of brain recordings, comprising 2.3 million trials across EEG, MEG, and fMRI scanners.
What they found is AI’s Bitter Lesson14 playing out in biological real time:
- The Straight Line. Decoding accuracy scales log-linearly with data, with no plateau in sight. Just like a language model, if you feed the decoder more data, it gets predictably better at reading the mind.
- The Hardware Offset. Higher-fidelity sensors (like an fMRI) simply raise the line compared to lower-fidelity sensors (like an EEG). The scaling law holds regardless of the sensor.







- The Data Shape. At fixed total hours, recording one subject for 100 hours drastically outperforms recording 100 subjects for one hour. This is expected: below a certain amount of data per brain, a model cannot yet generalize across people, much like at a small scale a language model would learn far more if you trained it on 100 books in one language rather than one book from 100 languages. With scale, this constraint will disappear.
That is the structural reality that will shape the future of the entire field. The scaling laws do not care about scattered, one-off clinical trials; they demand deep, continuous data. What we currently need for decoding the human mind is not a multimillion-dollar hospital scanner collecting sparse but highly accurate data, what we need is a cheap wearable BCI that can collect data at scale.
We should not wait 70 years to learn the lesson that AI had to learn the hard way. We can skip the hand-crafted era of BCI entirely and jump straight to the scaling playbook.
Part II
Part IIThe ScalingDoes the AI scaling playbook work on the brain?⁠▸ of this essay will focus on exploring how the AI scaling strategy can be applied to non-invasive BCI in order to create a first generation of consumer products and set a powerful flywheel effect
Section 4The Fastest Path to ScalePart I · The Signal: Can the brain be read?⁠▸ in motion. Over the next 5 to 10 years, this flywheel effect will in turn accelerate the momentum of the field significantly and help BCI close the gap with AI.
Whiteboard Skepticism
Of course, this is not a guaranteed victory. Brain data is noisier and more variable across subjects than text. It is time-correlated across multiple scales, and the labels are much weaker than simply predicting the next token. Yes, when scaling laws have been tested at a small scale in BCI (through datasets like ZuCo, DEAP, and AllJoined), they appear to hold, but small-scale trends do not automatically extrapolate.
On top of it, right now, there is no such thing as a decoder that works on a new brain out of the box. Every model built so far still requires a calibration phase specifically tuned to your brain before it actually works. The field has never had to merge tens of thousands of unique human brains into one universal training set without relying on per-subject calibration.
The most counterintuitive aspect of our current ability to decode the mind is that we don’t even fully know how the brain computes. We built the internet, so we understand how it works but we do not know for sure if cognition is reducible to simple neuron electrical impulses, or if sub-neuronal physics actually matters. Microtubules, glial signaling, and even quantum effects are all still live hypotheses as underlying mechanisms for thoughts. All we know is that if we put a sensor on a head, grab a bunch of messy data, and ask an AI to find the pattern, it actually starts reading minds.
The scaling playbook is a high-stakes wager, it is a bet that brute-force compute can extract enough usable structure from the noisy macro-signals of the brain without needing a perfect map of human cognition first. A bet that we can decode the mind without fully understanding it.15 If this turns out to be right, the scaling lines will continue pointing upwards and new breakthroughs will be achieved in fast succession as data and compute scale.
If the last decade of AI taught us anything, it is that betting against compute makes you sound incredibly smart, right up until it does not.
Theoretical counter-arguments can only go so far here. All the caveats above simply tell us that the brain is not exactly the internet, which is already obvious from the premise. The thesis is that it is similar enough that the exact same laws of compute apply. But you fundamentally cannot disprove scaling laws with a whiteboard. It is inherently an empirical problem. The only way to prove that scaling fails is to actually run the experiment and hit a wall.
And so far, every attempt to scale has only reinforced the evidence.
The Bitter LessonThe Bitter LessonTwo pages arguing that across 70 years of AI research, the methods that win are the ones that lean on raw compute and data over hand-crafted human knowledge. Hard-coded cleverness is repeatedly outperformed by general-purpose learning at scale. The frame this entire essay borrows.⁠▸, as Sutton put it (or its older sibling, Peter Norvig’s The Unreasonable Effectiveness of DataThe Unreasonable Effectiveness of DataThe companion case to The Bitter Lesson, made a decade earlier from inside Google: simple models trained on a lot of data consistently trump elaborate models trained on less. The empirical kernel modern AI scaling rests on.⁠▸, which made the exact same point in 2009), will come for every field eventually. It has not yet come for BCI; it is overdue.
Why no one has run the playbook yet
If the scaling bet is so obvious, a glaring question remains. BCI has had decades, why has nobody actually run the playbook?
There are a couple of reasons.
First and foremost, the specific AI architectures and self-supervised techniques required to actually execute this strategy were only proven in the last few years.
But this is not the only reason, this stopped being a blocker a few years ago. What explains the delay in adopting this strategy is institutional inertia. Culture evolves at a glacial pace, or as Max Planck colorfully puts it: “Science advances one funeral at a time”.
The entire BCI field was born in the world of invasive medical devices. Going from a culture shaped by FDA approvals, extreme patient safety and statistical rigor on small datasets to one of brute-force and scale is a brutal paradigm shift. Researchers who spent their careers learning to carefully hand-craft algorithms for specific patients do not naturally reach for the “throw 100K hours at a wall of GPUs” strategy. It certainly has a lot less scientific appeal. This heavy medical baggage was built for brain surgery, but it has completely bled over into the non-invasive space, and it is a real drag on the field.
Furthermore, a lot of the hardware for non-invasive BCI only recently became cheap and powerful enough for this purpose. Historically, the highest-performance, most “serious frontier” BCI results have been invasive or semi-invasive. Which explains why non-invasive hardware took a while to get the attention it deserved. There has also been a lab-fit hardware culture. Research-grade hardware optimizes for signal quality in a controlled setting; scaling-grade hardware optimizes for daily wear at the expense of signal quality.
Lastly, large coordinated data collection efforts only started to come to life recently. There is no Common Crawl16 or ImageNet17 for the human brain. We have several partial equivalents, but they are fragmented, small, clinical, lab-specific, and rarely track the same person over a long period of time. Data fragmentation is a major issue. Each lab collects its own dataset with its own preprocessing, electrode layouts, and approval pipelines. Some initiatives like BIDS and NWB are nascent but adoption is uneven. A paper saying it trained on “eight datasets” usually means a poor graduate student spent six miserable weeks reformatting files before training even began.
A lot of these are coordination problems. Coordination problems get fixed the moment one or two actors finally make the new optimal visible, forcing the rest of the field to follow, just like OpenAI and Google did for AI. The reason the scaling playbook has not run yet is not because it violates the laws of physics. It is because no coordinated push has been done in this direction, but that will change this decade.
The best strategy for the BCI field is not to try to be clever about brain decoding, it is to be willing to scale.
The natural next question is: how do we actually apply scaling in practice. To map the execution out, we are going to use a simple mental model, the BCI pyramid.
07The Pyramid
Before diving into the details of how to scale non-invasive BCI, we need to break it down to its core. Fundamentally, what is the field of non-invasive BCI made of?
The simplest mental model is a pyramid. It has four layers, with every layer built directly upon the foundation of the one below it: Hardware, Data, Algorithms, and Product.
Hardware is the base. Sensors, magnetometers, transducers, hardware sets the ceiling on signal quality. The rule is simple: anything you cannot record, you cannot decode. Right now, this layer is severely underserved. The equipment is often bespoke, expensive, and hacky. The foundation of the entire field is still very underdeveloped.
Data sits directly above that. The corpora of recordings, the protocols, the formatting conventions. Data sets the ceiling on how smart your model can get. You can only perform so much mathematical magic with 100 hours of recordings. Just like the hardware beneath it, this layer is critically immature and fragmented.
Algorithm sits above the data. These are the AI models that turn raw signal into actual meaning. Better algorithms can compensate for noisy data, but only up to a point. Today, a massive amount of the field’s time and talent is wasted here, trying to squeeze every last drop of signal out of tiny datasets.
Product sits at the very top. It is entirely constrained by the three layers below it. Most of the products built in non-invasive BCI today have severely limited capabilities because the foundations they sit on are not robust. Building a consumer BCI right now is like trying to build ChatGPT when you only own a hundred GPUs and ten gigabytes of text. It is a structurally impossible task. The companies doing it are working miracles, but they are fighting a brutal systemic headwind.
The core dynamic here is simple, each layer is bottlenecked by the one below it. You cannot ship a product that requires data your hardware cannot collect. You cannot train a foundation-grade decoder on a couple hours of a graduate student staring at a screen in a basement.
The problem with the field today is that too much attention is at the top of the pyramid and not enough at the bottom. Everyone is obsessing over the penthouse while the foundation is sinking. The hardware layer is starved of well-funded and focused companies. The data layer needs aggressive scaling. And yet, every month brings dozens of beautiful papers proposing new decoders trained on tiny datasets. The algorithm is not the main problem. Data is the problem, and the hardware that collects the data is the problem before that.
Take the ZuCo datasetZuCo, a simultaneous EEG and eye-tracking resource for natural sentence reading12 native English speakers reading natural sentences with simultaneous high-density EEG and eye-tracking. Roughly 70 hours total. The dataset every cross-subject EEG paper has trained against; the field’s existence proof at small scale.⁠▸, every time a research team works a mathematical miracle to squeeze a new paper out of the brain data of 12 people reading for six hours18, they are demonstrating the field’s problem more than they are solving it. Squeezing blood from a stone is impressive, but at a certain point, the miracle just becomes proof of the bottleneck. We need to stop praising the squeezing and start asking why we only have stones.
Now that we have a clear mental model of the field, we are going to dissect exactly what is causing the bottleneck at the top. We will walk up the pyramid layer by layer. First the hardware, then the data, then the AI algorithms, and finally, the complex coordination required across the entire pyramid to actually pull this off.
08Level One: Hardware Is Hard
“People who are really serious about software should make their own hardware.” — Alan Kay, 1982
Let us begin at the bottom of the pyramid: Hardware.
AI hardware has been all over the news recently. We watch trillions of dollars flow into massive GPU data centers because everyone now understands that compute is the ultimate bottleneck for artificial intelligence. But hardware is not just the blocker for AI, it is also one of the key gating factors for brain-computer interfaces, and right now, the BCI field desperately needs its own hardware paradigm shift.
A device engineered to gather massive datasets is a fundamentally different species of machine than a device engineered for a laboratory, and one of the most important differences is portability. If you are optimizing for the absolute cleanest signal possible, you demand precision at all costs. You accept the electromagnetically shielded room, the liquid helium cryogenics, and the completely immobile subject. However, if you are optimizing for the largest dataset possible, the math changes, you still care about signal quality, but the trade-off tilts heavily toward portability. Slightly noisier signal collected at a thousand times the volume beats a clean signal you cannot collect outside a shielded room. That is the trade-off the field has not yet fully internalized.
A slightly noisier signal collected at a thousand times the volume completely beats a pristine signal you cannot collect outside a shielded room
Today, the current state of the art is heavy, expensive, and lives in a room with no windows, and not just metaphorically. That is fine for a paper, but it is incompatible with collecting tens of thousands of hours per subject, across thousands of subjects, in real-world conditions. The field needs a person to be able to put a helmet on in the morning, go about their day, and come home with eight hours of multi-modal neural data perfectly paired with the ambient video and audio.
This does not necessarily imply that all data should be collected in fully uncontrolled environments with hardware that could withstand anything. Reality has a long tail of edge cases. In practice, in the short term, a significant amount of data will still be collected in semi-controlled environments. But even with marginal improvements, more portable hardware unlocks the flexibility to seamlessly record across a wide range of scenarios, that flexibility is exactly what can compound into vast, robust datasets.
Pristine lab data will still matter, doing model fine-tuning and evaluation with high quality, low noise data still provides a significant advantage19, and we will always need those clinical devices. But for the vast bulk of the training, having slightly noisier data at massive scale is much more valuable.
Of course, there is a limit to how much noise a dataset can have while remaining useful, and historically, consumer neural recordings have been 10x or even 100x noisier than their lab counterparts, reducing their practical usefulness for training. However, as hardware is now improving quickly, that limitation is disappearing fast.
Two things follow.
First, the field needs more specialized hardware companies. Companies that focus on the scaling paradigm rather than the highest-signal-quality-at-any-cost paradigm. Hardware iteration is its own brutal discipline, treating it as a side-effect of an algorithm paper has held the field back for two decades. Conway’s Law20 applies here: organizations ship their structure, and an academic field has shipped academic hardware. Thankfully, more specialized hardware companies are finally seeing the light of day, the momentum is already shifting away from signal purity toward scaling.
Second, the right design choices are different when scale is the goal. If portability is the first pillar of scale-fit hardware, multimodality is the second.
Many modalities can be used in tandem. MEG, fNIRS, and EEG measure entirely different physical phenomena; they do not interfere with each other, so you can stack them. When you are aiming for massive, robust datasets, capturing multiple biological signals at a low marginal cost is an obvious choice. The same way that modern LLMs are trained on text, videos and audio, brain foundation models should learn from multiple complementary biological signals. Multimodality should not be just a clever trick, it should be the sensible default.
Additionally, in the past, models and sensors were mostly developed independently, with people making educated guesses about which changes to noninvasive sensors would most improve them. But if we now train large models directly on the sensor’s output, we don’t have to guess anymore, we can co-design sensors and models to create a very tight feedback loop.
Once you accept portability and multimodality as your foundation, the rest of the engineering constraints follow. Building for scale means solving for:
- Manufacturing. How do you actually ship a million units?
- Context. Embedding cameras and microphones to give the neural data real-world grounding.
- Tolerance. Ensuring your sensor array can handle hair, sweat, physical motion, and the chaos of daily life.
- Standardization. One of the most critical and underweighted components. If your hardware outputs data that cannot be cleanly merged with another lab’s recordings, the data layer above it stays hopelessly uncoordinated.
The hardware bottleneck is slowly being overcome, but unevenly.
Each of the major non-invasive modalities (EEG, ultrasound, fNIRS, OPM-MEG, eventually NV-magnetometry) has started making the jump from “lab” to “wearable”. Butterfly Network put ultrasound on a chip, Kernel built a wearable fNIRS rig, Cerca Magnetics ships OPM arrays you wear.
Appendix DThe Modality MapAppendix: Predictions, acknowledgments, resources, and the modality map.⁠▸ describes each one in full.Hardware is also where the geopolitical reality sets in. China has been investing aggressively in BCI hardware over the last several years. Their five-year plan put BCI squarely inside the national strategic technology stack. China does not see BCI as a neuroscience curiosity anymore, it already sees it as a strategic asset.
We have seen this movie before, the AI decade was a fight over chips; the defining geopolitical fact has been who can make them; this has come down to one Taiwanese fab, one Dutch lithography company, a handful of US export controls, and a trillion dollars of strategic anxiety on top. This race has already started for BCI hardware, if brain data becomes strategically important, the machines that collect it will become strategically important too. Within a few years, the TSMCs of brain hardware will start taking form, and China is clearly trying to make sure the answer is not automatically American.
To sum it up, the point of this section is that the hardware unblock matters more than another decoder paper. The next leap in non-invasive BCI will not come from a smarter algorithm trained on a tiny dataset. It will come from sensors that can collect data at scale, sensors which can be worn by many people, for extended periods of time.
But sensors alone only produce raw numbers. Once the hardware to capture the signal at scale is no longer the bottleneck, the next one becomes: how to collect it at scale.
09Level Two: The Data Scaling Argument
“Simple models and a lot of data trump more elaborate models based on less data.” — Halevy, Norvig, Pereira, The Unreasonable Effectiveness of Data, 2009
The second layer of the pyramid is: Data.
For decades, the AI field chased clever symbolic methods, largely because it did not realize that brute force was a real option. The Bitter Lesson highlighted that mistake. The most uncomfortable thing about this outcome is that it is a deeply unsatisfying scientific result: the elegant theory of language did not produce GPT, scaling mindlessly did. That is in part why fields like BCI take so long to internalize the move.
BCI today has a significant advantage over AI in 2017: the AI model already exists. It simply needs to collect the data. The field is starting to awaken to this and is currently crossing into its foundation-model era.21
Decoded · ~10 hours
he saw a dog then looked out the window
Decoded · ~100 hours
I woke from bed and put my face to the glass window
Decoded · ~25,000 hours
I got up from the air mattress and pressed my face against the glass of the bedroom window
As the hardware is improving, scaling data collection becomes easier and datasets are growing, but BCI datasets are still nowhere near the scale that earned foundation models their name. The largest publicly available EEG datasets are tens of subjects, tens of hours each, single modality, mostly in a lab (ZuCoZuCo, a simultaneous EEG and eye-tracking resource for natural sentence reading12 native English speakers reading natural sentences with simultaneous high-density EEG and eye-tracking. Roughly 70 hours total.⁠▸, DEAPDEAP: A Database for Emotion Analysis using Physiological Signals32 subjects watched 40 one-minute music videos with EEG and peripheral signals recorded, rated for valence, arousal, dominance, and liking. ~21 hours; the de facto benchmark for emotion decoding from EEG.⁠▸, the Stanford handwriting corpusHigh-performance brain-to-text communication via handwritingFirst demonstration that thought-to-text could clear 90 characters per minute, decoded from intracortical recordings while a paralysed subject imagined writing letters by hand. The handwriting-imagery substrate the subsequent generation of speech-decoding papers built on.⁠▸, the BCI Competition datasetsBCI Competition IV datasetsThe community benchmark series for motor-imagery and ERP decoding. Small N, well-curated, used by hundreds of papers as the apples-to-apples comparison baseline. Total volume under 1,000 hours across the series.⁠▸). The most ambitious in-house corpora I am aware of are still under 10K hours total, and single modality.
However, that picture is starting to change. NeuralBench v1.0NeuralBench: A Unifying Framework to Benchmark NeuroAI ModelsConsolidates 94 datasets, 9,478 subjects, 13,603 hours of EEG; the first serious shared evaluation substrate for non-invasive neural decoding.⁠▸ (Meta, May 2026) consolidates 13,603 hours of EEG across 9,478 subjects from 94 datasets. Models are scaling alongside: NeuroLM (Jiang et al., ICLR 2025)NeuroLM: A universal multi-task foundation model for bridging brain signalsPretrained on 25,000 hours of EEG at 1.7B parameters; current open-source record for brain foundation models.⁠▸ at 1.7B parameters on 25,000 hours of EEG; LaBraM (Jiang et al., ICLR 2024)Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIEEG foundation model trained on 2,500 hours of EEG across 369M parameters.⁠▸ at 369M and 2,500 hours.
For comparison: GPT-3 (2020) was trained on roughly 300 billion tokens. Llama 3 (2024) used 15 trillion; Llama 4 (April 2025), 30 trillion. The effective stock of public text caps at roughly 300 trillion tokensWill we run out of data? Limits of LLM scaling based on human-generated dataEffective stock of quality-adjusted public text is ~300T tokens; full saturation between 2026 and 2032 at current growth rates.⁠▸. BCI has still a long way to go to be in AI scale territory.
If you are inside the BCI field, this data scaling argument may sound obvious. To a meaningful chunk of your colleagues, it is not. The field’s culture has been “do clever things on small datasets, publish, repeat” for a long time. The move to “build the boring foundation first, and scale until the Bitter Lesson kicks in” can read as a step away from real science, it can feel less beautiful or clean, but it is also the move that worked everywhere else and that the BCI field needs.
The data collection argument compounds in one more way. The massive BCI datasets produced by following this playbook are also very valuable to AI labs. Indeed, as we get closer to the data wall and alignment becomes a higher priority, AI labs are becoming more and more interested in alternative data sources. Therefore, a large dataset of human cognitive processes is both valuable to the BCI field as well as the AI field. The thesis is not just that BCI scales, it is that BCI scales and AI labs scale toward each other, following a self-reinforcing cycle. We will discuss this point more in depth in Part IIIPart IIIThe CollisionWhy does BCI matter for AI?⁠▸.
As datasets grow larger, training large models to decode the brain becomes more feasible. But to actually extract meaning effectively from this scale of data, the field has to abandon its custom algorithms and focus on the architectures and strategies that conquered language.
10Level Three: Foundation Models for the Brain
Data without models is not of much use, and so the next layer of the pyramid is the AI layer trained on the neural dataset.
The massive advantage BCI has today is that we do not need to invent the math. The AI playbook is already written. Little creativity is required on the architecture side, you do not need to reinvent the wheel, you take the exact transformer models that mastered language, tweak the ingestion layer to accept neural data, and hit run.22
Another page we must steal from the AI playbook is to build the foundation in the open. Right now, the BCI field lacks the data to even test the upper limits of a non-invasive brain decoder through scaling. But trying to hit that scale in a silo would be unwise, the way the AI field tackled this situation in its early days is through open collaboration. The foundational work at OpenAI, DeepMind, and FAIR was strikingly collaborative. The datasets were shared, the papers were openly published, and the models were generally released. The AI field eventually moved toward closed weights, but only after the foundation was laid. BCI is still firmly in the foundations-laying phase. Starting closed is bound to slow progress.
A small set of startups and academics have begun pushing in this direction23, the ecosystem is growing. Because the AI field is already so mature, the models themselves are not the blocker right now. What it lacks is the massive corpus to train them at frontier scale and the empirical proof that frontier scale produces the same capability uplift for brains that it did for text.
There is also a compute tailwind that benefits BCI specifically. GPU performance per dollar has been roughly doubling every two years, and the cost of training any given model capability has collapsed by orders of magnitude in five years. A GPT-3-equivalent run that cost millions in 2020 now sits in the low six figures and keeps falling. BCI does not need to fund frontier-AI-scale compute today. By the time a brain corpus justifies frontier compute, the compute will cost a fraction of what frontier labs pay now. The hard problem of this decade is the data side; the compute side gets cheaper simply by waiting.
We are now one Common Crawl and one foundation model away from BCI’s ChatGPT moment.
However, the biggest challenge in running the strategy that I have described is not the individual layers of the pyramid, it is the coordination across the field. As is often the case with deep tech, the hardest engineering problem is usually not the silicon or the math, it is human coordination.
11The Coordination Conundrum
Coordinated scaling is something BCI has not yet seriously attempted. It runs through the three independent layers of the BCI pyramid (hardware, data and models). Without it, every team is solving the same problem from scratch.
To understand why coordination is so important, we can have a look at the data itself. Brain data is inherently more complex than the text data used to train modern LLMs. Text is just discrete, standardized words whereas brain data is continuous, messy physics; it varies wildly depending on the hardware, the environment, and the human wearing the device.
Because of this inherent complexity, most labs structure their data completely differently. There are attempts at standard formats (like BIDS (Brain Imaging Data Structure) and NWB (Neurodata Without Borders)), but their coverage is partial; outside their reach, formats range from MATLAB MAT files to custom HDF5 to bespoke CSV exports. The cost of this fragmentation is invisible to outsiders but painful to insiders: a paper trained on “eight datasets” usually means a poor soul spent six grueling weeks just reformatting files before training even began. Multiply that friction across every lab in the world and consolidating a unified dataset becomes a real challenge.
The AI data revolution did not happen by accident. Common Crawl, ImageNet, and a handful of other artifacts standardized data collection, and the field compounded on top of them. This meant researchers could stop wasting time collecting images and start competing on algorithms. ImageNet alone seeded a decade of computer-vision progress because everyone trained on the same images and reported the same numbers. BCI desperately needs its ImageNet. But because brain data is so much harder to standardize than text or photos (and this is where my earlier analogy comparing the brain and the internet graph bends), establishing these conventions is going to be ten times harder, and ten times more important.24
Underneath the file-format problem sits a deeper structural one. The traditional neuroscience lab is highly siloed. A lab buys a few commercial sensors, runs a highly controlled experiment on 20 people, cleans the data just enough to get a result, and publishes a paper. That structure is fantastic for publishing niche scientific insights. But it is completely useless for building industrialized datasets of billions of hours. Academia simply is not built to be a massive data factory. Furthermore, the academic incentive structure still rewards researchers for building clever algorithms on tiny datasets, not for doing the boring, unglamorous work of standardizing massive ones. Expecting this system to spontaneously generate a unified, global data standard is like trying to herd academic cats. Tenured principal investigators have spent their entire careers being rewarded for being unique and clever, not for acting like a compliant, standardized cog in a massive industrial data factory. This is in some ways a brain-coded Tragedy of the Commons.
Coordination is not just a data-layer concern. Hardware standardization matters too: if your device’s output cannot be cleanly merged with another lab’s recordings, the data layer above stays uncoordinated. Likewise, open foundation models are coordination artifacts in their own right; they let the field share the cost of training one decoder rather than every lab training their own. Each of the three layers has a coordination dimension.
And because the field is fragmented, the three layers of the pyramid, hardware, data, and modelsThe BCI pyramid
⁠▸, are stuck in a classic cold-start trap.
- Hardware companies cannot justify scale-out until the data layer creates demand for it.
- Data collectors cannot gather massive volume until hardware becomes cheap and portable enough, and before BCI demand for data increases.
- Model builders cannot justify frontier compute costs until the massive datasets exist to train on.
Each layer is rationally waiting for the layer below to move first. As a result, the whole stack stays stuck at lab scale by default. The way out is for every layer to start building for the anticipated scale the layer above will eventually require, not the scale current demand rewards. That is the coordinated strategy transition this decade needs to unstick the cold start.
The underlying technology is finally ready. The remaining challenge is to coordinate scaling as BCI transitions from a scientific problem into an engineering problem. To succeed, non-invasive BCI now requires founder energy, massive capital, and a tolerance for brute-force engineering that the academic system simply was not designed to deliver.
Now, if we assume the field successfully executes this playbook, where would that actually lead us? And what exactly are we building?
12The Two Ladders
If the playbook works and compounds, how do we actually measure success? To answer this question, we will use two distinct “ladders of progress”. The first measures what a device can actually read. The second measures how much effort a user has to invest before the device works for them.
We will call them the capability ladder and the calibration ladder. They represent the two main dimensions which define BCI progress.
The Capability Ladder
The capability ladder represents the raw decoding power of the BCI: how detailed the extracted information is, and how reliably the device can read it. Each rung on the capability ladder corresponds to roughly one order of magnitude better signal quality (driven by better hardware, more data, larger models), which in turn unlocks a new class of product that was previously impossible.
Non-invasive hardware fundamentally has a lower ceiling on this ladder than invasive surgery. At the limit, a sensor outside the brain will never be able to beat one within the brain. A wearable sensor will never read a single, isolated neuron the way a surgically implanted electrode array does. But the crucial reality of the BCI field today is that no one actually knows where that non-invasive ceiling is, and we have barely begun to climb the ladder.
These ladders illustrate the technological milestones BCI can achieve as it matures:
Emotional and physiological biomarkers. Your device could record and monitor your emotional state continuously. Knowing when you are happy, sad, angry, stressed, etc.… We are already scratching the surface of this today.
Attention and focus tracking. The device knows when you are in deep thoughts versus when your mind has wandered. Tracking focus could help improve the ability to do deep work and has numerous health applications. This is also starting to become possible.
Sub-vocal text. This is the rung that changes daily life. Almost-mouth a word, and watch it appear on the screen before your jaw has finished moving with a lag below the threshold of human notice. This allows people to work hours without touching the keyboard once. We already have early lab instances of this technology.
Once you reach higher rungs, things get stranger.
Silent inner-speech recognition. The words you are thinking, decoded without any muscle activation at all. You compose a message in a meeting and nobody around you sees a single muscle move. There are very early, successful lab demonstrations of inner-speech decoding.
Visual concept reconstruction. Hold an image in your mind’s eye and your phone shows you a still of it. You can sketch by remembering. Surprisingly, there are some early successes here too; this is less sci-fi than it might seem.
Multi-modal imagination decoding. Not just a still image anymore, the full imagined scene with sound and motion. What a memory feels like to you could be exported with reasonable fidelity to a file that other people can open.
Episodic memory replay. A recording, in your own native cognitive format, of an hour from last Saturday, which could be played back by you as cleanly as you played it back from inside your own head.
Continuous semantic transcription. A real-time transcript of your thoughts. Or in other words: mind reading. The moment this becomes technically possible, every privacy question surrounding this topic will come forward at once. Before this rung can ever reach a consumer, the field has to design an entirely new kind of social firewall, because navigating the societal implications of mind reading will be vastly harder than writing the algorithm itself at this stage.
The top of the ladder is something we do not yet have a clean word for: a full real-time bidirectional channel between human cognition and machine reasoning, where the boundary between “what I thought” and “what we computed together” stops being legible. This is what some call the Merge.
The point of the ladder is to illustrate how far a non-invasive BCI could get us and what it could unlock. You only need the bottom rung to ship at consumer quality to have the rest follow through the flywheel effect
Section 4The Fastest Path to ScalePart I · The Signal: Can the brain be read?⁠▸. The first wave of non-invasive BCI will likely look like something modestly useful, probably wellness wearables, and then from there, every generation will have exponential gains in capabilities. We are now standing at the edge of the first wave.
The Calibration Ladder
But there is a hidden catch. The capability ladder assumes the device already understands you.
In reality, brains are like fingerprints, no two are wired exactly the same. That means every single brain has to be decoded differently. Right now, a BCI is like a custom-tailored suit; it only fits the exact person it was measured for. Before a device can decode your thoughts, there is a calibration period where the model has to learn your specific neural “accent”.
This added friction is the single biggest bottleneck to building a consumer device. In a lab, a graduate student might be willing to patiently train a decoder for six hours but, in the real world, a consumer has zero patience. They will throw the device in a drawer if it does not work almost immediately.
That brings us to the second ladder, running along the capability one, and arguably more important for having non-invasive BCI used outside of the lab. I call it the calibration ladder; how much effort each new user has to invest before the device works for them.
The capability ladder is what the device can read; the calibration ladder is how much friction the read has.
The expectation is that pretraining across many brains collapses the marginal cost of decoding a new brain, exactly the way zero-shot learning solved this exact problem for language models.
The strategy described in this essay is to scale non-invasive BCI, using the same scaling strategies that AI has used over the past decade, to build a first generation of BCI consumer products in order to ignite a flywheel effect which will in turn accelerate the BCI field as a whole.
The biggest unaddressed remaining unknown is whether non-invasive BCI has a high enough ceiling on the capability and calibration ladders to achieve this goal. The honest answer is that the entire field is currently staring into a fog of war.
13The Unknown Ceiling
“We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run.” — Roy Amara (Amara’s Law)
How far does this playbook actually take us? With what timeline? And will the physics even allow it? Well, nobody actually knows.
The historical track record of high-confidence BCI predictions is mostly a graveyard of wrong timelines. I am offering you mine anyway.
The bear case is that the skull, the scalp, and the Earth’s magnetic field combine into a noise floor that may cap how much signal we can pull from outside the head. If that worst case turns out to be true, non-invasive caps out at “modestly useful for medical contexts”, and the only path to high-bandwidth thought decoding is invasive surgery.
But here is the reality on the ground: we are nowhere near hitting that wall in actual experiments. The biggest non-invasive datasets are still single-modality, with under 10K hours on just a few dozen subjects, no theoretical argument can override this empirical fact. All the current empirical experiments we have point in the exact same direction: straight up the log-linear scaling law curve, the only way to know where the summit lies is by climbing towards it.
Earlier in this essay, if you remember the three S-curvesThree S-curves of capability
⁠▸, we illustrated the theoretical upper bounds of what each modality could achieve. This was just an illustration; in practice, we are standing at the absolute bottom of the curve, looking up into the dark.
For this scaling thesis to work and truly compound, the hardware must clear the friction bar for everyday consumer adoption, and we already have early evidence that this is happening. Just as we saw with language models, BCI scaling laws are already becoming visible in the lab, and more predictable.
If we assume the trajectory holds. In the best-case scenario, within three to six years we will see the first ten million non-invasive brain interfaces deployed in the wild. Within six years, they will evolve from early-adopter tech into a ubiquitous layer of everyday computing. The scaling laws will feed on this massive new data stream, creating a flywheel that makes the field exponentially more attractive to capital and talent.
So where does that actually leave us on our two ladders? How high do we expect to be able to climb before hitting the ceiling?
- The Capability Ladder: Inner intent, inner emotion, silent speech recognition, and visual concept reconstruction are all high-probability outcomes without surgery. There are already early signals of this working today. Full multi-modal imagination decoding (exporting a memory with sound and motion) has a low but non-zero probability non-invasively.
- The Calibration Ladder: Thanks to foundation models trained on population-scale data, the calibration tax will collapse. Just as modern speech AI instantly understands your voice without calibration the first time it hears you, BCI foundation models will recognize your neural patterns with minimal calibration because they have already seen thousands of other brains. We will reach a point where a new device takes only minutes to adapt to a new brain.
This manifesto relies on the conviction that the ceiling reachable with non-invasive BCI is much further away than current consensus assumes. Since theory won’t settle it, the only way to verify this conviction is to actually run the experiment.
But running the AI playbook requires a massive amount of data, and while for the first few orders of magnitude we can collect the data in the lab, at a certain scale, it becomes intractable. It will only scale if millions of people actually wear the hardware. To trigger this rapid scaling loop, the field has to step out of the lab and answer a very practical question: what actually convinces the first ten million everyday consumers to put a brain-computer interface on their heads?
14The Consumer Wedge
“Any sufficiently advanced technology is indistinguishable from magic.” — Arthur C. Clarke
We have spent a lot of time on theory, but it helps to ground our thoughts in reality. And even if we alluded to it, we have not yet clearly illustrated what the first wave of non-invasive BCI products could actually look like.
In Section 9
Section 9Level Two: The Data Scaling ArgumentPart II · The Scaling: Does the AI scaling playbook work on the brain?⁠▸, we stared down a four-order-of-magnitude gap between the data AI uses for frontier models and the data BCI currently has. To close this gap, we need orders of magnitude more users, millions of people using BCI on a regular basis.
This raises an obvious question. Who on earth are these 10 million people, and what exactly are they putting on their heads?
Normal people do not buy abstract scientific milestones, they do not care about “log-linear scaling laws” or “foundation models”, they buy a product because it solves a specific problem or feels like magic.
Big tech is already positioning itself for this transition. In 2025 Apple introduced the “Human Interface Device” protocol which allows BCI to register as first-class input on the iPhone.
The wedge that gets 10 million devices onto 10 million heads is probably not going to be one single, monolithic killer app. It will be a cluster of “first-good-apps”, products that offer enough immediate, undeniable utility to overcome the physical friction of strapping a new piece of hardware to your skull. Early adopters might tolerate a clunky wearable just because of the sci-fi novelty, but to hit true consumer scale, the utility still has to be obvious.
Two primary breakout products look most plausible to cross the chasm first.
The first is frictionless communication, aka telepathy, talking to machines without using your thumbs or your vocal cords. It would unlock sub-vocal text and intent decoding for smart glasses, coding companions, and AI assistants; making it faster than touch, quieter than voice, and ambient by default. It also provides hands-free augmented reality control (like the Meta Neural Band shape) and forms a critical accessibility bridge for users who cannot speak or type comfortably.
The advantage comes from bandwidth and seamlessness, reading directly from the brain allows us to capture much higher-fidelity intent than a keyboard or microphone ever could. In the beginning, the raw read will not be perfectly accurate, but it does not have to be. Smart product design will mask the early hardware limitations, and flawless, paragraph-length thought-to-text will come later; the first wave will likely rely on high-level intent reading and coarse-grain indicators. As we pair coarse telepathy with an AI that is smart enough to infer the rest of your context, it will start feeling like magic25.
The second is wellness and recovery, the Whoop for the brain. The longevity movement brings intense consumer appetite for neural interfaces. Users have already adopted rings and wristbands to track their physical recovery; a neural wearable provides the missing supervision layer by tracking the central nervous system readiness, emotional states and passive continuous stress tracking. This enables numerous other adjacent applications requiring visibility into the cognitive state such as professional fatigue auditing in high-stakes environments like surgery or air traffic control.
Other plausible first good applications of non-invasive BCI include:
- Brain-Aware Environments: Software that is no longer blind to your internal state. An empathetic AI assistant can ingest your cognitive load as part of its prompt context, giving shorter, punchier answers when you are exhausted, providing deeper analogies when it senses you are confused, or entirely suppressing notifications when you are in a flow state. Similarly, study software can track attention and focus, adapting a lesson plan the moment your mind wanders to lunch.
- The Inner Space Stack: Tools designed for exploring and optimizing human consciousness. This includes closed-loop meditation apps that track a genuine neural “depth score” and ring a gentle mindfulness bell when focus breaks; sleep and dream products that track REM cycles for lucid-dreaming cues or trigger smart alarms at your lightest sleep moment26; and therapeutic tools for optimizing mental health and clinical depression.
- The Creative Engine: Thought-to-image generation as a professional creative tool. Instead of typing complex text prompts into an image generator, a designer holds a visual concept in their mind’s eye and watches the system render a rough draft on screen. This translates to music and movie creation as well, any medium that captures a mental state and tries to translate it into creative software could benefit from this.
Once 10 million devices are used in the wild everything accelerates, manufacturing costs collapse, the sensor stack becomes commoditized, data becomes plentiful and the foundation decoder models become shared public goods.
This consumer flywheel strategy can shorten the timeline for progress significantly. By shifting the field out of the slow, methodical timeline of academic tenure tracks and clinical trials, non-invasive BCI gains a massive structural advantage, it allows data to accumulate at scale. Given the staggering velocity of frontier AI development, this consumer-led explosion is the only path moving fast enough for BCI to have a shot at keeping up with frontier AI, and for humans to remain in the loop.
If BCI, and neurotechnology overall, is going to play a role in keeping humans in the loop during the transition to superintelligence, it has to catch up, and it has to do it fast. It certainly cannot move at the agonizingly slow pace of invasive surgery, which is why the non-invasive consumer loop is the only vector fast enough to outrun an impending AI takeoff within the narrow window left to pull it off.
15Phase Lock
“The hope is that, in not too many years, human brains and computing machines will be coupled together very tightly, and that the resulting partnership will think as no human brain has ever thought …” — J. C. R. Licklider, Man-Computer Symbiosis, 1960
AGI is coming soon. At this point it is no longer a distant thought experiment, we are all expecting the future to arrive early, and that accelerated timeline raises a foundational question: how do we actually keep humans in the loop?
Frontier AI is moving ever faster, with flagship models now released monthly, sometimes faster. The window of opportunity to shape what happens next is terrifyingly narrow. These are the last few years of the world as we know it. Metaculus, a prediction website with an impressive track record, currently estimates that weak AGI will be reached by mid 2028. That leaves us with a couple of years, at most, before we fully cross the intelligence threshold and enter the New World.
The critical decisions, what these models learn to optimize, whose preferences carry through, and how they ultimately interface with humanity, are being locked in right now. If our goal is to maintain agency and actually keep humans in the loop as intelligence keeps improving, human-machine interfaces need to arrive within the next few years.
I call this brief interval between humanity recognizing that AI will surpass us and the moment it does the Phase Lock Window: the narrow span in which we still have the agency to shape the relationship between human and machine intelligence before it hardens into something far more difficult to change. The Phase Lock Window is roughly between now and the mid-2030s.
To be clear, BCI will not magically solve every problem raised by AGI. There is a mountain of socio-economic, philosophical, and technical alignment problems we also have to figure out; and the technology, as Section 22Section 22Mind Reading at Scale: What Could Possibly Go WrongPart IV · The Aftermath: What are the implications for the future?⁠▸ lays out, is itself very dangerous. However, BCI, and neurotechnology, are the only structural fix to the human-machine communication and cognition bottleneck, the sole solution expanding human capabilities rather than constraining machines. This makes neurotechnology one of the strongest long term levers for AI alignment. But it can only play that role if it scales in time, we only have a few years to make BCI part of the solution.
That urgency is why this manifesto proposes a concrete strategy for accelerating the field through non-invasive BCI. If establishing a high bandwidth brain-computer interface requires invasive surgery, it will not land this decade. And by the time it does, the transition to artificial superintelligence will have passed, and most of the questions worth asking will have already been answered using whatever limited tools we had on hand. Because the surgical route cannot close the data gap inside the Phase Lock Window, the explosive adoption of non-invasive hardware is our only remaining vector to integrate biology into the alignment toolbox before the future is permanently locked in.
When mapped against the arrival of superintelligence, this consumer-led BCI explosion is not just a cool product milestone, it is more akin to an urgent rescue mission
Blueprint and Convergence
Throughout Part II
Part IIThe ScalingDoes the AI scaling playbook work on the brain?⁠▸, we have asked one core question: Does the AI scaling playbook work for the brain?
The answer, as far as the evidence can take us for non-invasive BCI, is yes.
When you strip away the messy biology and treat the brain as a computer science problem, the similarities to the early days of the AI era are flagrant. The data speaks for itself; BCI has the same footprint as the early deep learning boom, with an eight-year lag (see the eight-year-offset chartAI vs BCI, the eight-year offset
⁠▸).
We now have a complete, cohesive blueprint for scaling non-invasive BCI. The final takeaway of Part II
Part IIThe ScalingDoes the AI scaling playbook work on the brain?⁠▸ is that we are standing in a short, weird and fragile window where BCI and neurotechnology can still matter for shaping the future of humanity, but this Phase Lock Window is remarkably narrow.
Yet scaling AI and BCI in isolation tells only half the story. To understand why BCI and neurotechnology as a whole become more important as AI grows more capable, we must examine how the two technologies converge. Neurotechnology inspired modern AI; AI now accelerates neurotechnology; and as BCI matures, it will become foundational to AI once again. BCI is about to become the infrastructure we use to both build and operate frontier models, providing the multimodal neural data needed to inform the training and alignment of the next-generation of AI models, and creating a high-bandwidth interface to control them directly through thought rather than words. Crucially, this bridge is one of the few alignment levers capable of scaling alongside AI itself.
AI and BCI are now on a direct and imminent collision course.
Zooming out, this is where we stand.
In Part I
Part IThe SignalCan the brain be read?⁠▸, we asked ourselves: Can the brain actually be read?
In Part II
Part IIThe ScalingDoes the AI scaling playbook work on the brain?⁠▸ we took the AI playbook and asked: what happens when you run it on the brain? We explored how applying AI’s scaling strategy to non-invasive BCI could unlock the first generation of BCI consumer products.
To answer that, we built a four-layer pyramid, hardware at the base, then data, then foundation models, then product at the top, and climbed it from the ground up. The hardware layer has to shift from lab-fit to scale-fit. The data layer has to close a four-order-of-magnitude gap with frontier AI corpora. The foundation-model layer has to inherit the open-collaboration and scaling culture that made language models take off.
Then we stepped back from the layers and looked at the system as a whole. We discussed the coordination problem that has kept the field stuck, the cold-start trap where every layer is rationally waiting for the one below to move first. We mapped the two ladders through which to measure progress: the capability ladder (what a device can read) and the calibration ladder (how cheap it is to make it work for a new brain). We addressed the bear case, that the skull might cap the signal earlier than we think, and concluded that no whiteboard argument can settle the question; the only way to find the ceiling of non-invasive BCI is to empirically test its limits. And finally, we looked into the two consumer wedges, frictionless communication and continuous wellness tracking, that have the best chance of sneaking this hardware out of the lab and onto ten million heads.
Along the way, I have made three fundamental assumptions:
- Non-invasive brain signals carry enough information. The brain needs to emit enough information for a sensor outside of the skull to be able to decode it. Early results suggest they do.
- Scaling laws hold for brain data. More data must continue improving performance. The exponential curve reported by Banville et al.Scaling laws for decoding images from brain activity8 datasets, 84 volunteers, 498 hours across EEG, MEG, 3T fMRI, and 7T fMRI. The first published scaling law for non-invasive neural decoding: log-linear in data, no plateau observed.⁠▸ must persist at 100K-hour and 1M-hour corpora rather than plateau early.
- Brain foundation models generalize across people. Models trained across enough subjects must reduce per-user calibration from hours to minutes or eliminate it entirely.
If all three hold, telepathic brain-machine communication will become a consumer reality this decade, setting in motion a powerful flywheel that accelerates the field and helps neurotechnology close the gap with AI.
13. Despite its recent missteps in the metaverse and AI, Meta has been impressively ahead of the curve in neurotechnology, as its latest research demonstrates.
14. The Bitter Lesson, articulated by Rich Sutton in 2019, is that general methods that scale with data and computation eventually outperform handcrafted human knowledge. It is “bitter” because researchers repeatedly resist the discovery that their clever domain expertise matters less than brute-force scale.
15. There is a deeper question lurking here: whether “understanding the brain” is even a well-defined milestone. I believe it is not: understanding is compression, and there is only so far a system as complex as the brain will compress. Tal Yarkoni makes the fuller case, and argues we might not even notice reaching it, in If we already understood the brain, would we even know it?
16. Common Crawl is a non-profit that has been continuously crawling and publishing the web since 2008; the resulting petabyte-scale dataset is what most large language models trained on.
17. ImageNet is a labeled dataset of ~14 million images across 20K+ categories. The 2012 ImageNet competition is widely credited as the beginning of the deep-learning era in computer vision.
18. Yes, I promise you this is not a joke. One of the most commonly used datasets in the entire field literally consists of 12 people recorded for six hours each. That is 72 total hours of data. We are trying to build the ultimate foundation model for the human brain using the equivalent of one long weekend of recordings. It is like teaching an AI Japanese by showing it one sushi menu.
19. The best strategy for training large brain models probably mirrors the one used in autonomous driving: train the base model on abundant but imperfect data (in their case simulated data), then fine-tune and validate it on smaller volumes of high-quality data (in their case real-world data). For BCIs, that would translate to pretraining on noisy consumer recordings, then fine-tuning and evaluating on cleaner lab data; there is already evidence that this works.
20. Conway’s Law: software systems mirror the communication structures of the organizations that build them. Coined by Melvin Conway in 1968.
21. I have been circling this idea for a while, it just took my distracted self a while to actually get it on paper. Here are some receipts from 2024: “We need a new Moore’s law for brain mapping. Full brain mapping will arrive sooner than we think.” and “Direct brain to LLM communication probably coming sooner than we think.”
22. It will obviously be more complex than this but the point is that no breakthrough is required, mostly engineering.
23. The broader open ecosystem has produced roughly a dozen brain foundation models since 2023: BrainLM, LaBraM, NeuroLM, REVE, POYO+, NDT3, LUNA, BENDR, BIOT, CBraMod, EEG-GPT, Neuro-GPT among them.
24. An early example of the kind of standardization this argument calls for: Meta AI’s NeuralBench, released in May 2026, is an open-source benchmarking framework that unifies 94 brain-data datasets across 14 deep-learning architectures, with initial focus on EEG and extensions to MEG and fMRI. One of the first attempts at a shared evaluation dataset for the field.
25. Once embedded into consumer devices, BCI will also come with numerous new UX questions: When should the system act on a thought, and how fast? Which thoughts matter, and which are just passing cognition? How to build a thought native UI interface? How to translate different mental mechanisms across people into actions? Overall, this will demand a UX overhaul of our way to interact with the digital world even more significant than the ones the internet and AI brought.
26. If dream data proves useful at all for training brain decoders, then sleep could become one of the most effective ways to collect neural data at scale.

