A Simple Framework for Finding the Next Breakthrough
The Dimension & Dynamics Framework: dynamize, add a dimension, go meta
How I got here
For a long time I kept noticing the same pattern. Silent films gained sound. Black-and-white film gained colour. In AI, fixed transformations gave way to attention, which adapts to every input.
Each time, technology moved forward in one of two ways: something static became dynamic, or the technology gained a new dimension. It felt like a universal question you could ask of anything: what is static here, and what dimension is missing?
I wondered whether this framework already had a name. Talking it through with Claude, Anthropic’s AI assistant, I found that much of it does. Genrich Altshuller’s TRIZ, a theory of invention built from patent analysis, describes both moves as inventive principles.
So this piece is my attempt to simplify that formalism into three operators and map each one to real innovations. Along the way, some of my intuitions turned out to be published results, and one of my favourite examples turned out to belong somewhere else. I have kept both lessons in.
Core idea
Most technology leaps come from three moves: make something static dynamic, add a dimension, or apply the system to itself. Every technology can be probed with the same two questions: what is fixed here, and what axes does it live on?
This framework is not new in spirit. It formalises two of the 40 inventive principles in TRIZ, Genrich Altshuller’s theory of invention built from tens of thousands of patents:
- Principle 15, Dynamics: make a rigid or fixed system adjustable, adaptive, or self-optimising.
- Principle 17, Another Dimension: move from a point to a line, a line to a plane, a plane to a volume, or add layers.
TRIZ also describes a long-run “trend of increasing dynamism” in how technical systems evolve. I add a third operator, Go meta, because it is the move behind many of the biggest leaps in AI.
| Operator | Question to ask | One-line example |
|---|---|---|
| Dynamize | What is fixed that could adapt? | Fixed-focus camera → autofocus |
| Add a dimension | What axis is missing? | Photograph → film (adds time) |
| Go meta | What if the system helps build a better version of itself? | Machine that makes tools → machine that makes better tools for itself |
Operator 1: Dynamize
Dynamizing means replacing a fixed value with one that changes in response to something. The key design choice is not whether to make it dynamic, but what signal drives the change.
The dynamism ladder
Things rarely jump from fully fixed to fully adaptive. They climb a ladder, and naming the rung tells you where the next innovation is.
the dynamism ladder · bottom row: the controller each rung needs
Each rung needs a richer controller: a human, a clock, a sensor, a model, then the system itself.
My starting intuition, made precise
In a classic feed-forward layer, the weights are frozen after training. Every input is transformed by the same matrix, whatever it contains.
In attention, the mixing matrix is computed from the input itself. The queries and keys decide, on the fly, how the values get combined. So part of the network is effectively writing the transformation the next step applies.
This has a name: fast weights. Jürgen Schmidhuber proposed “fast weight programmers” in 1992, and Schlag, Irie and Schmidhuber showed in 2021 that linear transformers are mathematically equivalent to them. What I thought I had noticed turned out to be a published result.
Examples by rung
| Domain | Static version | Dynamic version | Rung reached | What drives the change |
|---|---|---|---|---|
| Neural nets | Feed-forward weights | Attention weights | 3 | The input sequence |
| Word meaning | word2vec (one vector per word) | Contextual embeddings (BERT, GPT) | 3 | Surrounding words |
| Sequence models | S4 state-space model (fixed dynamics) | Mamba (selective, input-dependent dynamics) | 3 | Each token |
| Convolution | Fixed kernel grid | Deformable / dynamic convolution | 3 | Image content |
| Optimisation | Fixed learning rate | Schedules, then Adam (per-parameter adaptive) | 2 → 3 | Time, then gradient history |
| Compute per input | Every token uses all layers | Mixture-of-experts, early exit, adaptive computation time | 3 | Difficulty of the token |
| Software | Ahead-of-time compilation | Just-in-time compilation | 3 | Observed runtime behaviour |
| Memory | Static allocation | Dynamic allocation, garbage collection | 3 | Program demand |
| Networking | Static routing tables | Dynamic routing (OSPF, BGP) | 3 | Link state |
| Web | Static HTML pages | Server-rendered, then personalised pages | 3 | User and context |
| Pricing | Fixed price tag | Dynamic pricing (airlines, ride-hailing) | 3 | Demand |
| Traffic | Fixed-timer lights | Adaptive signal control | 3 | Sensed traffic |
| Aircraft | Fixed wing | Variable-sweep wing (F-14), flaps | 1 → 3 | Speed and flight phase |
| Bicycles | Fixed gear | Derailleur, then automatic shifting | 1 → 3 | Rider, then cadence |
| Optics | Fixed-focus lens | Zoom, then autofocus | 1 → 3 | Scene distance |
| Medicine | One dose for all | Closed-loop insulin pumps | 3 | Blood glucose |
| Cinema | Silent film | Film with synchronised sound | 3 | The image track |
The last row shows that sound in film is both dynamization and a new dimension. Many real leaps are both at once.
Operator 2: Add a dimension
Adding a dimension means giving the system a new axis along which it can vary, store, or carry information. “Dimension” is wider than geometry; it is any independent axis.
Dimension catalogue
When analysing a technology, check it against these seven axes and ask which ones it ignores.
| Axis | Question | Examples of adding it |
|---|---|---|
| Space | Can it extend into another spatial direction or layer? | 1D barcode → 2D QR code; planar chips → 3D-stacked memory (HBM, 3D NAND); 2D printing → 3D printing |
| Time | Can it capture or act across time? | Photograph → film; paper map → live traffic map; single answer → chain-of-thought reasoning |
| Spectrum / frequency | Can it use many channels at once? | Black-and-white → colour film; single-carrier radio → OFDM (Wi-Fi, 4G) |
| Modality | Can it handle another kind of signal? | Silent → sound film; text-only LLMs → multimodal models |
| Parallelism | Can many copies run side by side? | Single-core → multi-core CPU; one antenna → MIMO; single-head → multi-head attention |
| Scale / resolution | Can it work at several scales at once? | Single-scale images → image pyramids, U-Net; flat indexes → hierarchical search |
| Context / audience | Can it vary by who or where? | One edition for all → personalised feeds; mono → stereo → spatial audio |
The pattern inside the examples
A new dimension usually arrives as a separate add-on, then becomes native. Sound was first played from a disc beside the projector, then printed on the film strip. Multimodal AI followed the same path: vision encoders bolted onto text models, then models trained on all modalities from the start.
In AI specifically, scaling itself has gained dimensions. First came parameters, then data (Chinchilla scaling, 2022), then inference-time compute, where models “think longer” on hard problems. Each new axis reopened progress when the previous one slowed.
Operator 3: Go meta
Going meta means applying the system’s own kind of process to the system itself. It is rung 5 of the dynamism ladder, applied to the whole system, and it deserves its own operator because it compounds.
Here it is at increasing levels of self-reference:
| Level | What controls what | Examples |
|---|---|---|
| Learning to learn | A model learns how to update a model | MAML meta-learning (2017); learned optimisers |
| Designing the designer | A system searches for its own architecture or data | Neural architecture search; models generating their own training data |
Outside AI the same move appears whenever a tool is used to build better tools. Compilers that compile themselves, machine tools that manufacture machine tools, and CAD software used to design chips that run CAD software are all “go meta” loops.
The benefit is compounding improvement. The risk is loss of predictability, since the behaviour now depends on a controller that is itself changing.
Closing the loop: recursive self-improvement
Push the last row of the table one step further and the loop closes completely. The system improves not only itself but the process that improves it, so each better version is also a better improver. This is recursive self-improvement (RSI), and in the era of LLMs it has moved from a thought experiment to an active research trend, and a lot of hype.
The idea is old. In 1965, I. J. Good wrote that a machine able to design better machines could set off an “intelligence explosion”. In 2003, Schmidhuber’s Gödel machine formalised a program that rewrites its own code once it can prove the rewrite is an improvement.
What is new is that LLMs now take part in several parts of their own improvement loop:
| Part of the loop | What the model improves | Examples |
|---|---|---|
| Data | Its own training examples | STaR (2022) fine-tunes a model on its own correct reasoning; SEAL writes its own fine-tuning data |
| Feedback | The judge that scores it | Constitutional AI (2022) replaces human labels with AI feedback |
| Code and scaffolding | The agent built around the model | Sakana’s Darwin Gödel Machine (2025) rewrites its own agent code and keeps the versions that score better |
| Infrastructure | The kernels it trains on | Google DeepMind’s AlphaEvolve (2025) found a faster kernel that cut training time for the Gemini models it runs on |
Seen through this framework, RSI is rung 5 of the dynamism ladder applied to the whole development process, not only to the weights. And the question from Operator 1 still decides everything: what signal drives the change? A self-improving loop is only as good as its evaluator.
That is why progress shows up first where results can be checked automatically: code with tests, maths with proofs, games with scores. Where no reliable check exists, a model trained on its own output can drift and degrade, an effect known as model collapse (Shumailov et al., 2024).
So the honest reading of the hype is this. RSI is a real direction, but today it is many partial loops, each closed around one component and one verifier, with humans still holding the outer loop. Whether compounding becomes an “explosion” depends on whether the verifiers can scale as fast as the models they check.
Supporting principles
Three more ideas make the operators work in practice. They are kept short on purpose.
Segment (TRIZ Principle 1)
Dynamizing is often impossible until a monolith is split into parts that can vary independently. Dense networks became mixture-of-experts; monolithic software became microservices; one big antenna became phased arrays of small ones.
Merge into a super-system (TRIZ trend)
Mature devices get absorbed into a larger system that adds dimensions for free. The phone swallowed the camera, GPS, music player and wallet, and each gained connectivity and context it never had alone.
The complexify-then-compress cycle
This is where FlashAttention belongs. I first filed it under static-to-dynamic, but it computes exactly the same attention, just far more efficiently. Dynamizing or adding a dimension almost always makes a system more expensive. A second innovation then compresses the cost without losing the new capability.
the complexify-then-compress cycle
| Complexify step | Compress step |
|---|---|
| Attention (quadratic cost in sequence length) | FlashAttention (same exact result, far less memory traffic) |
| Large dense models | Quantisation, distillation |
| Colour film (costly Technicolor process) | Cheap single-strip colour stock |
| Mixture-of-experts (more parameters) | Sparse routing (only a few experts run per token) |
TRIZ calls the goal of this cycle ideality: more useful function per unit of cost and harm. It is the compass that tells you whether a new dimension is worth keeping.
The method
Apply these six steps to any technology. Each step is a question with a concrete output.
- Inventory the parameters. List every quantity the system uses, then mark each as fixed or variable, and note its rung on the dynamism ladder.
- Map the dimensions. Check the system against the seven axes in the catalogue. Mark which ones it uses and which it ignores.
- Find the mismatch. Pick the fixed parameter or missing axis that fails most often across real situations. Dynamism pays exactly where conditions vary.
- Choose the controller. Decide what signal will drive the change: a person, time, the input, the context, or another model. Then pick the target rung.
- Pay, then compress. Build the capable but expensive version first. Then look for the compression step that makes it affordable.
- Check ideality. Compare function gained against cost, complexity and new failure modes. Keep it only if the ratio improves.
Worked example: a fixed-size LLM context window
- Step 1: the context length is a fixed parameter, rung 0.
- Step 2: the model ignores the time axis beyond one conversation; memory does not persist.
- Step 3: the mismatch is long projects, where the needed context varies from 1 page to thousands.
- Step 4: the controller could be the model itself, deciding what to store and retrieve.
- Step 5: the expensive version is huge context windows; the compression is learned memory or retrieval.
- Step 6: a gain if recall quality holds and cost per task falls.
When not to apply it
Static is often the right answer. Dynamizing trades predictability for adaptability, and some systems need predictability most.
- Safety-critical systems: aviation and medical software favour fixed, verifiable behaviour, because adaptive behaviour is hard to certify.
- Stable environments: if conditions never vary, adaptation adds cost for no gain. A fixed-gear bike still wins on a velodrome.
- Standards and interfaces: protocols, file formats and plugs are valuable because they do not change.
- Debuggability: every dynamic part is a new thing that can drift, oscillate or be manipulated, as dynamic pricing and recommendation feeds show.
A useful rule: dynamize the parts that face a varying world, and keep fixed the parts others depend on.
Expected upgrades and innovations for state-of-the-art LLMs
The biggest remaining static part of today’s LLMs is their weights at inference time. Almost every promising direction is a way to dynamize that, add a missing axis, or close a meta loop.
Inventory: what is still static
| Component | Current rung | Status |
|---|---|---|
| Weights during use | 0 (frozen after training) | Largest open opportunity |
| Tokenizer | 0 (fixed vocabulary) | Early research |
| Compute per token | 3 in MoE models, 0 in dense ones | Partly dynamized |
| Reasoning length | 3 (reasoning models choose how long to think) | Recently dynamized |
| Memory across sessions | 0–1 (external notes, retrieval) | Bolted on, not native |
| Reasoning medium | Fixed to text tokens | Early research |
Direction 1: Dynamize the weights (rung 4–5)
Models that update part of themselves while running, instead of relying only on the prompt. Google’s Titans architecture adds a memory module that keeps learning at test time. Sakana’s Transformer² adjusts weight components per task on the fly, and MIT’s SEAL has a model write its own fine-tuning data and updates. This is the fast-weights idea taken to its conclusion.
Direction 2: Dynamize the input units
Fixed tokenizers split text the same way whatever it contains. Meta’s Byte Latent Transformer groups raw bytes into patches sized by how predictable the next bytes are, spending more compute where the text is hard.
Direction 3: Dynamize depth and width per token
Mixture-of-experts varies which parameters run; Mixture-of-Depths varies how many layers a token passes through. The next step is models that allocate compute per token, per step and per task from one unified budget.
Direction 4: Add a reasoning dimension beyond text
Today, reasoning happens in written tokens. Meta’s Coconut work lets models reason in continuous hidden states instead, a richer space than words. Parallel reasoning, exploring several branches at once and merging them, adds a width axis to thinking.
Direction 5: Add the time axis for real
Persistent memory and long-running agents give models a life beyond one conversation. The open problem is making memory native and learned, rather than a retrieval system attached on the side.
Direction 6: Add the action and world axes
Text describes the world; actions change it. World models and embodied training add a grounded dimension: predicting consequences, not just words.
Direction 7: Close the meta loop carefully
Models already help design training data, evaluate other models and write code for their own tooling. A full loop, where a model improves its own learning process, is the ultimate “go meta” move: the recursive self-improvement loop described in Operator 3. It is also where the “when not to apply it” section matters most, since predictability and oversight get harder at every rung.
Compress what follows
Each direction will be expensive at first. Expect the matching compress step to follow, as FlashAttention followed attention: cheaper test-time updates, sparse memory, and distilled reasoning.
The research examples are summarised briefly here; read the original papers before relying on the details.