A Simple Framework for Finding the Next Breakthrough

The Dimension & Dynamics Framework: dynamize, add a dimension, go meta

· 18 min read · ai-ml , futures

How I got here

For a long time I kept noticing the same pattern. Silent films gained sound. Black-and-white film gained colour. In AI, fixed transformations gave way to attention, which adapts to every input.

Each time, technology moved forward in one of two ways: something static became dynamic, or the technology gained a new dimension. It felt like a universal question you could ask of anything: what is static here, and what dimension is missing?

I wondered whether this framework already had a name. Talking it through with Claude, Anthropic’s AI assistant, I found that much of it does. Genrich Altshuller’s TRIZ, a theory of invention built from patent analysis, describes both moves as inventive principles.

So this piece is my attempt to simplify that formalism into three operators and map each one to real innovations. Along the way, some of my intuitions turned out to be published results, and one of my favourite examples turned out to belong somewhere else. I have kept both lessons in.

Core idea

Most technology leaps come from three moves: make something static dynamic, add a dimension, or apply the system to itself. Every technology can be probed with the same two questions: what is fixed here, and what axes does it live on?

This framework is not new in spirit. It formalises two of the 40 inventive principles in TRIZ, Genrich Altshuller’s theory of invention built from tens of thousands of patents:

  • Principle 15, Dynamics: make a rigid or fixed system adjustable, adaptive, or self-optimising.
  • Principle 17, Another Dimension: move from a point to a line, a line to a plane, a plane to a volume, or add layers.

TRIZ also describes a long-run “trend of increasing dynamism” in how technical systems evolve. I add a third operator, Go meta, because it is the move behind many of the biggest leaps in AI.

OperatorQuestion to askOne-line example
DynamizeWhat is fixed that could adapt?Fixed-focus camera → autofocus
Add a dimensionWhat axis is missing?Photograph → film (adds time)
Go metaWhat if the system helps build a better version of itself?Machine that makes tools → machine that makes better tools for itself

Operator 1: Dynamize

Dynamizing means replacing a fixed value with one that changes in response to something. The key design choice is not whether to make it dynamic, but what signal drives the change.

The dynamism ladder

Things rarely jump from fully fixed to fully adaptive. They climb a ladder, and naming the rung tells you where the next innovation is.

0FIXEDNEVER CHANGES—1ADJUSTABLEBY HANDHUMAN2SCHEDULEDCHANGES WITH TIMECLOCK3ADAPTIVECHANGES WITH INPUTSENSOR4GENERATEDSET BY ANOTHER SYSTEMMODEL5SELF-MODIFYINGLEARNS WHILE RUNNINGITSELF

the dynamism ladder · bottom row: the controller each rung needs

Each rung needs a richer controller: a human, a clock, a sensor, a model, then the system itself.

My starting intuition, made precise

In a classic feed-forward layer, the weights are frozen after training. Every input is transformed by the same matrix, whatever it contains.

In attention, the mixing matrix is computed from the input itself. The queries and keys decide, on the fly, how the values get combined. So part of the network is effectively writing the transformation the next step applies.

This has a name: fast weights. Jürgen Schmidhuber proposed “fast weight programmers” in 1992, and Schlag, Irie and Schmidhuber showed in 2021 that linear transformers are mathematically equivalent to them. What I thought I had noticed turned out to be a published result.

Examples by rung

DomainStatic versionDynamic versionRung reachedWhat drives the change
Neural netsFeed-forward weightsAttention weights3The input sequence
Word meaningword2vec (one vector per word)Contextual embeddings (BERT, GPT)3Surrounding words
Sequence modelsS4 state-space model (fixed dynamics)Mamba (selective, input-dependent dynamics)3Each token
ConvolutionFixed kernel gridDeformable / dynamic convolution3Image content
OptimisationFixed learning rateSchedules, then Adam (per-parameter adaptive)2 → 3Time, then gradient history
Compute per inputEvery token uses all layersMixture-of-experts, early exit, adaptive computation time3Difficulty of the token
SoftwareAhead-of-time compilationJust-in-time compilation3Observed runtime behaviour
MemoryStatic allocationDynamic allocation, garbage collection3Program demand
NetworkingStatic routing tablesDynamic routing (OSPF, BGP)3Link state
WebStatic HTML pagesServer-rendered, then personalised pages3User and context
PricingFixed price tagDynamic pricing (airlines, ride-hailing)3Demand
TrafficFixed-timer lightsAdaptive signal control3Sensed traffic
AircraftFixed wingVariable-sweep wing (F-14), flaps1 → 3Speed and flight phase
BicyclesFixed gearDerailleur, then automatic shifting1 → 3Rider, then cadence
OpticsFixed-focus lensZoom, then autofocus1 → 3Scene distance
MedicineOne dose for allClosed-loop insulin pumps3Blood glucose
CinemaSilent filmFilm with synchronised sound3The image track

The last row shows that sound in film is both dynamization and a new dimension. Many real leaps are both at once.

Operator 2: Add a dimension

Adding a dimension means giving the system a new axis along which it can vary, store, or carry information. “Dimension” is wider than geometry; it is any independent axis.

Dimension catalogue

When analysing a technology, check it against these seven axes and ask which ones it ignores.

AxisQuestionExamples of adding it
SpaceCan it extend into another spatial direction or layer?1D barcode → 2D QR code; planar chips → 3D-stacked memory (HBM, 3D NAND); 2D printing → 3D printing
TimeCan it capture or act across time?Photograph → film; paper map → live traffic map; single answer → chain-of-thought reasoning
Spectrum / frequencyCan it use many channels at once?Black-and-white → colour film; single-carrier radio → OFDM (Wi-Fi, 4G)
ModalityCan it handle another kind of signal?Silent → sound film; text-only LLMs → multimodal models
ParallelismCan many copies run side by side?Single-core → multi-core CPU; one antenna → MIMO; single-head → multi-head attention
Scale / resolutionCan it work at several scales at once?Single-scale images → image pyramids, U-Net; flat indexes → hierarchical search
Context / audienceCan it vary by who or where?One edition for all → personalised feeds; mono → stereo → spatial audio

The pattern inside the examples

A new dimension usually arrives as a separate add-on, then becomes native. Sound was first played from a disc beside the projector, then printed on the film strip. Multimodal AI followed the same path: vision encoders bolted onto text models, then models trained on all modalities from the start.

In AI specifically, scaling itself has gained dimensions. First came parameters, then data (Chinchilla scaling, 2022), then inference-time compute, where models “think longer” on hard problems. Each new axis reopened progress when the previous one slowed.

Operator 3: Go meta

Going meta means applying the system’s own kind of process to the system itself. It is rung 5 of the dynamism ladder, applied to the whole system, and it deserves its own operator because it compounds.

Here it is at increasing levels of self-reference:

LevelWhat controls whatExamples
Learning to learnA model learns how to update a modelMAML meta-learning (2017); learned optimisers
Designing the designerA system searches for its own architecture or dataNeural architecture search; models generating their own training data

Outside AI the same move appears whenever a tool is used to build better tools. Compilers that compile themselves, machine tools that manufacture machine tools, and CAD software used to design chips that run CAD software are all “go meta” loops.

The benefit is compounding improvement. The risk is loss of predictability, since the behaviour now depends on a controller that is itself changing.

Closing the loop: recursive self-improvement

Push the last row of the table one step further and the loop closes completely. The system improves not only itself but the process that improves it, so each better version is also a better improver. This is recursive self-improvement (RSI), and in the era of LLMs it has moved from a thought experiment to an active research trend, and a lot of hype.

The idea is old. In 1965, I. J. Good wrote that a machine able to design better machines could set off an “intelligence explosion”. In 2003, Schmidhuber’s Gödel machine formalised a program that rewrites its own code once it can prove the rewrite is an improvement.

What is new is that LLMs now take part in several parts of their own improvement loop:

Part of the loopWhat the model improvesExamples
DataIts own training examplesSTaR (2022) fine-tunes a model on its own correct reasoning; SEAL writes its own fine-tuning data
FeedbackThe judge that scores itConstitutional AI (2022) replaces human labels with AI feedback
Code and scaffoldingThe agent built around the modelSakana’s Darwin Gödel Machine (2025) rewrites its own agent code and keeps the versions that score better
InfrastructureThe kernels it trains onGoogle DeepMind’s AlphaEvolve (2025) found a faster kernel that cut training time for the Gemini models it runs on

Seen through this framework, RSI is rung 5 of the dynamism ladder applied to the whole development process, not only to the weights. And the question from Operator 1 still decides everything: what signal drives the change? A self-improving loop is only as good as its evaluator.

That is why progress shows up first where results can be checked automatically: code with tests, maths with proofs, games with scores. Where no reliable check exists, a model trained on its own output can drift and degrade, an effect known as model collapse (Shumailov et al., 2024).

So the honest reading of the hype is this. RSI is a real direction, but today it is many partial loops, each closed around one component and one verifier, with humans still holding the outer loop. Whether compounding becomes an “explosion” depends on whether the verifiers can scale as fast as the models they check.

Supporting principles

Three more ideas make the operators work in practice. They are kept short on purpose.

Segment (TRIZ Principle 1)

Dynamizing is often impossible until a monolith is split into parts that can vary independently. Dense networks became mixture-of-experts; monolithic software became microservices; one big antenna became phased arrays of small ones.

Merge into a super-system (TRIZ trend)

Mature devices get absorbed into a larger system that adds dimensions for free. The phone swallowed the camera, GPS, music player and wallet, and each gained connectivity and context it never had alone.

The complexify-then-compress cycle

This is where FlashAttention belongs. I first filed it under static-to-dynamic, but it computes exactly the same attention, just far more efficiently. Dynamizing or adding a dimension almost always makes a system more expensive. A second innovation then compresses the cost without losing the new capability.

COMPLEXIFYNEW DIMENSION OR DYNAMISMCAPABILITY JUMP,COST JUMPCOMPRESSSAME FUNCTION, CHEAPERIDEALITY

the complexify-then-compress cycle

Complexify stepCompress step
Attention (quadratic cost in sequence length)FlashAttention (same exact result, far less memory traffic)
Large dense modelsQuantisation, distillation
Colour film (costly Technicolor process)Cheap single-strip colour stock
Mixture-of-experts (more parameters)Sparse routing (only a few experts run per token)

TRIZ calls the goal of this cycle ideality: more useful function per unit of cost and harm. It is the compass that tells you whether a new dimension is worth keeping.

The method

Apply these six steps to any technology. Each step is a question with a concrete output.

  1. Inventory the parameters. List every quantity the system uses, then mark each as fixed or variable, and note its rung on the dynamism ladder.
  2. Map the dimensions. Check the system against the seven axes in the catalogue. Mark which ones it uses and which it ignores.
  3. Find the mismatch. Pick the fixed parameter or missing axis that fails most often across real situations. Dynamism pays exactly where conditions vary.
  4. Choose the controller. Decide what signal will drive the change: a person, time, the input, the context, or another model. Then pick the target rung.
  5. Pay, then compress. Build the capable but expensive version first. Then look for the compression step that makes it affordable.
  6. Check ideality. Compare function gained against cost, complexity and new failure modes. Keep it only if the ratio improves.

Worked example: a fixed-size LLM context window

  • Step 1: the context length is a fixed parameter, rung 0.
  • Step 2: the model ignores the time axis beyond one conversation; memory does not persist.
  • Step 3: the mismatch is long projects, where the needed context varies from 1 page to thousands.
  • Step 4: the controller could be the model itself, deciding what to store and retrieve.
  • Step 5: the expensive version is huge context windows; the compression is learned memory or retrieval.
  • Step 6: a gain if recall quality holds and cost per task falls.

When not to apply it

Static is often the right answer. Dynamizing trades predictability for adaptability, and some systems need predictability most.

  • Safety-critical systems: aviation and medical software favour fixed, verifiable behaviour, because adaptive behaviour is hard to certify.
  • Stable environments: if conditions never vary, adaptation adds cost for no gain. A fixed-gear bike still wins on a velodrome.
  • Standards and interfaces: protocols, file formats and plugs are valuable because they do not change.
  • Debuggability: every dynamic part is a new thing that can drift, oscillate or be manipulated, as dynamic pricing and recommendation feeds show.

A useful rule: dynamize the parts that face a varying world, and keep fixed the parts others depend on.

Expected upgrades and innovations for state-of-the-art LLMs

The biggest remaining static part of today’s LLMs is their weights at inference time. Almost every promising direction is a way to dynamize that, add a missing axis, or close a meta loop.

Inventory: what is still static

ComponentCurrent rungStatus
Weights during use0 (frozen after training)Largest open opportunity
Tokenizer0 (fixed vocabulary)Early research
Compute per token3 in MoE models, 0 in dense onesPartly dynamized
Reasoning length3 (reasoning models choose how long to think)Recently dynamized
Memory across sessions0–1 (external notes, retrieval)Bolted on, not native
Reasoning mediumFixed to text tokensEarly research

Direction 1: Dynamize the weights (rung 4–5)

Models that update part of themselves while running, instead of relying only on the prompt. Google’s Titans architecture adds a memory module that keeps learning at test time. Sakana’s Transformer² adjusts weight components per task on the fly, and MIT’s SEAL has a model write its own fine-tuning data and updates. This is the fast-weights idea taken to its conclusion.

Direction 2: Dynamize the input units

Fixed tokenizers split text the same way whatever it contains. Meta’s Byte Latent Transformer groups raw bytes into patches sized by how predictable the next bytes are, spending more compute where the text is hard.

Direction 3: Dynamize depth and width per token

Mixture-of-experts varies which parameters run; Mixture-of-Depths varies how many layers a token passes through. The next step is models that allocate compute per token, per step and per task from one unified budget.

Direction 4: Add a reasoning dimension beyond text

Today, reasoning happens in written tokens. Meta’s Coconut work lets models reason in continuous hidden states instead, a richer space than words. Parallel reasoning, exploring several branches at once and merging them, adds a width axis to thinking.

Direction 5: Add the time axis for real

Persistent memory and long-running agents give models a life beyond one conversation. The open problem is making memory native and learned, rather than a retrieval system attached on the side.

Direction 6: Add the action and world axes

Text describes the world; actions change it. World models and embodied training add a grounded dimension: predicting consequences, not just words.

Direction 7: Close the meta loop carefully

Models already help design training data, evaluate other models and write code for their own tooling. A full loop, where a model improves its own learning process, is the ultimate “go meta” move: the recursive self-improvement loop described in Operator 3. It is also where the “when not to apply it” section matters most, since predictability and oversight get harder at every rung.

Compress what follows

Each direction will be expensive at first. Expect the matching compress step to follow, as FlashAttention followed attention: cheaper test-time updates, sparse memory, and distilled reasoning.

The research examples are summarised briefly here; read the original papers before relying on the details.