Live, field-level questions that remain genuinely unsettled. Each is a place where a careful contribution could still move the discipline.
Deep LearningStatistics
Why Overparameterized Networks Generalize
Why do neural networks with far more parameters than training examples generalize well, when classical theory predicts they should overfit badly?
The bias-variance trade-off says that a model flexible enough to interpolate its training data should have ruinous variance. Modern networks routinely fit their training set exactly and still generalize, and test error has been observed to fall again beyond the interpolation point rather than continuing to rise. Explanations appeal to implicit regularization by gradient descent, to properties of the loss landscape, and to the structure of real data, but no account has settled the question. Until it is settled, the field cannot say in advance how large a model a given dataset can support, and capacity is chosen empirically.
Deep LearningGenerative AI
A First-Principles Account of Scaling Laws
Why does model performance improve as a smooth power law in parameters, data, and compute, and what determines the exponents?
The empirical regularity is strong enough to guide multi-million-dollar training decisions, yet it rests on curve fitting rather than derivation. Without a theory, the exponents cannot be predicted for a new architecture or domain, and there is no principled way to know whether a trend will continue or break. A derivation would turn the most consequential planning decision in the field from extrapolation into calculation.
Deep LearningGenerative AI
Mechanistic Interpretability of Large Models
Can the computation a large trained network performs be reverse-engineered into human-understandable algorithms?
A trained network is a very large array of weights that provably computes something useful, with no accompanying account of how. Progress has been made on identifying interpretable circuits and features in smaller models, but individual units frequently encode several unrelated concepts at once, which frustrates straightforward reading. Without interpretability there is no way to verify that a model relies on legitimate structure rather than a spurious correlation, which matters wherever the stakes are high.
Game TheoryMathematics
The Computational Cost of Finding Equilibria
Nash proved an equilibrium always exists, but how hard is it to actually find one, and what does that imply for equilibrium as a predictive concept?
Existence and computability are different properties. Computing a Nash equilibrium is now known to be complete for the complexity class PPAD, which is strong evidence that no general efficient algorithm exists. This raises a genuine question about the concept itself: if the players in a modelled situation could not feasibly compute the equilibrium, it is not obvious why their behaviour should be expected to reach it. The tension between existence, computability, and predictive relevance remains unresolved.
Reinforcement LearningMachine Learning
Sample-Efficient Reinforcement Learning
How can an agent learn effective behaviour from a realistic amount of experience, and transfer it when the environment shifts?
Reinforcement learning agents routinely need millions of interactions to learn tasks a human picks up in a handful of attempts, which confines the successes largely to simulation, where experience is cheap. Two obstacles compound: sparse rewards mean informative signal reaches early decisions only slowly, and policies learned in one environment often degrade sharply when the dynamics change. Closing the gap is what stands between current methods and reliable deployment in the physical world.
Deep LearningComputer Vision
Robustness to Adversarial Perturbation
Why are accurate models so easily fooled by tiny, deliberately chosen input perturbations, and can robustness be achieved without sacrificing accuracy?
Perturbations far too small for a person to notice can flip a confident classification entirely, which means high average accuracy does not imply the model has learned what a human would call the concept. Defences have tended to fail against stronger subsequent attacks, and there is evidence of a genuine tension between robustness and clean accuracy. The question is both practical, for any security-relevant deployment, and conceptual, since it suggests these models generalize in a different way than their accuracy suggests.
Machine LearningProbability
Learning Causal Structure from Observation
Can models learn causal structure, rather than correlation, from data that was mostly not gathered through controlled experiment?
Predictive models capture association, which is sufficient while the world stays as it was during training and insufficient the moment anything intervenes. Answering what would happen if a variable were changed requires causal structure, and that structure is generally not identifiable from observational data without additional assumptions. Establishing which assumptions are both plausible and sufficient is an open question, and it is what separates models that predict from models that support decisions.
Generative AIDeep Learning
Efficient Attention over Long Contexts
Can the quadratic cost of self-attention be avoided without losing the ability to relate any two positions directly?
Self-attention compares every position with every other, so cost grows with the square of sequence length, which is the binding constraint on context size. Numerous alternatives, sparse attention patterns, linear approximations, and recurrent state-space models among them, reduce the asymptotic cost, but typically give up some of the unrestricted pairwise access that makes attention effective. Whether the full capability can be retained at lower cost is unresolved.