Skip to content
Kudos AI

Research

Where the frontier is, and how it got here. A pinned board of open problems the field is actively working on, followed by the landmark papers that built the foundations.

Open problems today

The frontier

Live, field-level questions that remain genuinely unsettled. Each is a place where a careful contribution could still move the discipline.

Deep LearningStatistics

Why Overparameterized Networks Generalize

Why do neural networks with far more parameters than training examples generalize well, when classical theory predicts they should overfit badly?

The bias-variance trade-off says that a model flexible enough to interpolate its training data should have ruinous variance. Modern networks routinely fit their training set exactly and still generalize, and test error has been observed to fall again beyond the interpolation point rather than continuing to rise. Explanations appeal to implicit regularization by gradient descent, to properties of the loss landscape, and to the structure of real data, but no account has settled the question. Until it is settled, the field cannot say in advance how large a model a given dataset can support, and capacity is chosen empirically.

Deep LearningGenerative AI

A First-Principles Account of Scaling Laws

Why does model performance improve as a smooth power law in parameters, data, and compute, and what determines the exponents?

The empirical regularity is strong enough to guide multi-million-dollar training decisions, yet it rests on curve fitting rather than derivation. Without a theory, the exponents cannot be predicted for a new architecture or domain, and there is no principled way to know whether a trend will continue or break. A derivation would turn the most consequential planning decision in the field from extrapolation into calculation.

Deep LearningGenerative AI

Mechanistic Interpretability of Large Models

Can the computation a large trained network performs be reverse-engineered into human-understandable algorithms?

A trained network is a very large array of weights that provably computes something useful, with no accompanying account of how. Progress has been made on identifying interpretable circuits and features in smaller models, but individual units frequently encode several unrelated concepts at once, which frustrates straightforward reading. Without interpretability there is no way to verify that a model relies on legitimate structure rather than a spurious correlation, which matters wherever the stakes are high.

Game TheoryMathematics

The Computational Cost of Finding Equilibria

Nash proved an equilibrium always exists, but how hard is it to actually find one, and what does that imply for equilibrium as a predictive concept?

Existence and computability are different properties. Computing a Nash equilibrium is now known to be complete for the complexity class PPAD, which is strong evidence that no general efficient algorithm exists. This raises a genuine question about the concept itself: if the players in a modelled situation could not feasibly compute the equilibrium, it is not obvious why their behaviour should be expected to reach it. The tension between existence, computability, and predictive relevance remains unresolved.

Reinforcement LearningMachine Learning

Sample-Efficient Reinforcement Learning

How can an agent learn effective behaviour from a realistic amount of experience, and transfer it when the environment shifts?

Reinforcement learning agents routinely need millions of interactions to learn tasks a human picks up in a handful of attempts, which confines the successes largely to simulation, where experience is cheap. Two obstacles compound: sparse rewards mean informative signal reaches early decisions only slowly, and policies learned in one environment often degrade sharply when the dynamics change. Closing the gap is what stands between current methods and reliable deployment in the physical world.

Deep LearningComputer Vision

Robustness to Adversarial Perturbation

Why are accurate models so easily fooled by tiny, deliberately chosen input perturbations, and can robustness be achieved without sacrificing accuracy?

Perturbations far too small for a person to notice can flip a confident classification entirely, which means high average accuracy does not imply the model has learned what a human would call the concept. Defences have tended to fail against stronger subsequent attacks, and there is evidence of a genuine tension between robustness and clean accuracy. The question is both practical, for any security-relevant deployment, and conceptual, since it suggests these models generalize in a different way than their accuracy suggests.

Machine LearningProbability

Learning Causal Structure from Observation

Can models learn causal structure, rather than correlation, from data that was mostly not gathered through controlled experiment?

Predictive models capture association, which is sufficient while the world stays as it was during training and insufficient the moment anything intervenes. Answering what would happen if a variable were changed requires causal structure, and that structure is generally not identifiable from observational data without additional assumptions. Establishing which assumptions are both plausible and sufficient is an open question, and it is what separates models that predict from models that support decisions.

Generative AIDeep Learning

Efficient Attention over Long Contexts

Can the quadratic cost of self-attention be avoided without losing the ability to relate any two positions directly?

Self-attention compares every position with every other, so cost grows with the square of sequence length, which is the binding constraint on context size. Numerous alternatives, sparse attention patterns, linear approximations, and recurrent state-space models among them, reduce the asymptotic cost, but typically give up some of the unrestricted pairwise access that makes attention effective. Whether the full capability can be retained at lower cost is unresolved.

Landmark papers

The foundations

Each paper is summarised, placed in context, and linked to deeper reviews and reference entries, why it mattered, not just what it said.

  1. 1936Alan Turing

    On Computable Numbers, with an Application to the Entscheidungsproblem

    Introduces an abstract machine that reads and writes symbols on a tape according to a finite table of rules, and uses it to show that no general procedure can decide whether an arbitrary program halts.

    MathematicsArtificial IntelligenceProgramming
  2. 1948Claude E. Shannon

    A Mathematical Theory of Communication

    Defines information quantitatively, introduces entropy as the measure of a source’s uncertainty, and proves limits on lossless compression and on reliable transmission over a noisy channel.

    Information TheoryProbabilityMathematics
  3. 1950John Nash

    Equilibrium Points in N-Person Games

    Proves that every finite game with any number of players has at least one equilibrium point, provided players may use mixed strategies.

    Game TheoryMathematics
  4. 1950Claude E. Shannon

    Programming a Computer for Playing Chess

    Lays out how a machine might play chess: represent positions, generate legal moves, search the game tree with minimax, and evaluate non-terminal positions with a heuristic scoring function.

    Search & PlanningGame TheoryArtificial Intelligence
  5. 1950Alan Turing

    Computing Machinery and Intelligence

    Proposes replacing the question "can machines think?" with a behavioural test, in which an interrogator tries to distinguish a machine from a human through written conversation.

    Artificial Intelligence
  6. 1957Frank Rosenblatt

    The Perceptron: A Perceiving and Recognizing Automaton

    Introduces the perceptron, a trainable unit that computes a weighted sum of its inputs and fires if the sum exceeds a threshold, with a rule for adjusting weights from labelled examples.

    Machine LearningDeep LearningArtificial Intelligence
  7. 1967John C. Harsanyi

    Games with Incomplete Information Played by Bayesian Players

    Shows how games in which players are uncertain about one another’s payoffs can be transformed into games of complete but imperfect information, by treating each player as having a randomly assigned "type".

    Game TheoryProbability
  8. 1968Peter E. Hart, Nils J. Nilsson, Bertram Raphael

    A Formal Basis for the Heuristic Determination of Minimum Cost Paths

    Introduces the A* algorithm, which orders search by the sum of the cost already incurred and a heuristic estimate of the cost remaining, and proves it optimal when the heuristic never overestimates.

    Search & PlanningArtificial Intelligence
  9. 1984Leo Breiman, Jerome Friedman, Richard A. Olshen, Charles J. Stone

    Classification and Regression Trees

    Establishes the CART methodology: growing decision trees by recursively choosing the split that most improves node purity, then pruning back the fully grown tree using held-out data.

    Machine LearningStatistics
  10. 1986David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams

    Learning Internal Representations by Error Propagation

    Presents backpropagation as a general method for training multilayer networks, showing that hidden layers can learn useful internal representations rather than needing to be designed by hand.

    Deep LearningMachine LearningOptimization
  11. 1989Christopher J. C. H. Watkins

    Models of Delayed Reinforcement Learning

    Develops Q-learning, an algorithm that estimates the value of each action in each state directly from experience, requiring no model of the environment’s transition probabilities.

    Reinforcement LearningMachine Learning
  12. 1995Corinna Cortes, Vladimir N. Vapnik

    Support-Vector Networks

    Introduces the support vector machine with a soft margin, separating classes by the widest possible margin while permitting bounded violations, and using kernels to obtain non-linear boundaries.

    Machine LearningOptimizationMathematics
  13. 1996Leo Breiman

    Bagging Predictors

    Introduces bootstrap aggregating: fitting a model to many bootstrap resamples of the training data and averaging the predictions, which reduces variance without increasing bias.

    Machine LearningStatistics
  14. 2015Rico Sennrich, Barry Haddow, Alexandra Birch

    Neural Machine Translation of Rare Words with Subword Units

    Adapts byte pair encoding to text segmentation, building a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair, so that rare words decompose into known fragments.

    Natural Language ProcessingDeep Learning
  15. 2017Ashish Vaswani et al.

    Attention Is All You Need

    Introduces the transformer, an architecture built entirely from attention and feed-forward layers with no recurrence, originally developed for machine translation.

    Generative AIDeep LearningNatural Language Processing
  16. 2018Jacob Devlin et al.

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Pretrains a transformer encoder to predict masked tokens using context from both directions, then fine-tunes the same model on downstream tasks with a small task-specific head.

    Natural Language ProcessingDeep LearningGenerative AI
  17. 2020Tom B. Brown et al.

    Language Models are Few-Shot Learners

    Describes GPT-3 and shows that a sufficiently large decoder-only language model can perform new tasks from a handful of examples supplied in its prompt, with no gradient updates.

    Generative AINatural Language ProcessingDeep Learning
  18. 2020Alexey Dosovitskiy et al.

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    Applies a standard transformer directly to images by cutting each image into fixed-size patches and treating the sequence of patches as tokens, with no convolutions.

    Computer VisionDeep LearningGenerative AI