Andrej Karpathy on

memory

6 entries, 2 Aug 2026 to 10 Aug 2026

On the recordsourced and dated, oldest first

    1. posted

      Karpathy wonders if ngrams or decision trees give better log probs than neural models at 25KB constraint

      “@MattBeton so fun! :) at some point i wonder if ngram (tables) or even something like decision trees start to give superior log probs, and at much smaller program lengths overall (sum of program + weights). i.e. what is the best val loss model overall, for 25KB of user space. fun q!”

      2 Aug 2026 · X · source
    2. reported

      Karpathy compares human sleep to a distillation process that consolidates daily experiences into long-term memory, a capability he says LLMs lack since they restart with empty context windows.

      “I feel like when I’m awake, I’m building up a context window of stuff that’s happening during the day. But when I go to sleep, something magical happens… a process of distillation into the weights of my brain. We don’t have an equivalent of that in LLMs. When you boot them up, they have zero tokens in the window. They’re always restarting from scratch.”

      5 Aug 2026 · Bloss0m · Bloss0m · source
    3. reported

      Large language models cannot retain new information told to them during conversations.

      “Current large models lack continuous learning. You can’t tell them something and expect them to remember.”

      10 Aug 2026 · kucoin.com · source
    4. reported

      Karpathy proposes a small model of one to two billion parameters combined with external memory that can grow over time.

      “ (which he estimates needs only one or two billion parameters) paired with a structured external memory system capable of self-reinforcing growth. Following this idea, he released a model called ”

      10 Aug 2026 · kucoin.com · source
    5. reported

      Current large language models cannot retain information told to them across sessions.

      “have no continual learning. You can't tell it one thing and expect it to remember.”

      10 Aug 2026 · eu.36kr.com · source
    6. reported

      Karpathy proposes a model architecture using a small language model of one to two billion parameters combined with structured external memory that can grow independently.

      “ (he estimates that 1 to 2 billion parameters are enough) plus a set of structured external memory that can compound growth on its own. Based on this idea, he released a pattern called ”

      10 Aug 2026 · eu.36kr.com · source

Andrej Karpathy ontheir other subjects

Everything Andrej Karpathy is on record saying · RSS