[{"data":1,"prerenderedAt":1375},["ShallowReactive",2],{"subject:reinforcement-learning":3,"course-wordcounts":63,"nav:reinforcement-learning":975},{"id":4,"title":5,"blurb":6,"body":7,"brief":14,"category":46,"description":47,"draft":48,"extension":49,"meta":50,"module":11,"navigation":21,"path":51,"practice":52,"rawbody":53,"readingTime":54,"seo":57,"sources":58,"status":59,"stem":60,"summary":11,"topics":61,"__hash__":62},"course\u002F07.reinforcement-learning\u002Findex.md","Reinforcement Learning","Learning to act from reward alone — agents that try, fail, and improve by\ntrial and error, from bandits to AlphaZero.\n",{"type":8,"value":9,"toc":10},"minimark",[],{"title":11,"searchDepth":12,"depth":12,"links":13},"",2,[],[15,17,22,24,28,30,34,36,40,42],{"p":16},"Most machine learning trains on labeled answers. Reinforcement learning\ngets only a number — a reward — and must work out, from its own trial\nand error, which of its actions earned it.\n",{"fig":18,"n":19,"caption":20,"large":21},"rlloop","001","The loop: the agent acts, the environment answers with a state and reward.\n",true,{"p":23},"The setting is an \u003Cstrong>agent\u003C\u002Fstrong> acting in an\n\u003Cstrong>environment\u003C\u002Fstrong>: at each step it observes a state, picks an\naction, and receives a reward and a next state. The goal is not to match a\nlabel but to maximize reward accumulated over the long run — which means\ntrading a sure gain now against a larger one later.\n",{"fig":25,"n":26,"caption":27},"mdp","002","A Markov decision process: states, transitions, and the actions between.\n",{"p":29},"One idea organizes the subject: learn a \u003Cem>value\u003C\u002Fem>, the expected\nlong-run reward of a state, and act greedily toward it. The methods\ndiffer in how they estimate that value: sweeping a known model with\ndynamic programming, averaging sampled returns with Monte Carlo, or\n\u003Cstrong>temporal-difference learning\u003C\u002Fstrong>, which updates each\nestimate toward the next before the episode is even over.\n",{"fig":31,"n":32,"caption":33},"gridworld","003","A gridworld: value spreads from the goal and the greedy policy points home.\n",{"p":35},"Bootstrapping from one estimate to the next is what makes it work\nonline. It is also what the brain's dopamine neurons appear to compute:\na reward-prediction error, the same signal the algorithms are built on.\n",{"fig":37,"n":38,"caption":39},"tdbackup","004","Temporal-difference learning: the prediction error backs up along the chain.\n",{"p":41},"Scale the tables into function approximators and the same equations drive\ndeep reinforcement learning: the Q-network that first played Atari from\npixels, the policy gradients behind modern robot control, and the\nself-play search that took AlphaZero from the rules of Go to superhuman in\na day.\n",{"fig":43,"n":44,"caption":45},"bandit","005","A bandit: pulls sharpen each arm's estimate until the best one wins out.\n","computer-science","The formal problem of learning from interaction: Markov decision processes,\nvalue functions, and the Bellman equations; the three families of tabular\nsolution methods — dynamic programming, Monte Carlo, and temporal-difference\nlearning — and how they unify through bootstrapping and n-step returns;\nplanning with learned models; function approximation, eligibility traces, and\npolicy-gradient methods that scale RL past small tables; deep reinforcement\nlearning from DQN to actor–critic and AlphaZero; and the ties back to\npsychology and the dopamine system. Built on Sutton & Barto. Notes for this\nsubject are coming soon.\n",false,"md",{},"\u002Freinforcement-learning",[],"---\ntitle: Reinforcement Learning\nstatus: available\nblurb: |\n  Learning to act from reward alone — agents that try, fail, and improve by\n  trial and error, from bandits to AlphaZero.\ndescription: |\n  The formal problem of learning from interaction: Markov decision processes,\n  value functions, and the Bellman equations; the three families of tabular\n  solution methods — dynamic programming, Monte Carlo, and temporal-difference\n  learning — and how they unify through bootstrapping and n-step returns;\n  planning with learned models; function approximation, eligibility traces, and\n  policy-gradient methods that scale RL past small tables; deep reinforcement\n  learning from DQN to actor–critic and AlphaZero; and the ties back to\n  psychology and the dopamine system. Built on Sutton & Barto. Notes for this\n  subject are coming soon.\nbrief:\n  - p: |\n      Most machine learning trains on labeled answers. Reinforcement learning\n      gets only a number — a reward — and must work out, from its own trial\n      and error, which of its actions earned it.\n  - fig: rlloop\n    n: \"001\"\n    caption: |\n      The loop: the agent acts, the environment answers with a state and reward.\n    large: true\n  - p: |\n      The setting is an \u003Cstrong>agent\u003C\u002Fstrong> acting in an\n      \u003Cstrong>environment\u003C\u002Fstrong>: at each step it observes a state, picks an\n      action, and receives a reward and a next state. The goal is not to match a\n      label but to maximize reward accumulated over the long run — which means\n      trading a sure gain now against a larger one later.\n  - fig: mdp\n    n: \"002\"\n    caption: |\n      A Markov decision process: states, transitions, and the actions between.\n  - p: |\n      One idea organizes the subject: learn a \u003Cem>value\u003C\u002Fem>, the expected\n      long-run reward of a state, and act greedily toward it. The methods\n      differ in how they estimate that value: sweeping a known model with\n      dynamic programming, averaging sampled returns with Monte Carlo, or\n      \u003Cstrong>temporal-difference learning\u003C\u002Fstrong>, which updates each\n      estimate toward the next before the episode is even over.\n  - fig: gridworld\n    n: \"003\"\n    caption: |\n      A gridworld: value spreads from the goal and the greedy policy points home.\n  - p: |\n      Bootstrapping from one estimate to the next is what makes it work\n      online. It is also what the brain's dopamine neurons appear to compute:\n      a reward-prediction error, the same signal the algorithms are built on.\n  - fig: tdbackup\n    n: \"004\"\n    caption: |\n      Temporal-difference learning: the prediction error backs up along the chain.\n  - p: |\n      Scale the tables into function approximators and the same equations drive\n      deep reinforcement learning: the Q-network that first played Atari from\n      pixels, the policy gradients behind modern robot control, and the\n      self-play search that took AlphaZero from the rules of Go to superhuman in\n      a day.\n  - fig: bandit\n    n: \"005\"\n    caption: |\n      A bandit: pulls sharpen each arm's estimate until the best one wins out.\n---\n",{"text":55,"minutes":56,"time":56,"words":56},"0 min read",0,{"title":5,"description":47},[],"available","07.reinforcement-learning\u002Findex",[],"hy1SnrtieYUMjRWggvkTPcjc9YmRIlgWAyh0534vTig",{"\u002Falgorithms\u002Ffoundations\u002Fwhat-is-an-algorithm":64,"\u002Falgorithms\u002Ffoundations\u002Fproof-techniques":65,"\u002Falgorithms\u002Ffoundations\u002Fasymptotic-analysis":66,"\u002Falgorithms\u002Ffoundations\u002Fgrowth-rates-and-loop-analysis":67,"\u002Falgorithms\u002Ffoundations\u002Frecurrences":68,"\u002Falgorithms\u002Ffoundations\u002Famortized-analysis":69,"\u002Falgorithms\u002Fdivide-and-conquer\u002Fmergesort":70,"\u002Falgorithms\u002Fdivide-and-conquer\u002Fquicksort":71,"\u002Falgorithms\u002Fdivide-and-conquer\u002Fselection":72,"\u002Falgorithms\u002Fdivide-and-conquer\u002Ffast-multiplication":73,"\u002Falgorithms\u002Fsorting\u002Fheaps-and-heapsort":74,"\u002Falgorithms\u002Fsorting\u002Fsorting-lower-bounds":75,"\u002Falgorithms\u002Fsorting\u002Flinear-time-sorting":76,"\u002Falgorithms\u002Fsorting\u002Fexternal-sorting":77,"\u002Falgorithms\u002Fdata-structures\u002Felementary-structures":78,"\u002Falgorithms\u002Fdata-structures\u002Fhash-tables":79,"\u002Falgorithms\u002Fdata-structures\u002Fbinary-search-trees":80,"\u002Falgorithms\u002Fdata-structures\u002Favl-trees":81,"\u002Falgorithms\u002Fdata-structures\u002Fbalanced-trees":82,"\u002Falgorithms\u002Fdata-structures\u002Funion-find":83,"\u002Falgorithms\u002Fdata-structures\u002Ffenwick-and-segment-trees":84,"\u002Falgorithms\u002Fdata-structures\u002Fspatial-data-structures":85,"\u002Falgorithms\u002Fdata-structures\u002Fskip-lists-and-probabilistic-structures":86,"\u002Falgorithms\u002Fdata-structures\u002Fb-trees":87,"\u002Falgorithms\u002Fdata-structures\u002Fdata-stream-algorithms":88,"\u002Falgorithms\u002Fdata-structures\u002Fstreaming-sketches":89,"\u002Falgorithms\u002Fsequences\u002Ftwo-pointers-and-windows":90,"\u002Falgorithms\u002Fsequences\u002Fprefix-sums":91,"\u002Falgorithms\u002Fsequences\u002Fmonotonic-stacks":92,"\u002Falgorithms\u002Fsequences\u002Fbinary-search-on-the-answer":93,"\u002Falgorithms\u002Fsequences\u002Fstring-matching":94,"\u002Falgorithms\u002Fsequences\u002Fkmp-and-z-function":95,"\u002Falgorithms\u002Fsequences\u002Ftries":96,"\u002Falgorithms\u002Fsequences\u002Fsuffix-arrays-and-aho-corasick":97,"\u002Falgorithms\u002Fgraphs\u002Frepresentations-and-traversal":98,"\u002Falgorithms\u002Fgraphs\u002Fdepth-first-search":99,"\u002Falgorithms\u002Fgraphs\u002Ftopological-sort-and-scc":100,"\u002Falgorithms\u002Fgraphs\u002Fminimum-spanning-trees":101,"\u002Falgorithms\u002Fgraphs\u002Fkruskal-and-prim":102,"\u002Falgorithms\u002Fgraphs\u002Fshortest-paths":103,"\u002Falgorithms\u002Fgraphs\u002Fall-pairs-and-negative-weights":104,"\u002Falgorithms\u002Fgraphs\u002Fnetwork-flow":105,"\u002Falgorithms\u002Fgraphs\u002Fmax-flow-min-cut":106,"\u002Falgorithms\u002Fgraphs\u002Fbridges-and-articulation-points":107,"\u002Falgorithms\u002Fgraphs\u002Flowest-common-ancestor":108,"\u002Falgorithms\u002Fgraphs\u002Ftwo-sat":109,"\u002Falgorithms\u002Fgraphs\u002Feulerian-tours":110,"\u002Falgorithms\u002Fgraphs\u002Fbipartite-matching":111,"\u002Falgorithms\u002Fgreedy\u002Fthe-greedy-method":112,"\u002Falgorithms\u002Fgreedy\u002Fscheduling-and-intervals":113,"\u002Falgorithms\u002Fgreedy\u002Fhuffman-codes":114,"\u002Falgorithms\u002Fgreedy\u002Fmatroids":115,"\u002Falgorithms\u002Fgreedy\u002Fstable-matching":116,"\u002Falgorithms\u002Fdynamic-programming\u002Fprinciples":117,"\u002Falgorithms\u002Fdynamic-programming\u002Fsequence-dp":118,"\u002Falgorithms\u002Fdynamic-programming\u002Flongest-increasing-subsequence":119,"\u002Falgorithms\u002Fdynamic-programming\u002Fknapsack":120,"\u002Falgorithms\u002Fdynamic-programming\u002Fcoin-change-and-unbounded":121,"\u002Falgorithms\u002Fdynamic-programming\u002Finterval-dp":122,"\u002Falgorithms\u002Fdynamic-programming\u002Ftree-dp":123,"\u002Falgorithms\u002Fdynamic-programming\u002Fbitmask-dp":124,"\u002Falgorithms\u002Fdynamic-programming\u002Fdp-optimizations":125,"\u002Falgorithms\u002Fdynamic-programming\u002Fdp-on-graphs":126,"\u002Falgorithms\u002Fdynamic-programming\u002Fdigit-and-probability-dp":127,"\u002Falgorithms\u002Fbacktracking\u002Fbacktracking-fundamentals":128,"\u002Falgorithms\u002Fbacktracking\u002Fconstraint-search":129,"\u002Falgorithms\u002Fbacktracking\u002Fbranch-and-bound":130,"\u002Falgorithms\u002Fbacktracking\u002Fgraph-backtracking":131,"\u002Falgorithms\u002Fmathematical-algorithms\u002Fnumber-theory-basics":132,"\u002Falgorithms\u002Fmathematical-algorithms\u002Fmodular-exponentiation-and-primality":133,"\u002Falgorithms\u002Fmathematical-algorithms\u002Fsieve-and-factorization":134,"\u002Falgorithms\u002Fmathematical-algorithms\u002Fcombinatorics":135,"\u002Falgorithms\u002Fmathematical-algorithms\u002Fmatrix-exponentiation":136,"\u002Falgorithms\u002Fmathematical-algorithms\u002Ffast-fourier-transform":137,"\u002Falgorithms\u002Fmathematical-algorithms\u002Fgradient-descent":138,"\u002Falgorithms\u002Fcomputational-geometry\u002Fgeometric-primitives":139,"\u002Falgorithms\u002Fcomputational-geometry\u002Fconvex-hull":140,"\u002Falgorithms\u002Fcomputational-geometry\u002Fsweep-line":141,"\u002Falgorithms\u002Fcomputational-geometry\u002Fpolygons-and-proximity":142,"\u002Falgorithms\u002Fintractability\u002Fp-np-reductions":143,"\u002Falgorithms\u002Fintractability\u002Fnp-completeness":144,"\u002Falgorithms\u002Fintractability\u002Fcoping-with-hardness":145,"\u002Falgorithms\u002Fintractability\u002Fapproximation-algorithms":146,"\u002Falgorithms":147,"\u002Fcalculus\u002Flimits-and-continuity\u002Ffunctions-and-models":148,"\u002Fcalculus\u002Flimits-and-continuity\u002Fthe-limit-of-a-function":149,"\u002Fcalculus\u002Flimits-and-continuity\u002Flimit-laws-and-the-precise-definition":150,"\u002Fcalculus\u002Flimits-and-continuity\u002Fcontinuity":151,"\u002Fcalculus\u002Fderivatives\u002Fthe-derivative-and-rates-of-change":152,"\u002Fcalculus\u002Fderivatives\u002Fdifferentiation-rules-and-the-chain-rule":153,"\u002Fcalculus\u002Fderivatives\u002Fimplicit-differentiation-and-related-rates":154,"\u002Fcalculus\u002Fderivatives\u002Flinear-approximations-and-differentials":155,"\u002Fcalculus\u002Fapplications-of-derivatives\u002Fextrema-and-the-mean-value-theorem":156,"\u002Fcalculus\u002Fapplications-of-derivatives\u002Fhow-derivatives-shape-a-graph":157,"\u002Fcalculus\u002Fapplications-of-derivatives\u002Fcurve-sketching-and-optimization":158,"\u002Fcalculus\u002Fapplications-of-derivatives\u002Fnewtons-method-and-antiderivatives":159,"\u002Fcalculus\u002Fintegrals\u002Farea-and-the-definite-integral":160,"\u002Fcalculus\u002Fintegrals\u002Fthe-fundamental-theorem-of-calculus":161,"\u002Fcalculus\u002Fintegrals\u002Fthe-substitution-rule":162,"\u002Fcalculus\u002Fapplications-of-integration\u002Fareas-and-volumes":163,"\u002Fcalculus\u002Fapplications-of-integration\u002Fwork-average-value-and-arc-length":164,"\u002Fcalculus\u002Fapplications-of-integration\u002Fphysics-economics-and-probability":165,"\u002Fcalculus\u002Fexponential-logarithmic-and-inverse-functions\u002Finverse-functions-logarithms-and-exponentials":166,"\u002Fcalculus\u002Fexponential-logarithmic-and-inverse-functions\u002Fgrowth-decay-inverse-trig-and-hyperbolic-functions":167,"\u002Fcalculus\u002Fexponential-logarithmic-and-inverse-functions\u002Flhospitals-rule":168,"\u002Fcalculus\u002Ftechniques-of-integration\u002Fintegration-by-parts":169,"\u002Fcalculus\u002Ftechniques-of-integration\u002Ftrigonometric-integrals-and-substitution":170,"\u002Fcalculus\u002Ftechniques-of-integration\u002Fpartial-fractions-and-integration-strategy":171,"\u002Fcalculus\u002Ftechniques-of-integration\u002Fapproximate-and-improper-integrals":172,"\u002Fcalculus\u002Fparametric-and-polar\u002Fparametric-curves-and-their-calculus":173,"\u002Fcalculus\u002Fparametric-and-polar\u002Fpolar-coordinates":174,"\u002Fcalculus\u002Fparametric-and-polar\u002Fconic-sections":175,"\u002Fcalculus\u002Fsequences-and-series\u002Fsequences":176,"\u002Fcalculus\u002Fsequences-and-series\u002Fseries-and-the-integral-test":177,"\u002Fcalculus\u002Fsequences-and-series\u002Fthe-convergence-tests":178,"\u002Fcalculus\u002Fsequences-and-series\u002Fpower-series":179,"\u002Fcalculus\u002Fsequences-and-series\u002Ftaylor-and-maclaurin-series":180,"\u002Fcalculus\u002Fvectors-and-space-curves\u002Fvectors-and-the-dot-product":181,"\u002Fcalculus\u002Fvectors-and-space-curves\u002Fthe-cross-product-lines-and-planes":162,"\u002Fcalculus\u002Fvectors-and-space-curves\u002Fcylinders-and-quadric-surfaces":182,"\u002Fcalculus\u002Fvectors-and-space-curves\u002Fvector-functions-and-space-curves":183,"\u002Fcalculus\u002Fvectors-and-space-curves\u002Farc-length-curvature-and-motion":184,"\u002Fcalculus\u002Fpartial-derivatives\u002Ffunctions-of-several-variables":152,"\u002Fcalculus\u002Fpartial-derivatives\u002Fpartial-derivatives":185,"\u002Fcalculus\u002Fpartial-derivatives\u002Ftangent-planes-and-the-chain-rule":186,"\u002Fcalculus\u002Fpartial-derivatives\u002Fdirectional-derivatives-and-the-gradient":187,"\u002Fcalculus\u002Fpartial-derivatives\u002Foptimization-and-lagrange-multipliers":188,"\u002Fcalculus\u002Fmultiple-integrals-and-vector-calculus\u002Fdouble-integrals":189,"\u002Fcalculus\u002Fmultiple-integrals-and-vector-calculus\u002Ftriple-integrals-and-coordinate-systems":190,"\u002Fcalculus\u002Fmultiple-integrals-and-vector-calculus\u002Fvector-fields-and-line-integrals":191,"\u002Fcalculus\u002Fmultiple-integrals-and-vector-calculus\u002Fgreens-theorem-curl-and-divergence":192,"\u002Fcalculus\u002Fmultiple-integrals-and-vector-calculus\u002Fsurface-integrals":193,"\u002Fcalculus\u002Fmultiple-integrals-and-vector-calculus\u002Fstokes-and-the-divergence-theorem":194,"\u002Fcalculus":195,"\u002Fmechanics\u002Ffoundations\u002Fmeasurement-and-dimensions":196,"\u002Fmechanics\u002Ffoundations\u002Fvector-algebra":197,"\u002Fmechanics\u002Fkinematics\u002Fone-dimensional-motion":198,"\u002Fmechanics\u002Fkinematics\u002Fmotion-graphs":199,"\u002Fmechanics\u002Fkinematics\u002Fprojectile-motion":200,"\u002Fmechanics\u002Fkinematics\u002Frelative-motion":201,"\u002Fmechanics\u002Fkinematics\u002Fcircular-motion":202,"\u002Fmechanics\u002Fdynamics\u002Fnewtons-laws":203,"\u002Fmechanics\u002Fdynamics\u002Ffree-body-diagrams":204,"\u002Fmechanics\u002Fdynamics\u002Ffriction-and-curved-motion":205,"\u002Fmechanics\u002Fdynamics\u002Fnumerical-dynamics":206,"\u002Fmechanics\u002Fdynamics\u002Fcenter-of-mass-systems":207,"\u002Fmechanics\u002Fenergy\u002Fwork-and-kinetic-energy":208,"\u002Fmechanics\u002Fenergy\u002Fpotential-energy":209,"\u002Fmechanics\u002Fenergy\u002Fmultiparticle-work":210,"\u002Fmechanics\u002Fenergy\u002Fmass-energy-and-binding":211,"\u002Fmechanics\u002Fenergy\u002Fphotons-and-quantization":212,"\u002Fmechanics\u002Fmomentum\u002Fmomentum-and-collisions":213,"\u002Fmechanics\u002Fmomentum\u002Fcenter-of-mass-collisions":214,"\u002Fmechanics\u002Fmomentum\u002Frocket-propulsion":215,"\u002Fmechanics\u002Frotation\u002Frotational-inertia":216,"\u002Fmechanics\u002Frotation\u002Frotational-dynamics":217,"\u002Fmechanics\u002Frotation\u002Frolling-motion":218,"\u002Fmechanics\u002Frotation\u002Fangular-momentum":219,"\u002Fmechanics\u002Frotation\u002Frolling-resistance":220,"\u002Fmechanics\u002Frotation\u002Fgyroscopic-precession":221,"\u002Fmechanics\u002Fgravity-and-matter\u002Fkeplerian-orbits":222,"\u002Fmechanics\u002Fgravity-and-matter\u002Fgravitational-fields":223,"\u002Fmechanics\u002Fgravity-and-matter\u002Fstatic-equilibrium":224,"\u002Fmechanics\u002Fgravity-and-matter\u002Ffluid-statics":225,"\u002Fmechanics\u002Fgravity-and-matter\u002Ffluid-flow":226,"\u002Fmechanics\u002Fgravity-and-matter\u002Forbital-motion":227,"\u002Fmechanics\u002Fgravity-and-matter\u002Fstress-and-elasticity":228,"\u002Fmechanics\u002Foscillations-waves\u002Fdamped-oscillators":229,"\u002Fmechanics\u002Foscillations-waves\u002Ftravelling-waves":230,"\u002Fmechanics\u002Foscillations-waves\u002Fwave-superposition":231,"\u002Fmechanics\u002Foscillations-waves\u002Fstanding-waves":232,"\u002Fmechanics\u002Foscillations-waves\u002Fsound-waves":233,"\u002Fmechanics\u002Foscillations-waves\u002Fdoppler-effect":234,"\u002Fmechanics\u002Foscillations-waves\u002Fwave-packets":235,"\u002Fmechanics\u002Foscillations-waves\u002Fbeats-and-coupling":236,"\u002Fmechanics\u002Foscillations-waves\u002Fsimple-harmonic-motion":237,"\u002Fmechanics\u002Foscillations-waves\u002Fpendulum-motion":238,"\u002Fmechanics\u002Foscillations-waves\u002Fdriven-oscillators":239,"\u002Fmechanics\u002Foscillations-waves\u002Fwave-boundaries":240,"\u002Fmechanics\u002Fthermodynamics\u002Fkinetic-theory-of-ideal-gases":241,"\u002Fmechanics\u002Fthermodynamics\u002Ffirst-law-of-thermodynamics":242,"\u002Fmechanics\u002Fthermodynamics\u002Fentropy-and-the-second-law":243,"\u002Fmechanics\u002Fthermodynamics\u002Fthermal-processes":244,"\u002Fmechanics\u002Fthermodynamics\u002Fphase-changes":245,"\u002Fmechanics\u002Fthermodynamics\u002Fthermal-machines":246,"\u002Fmechanics":247,"\u002Felectricity-and-magnetism\u002Felectric-fields\u002Fcharge-and-conductors":248,"\u002Felectricity-and-magnetism\u002Felectric-fields\u002Fcoulombs-law":249,"\u002Felectricity-and-magnetism\u002Felectric-fields\u002Felectric-field-and-force":250,"\u002Felectricity-and-magnetism\u002Felectric-fields\u002Felectric-field-maps":251,"\u002Felectricity-and-magnetism\u002Felectric-fields\u002Felectric-dipoles":252,"\u002Felectricity-and-magnetism\u002Fcontinuous-charge-distributions\u002Fcontinuous-charge-fields":253,"\u002Felectricity-and-magnetism\u002Fcontinuous-charge-distributions\u002Fgauss-law-and-conductors":254,"\u002Felectricity-and-magnetism\u002Felectric-potential\u002Fpoint-charge-potential":255,"\u002Felectricity-and-magnetism\u002Felectric-potential\u002Fpotential-gradients-and-equipotentials":256,"\u002Felectricity-and-magnetism\u002Felectric-potential\u002Felectrostatic-energy-and-pressure":257,"\u002Felectricity-and-magnetism\u002Felectric-potential\u002Flaplace-boundary-problems":258,"\u002Felectricity-and-magnetism\u002Felectric-potential\u002Fcontinuous-charge-potentials":259,"\u002Felectricity-and-magnetism\u002Fcapacitance\u002Fcapacitance-fundamentals":236,"\u002Felectricity-and-magnetism\u002Fcapacitance\u002Fcapacitor-networks":260,"\u002Felectricity-and-magnetism\u002Fcapacitance\u002Fcapacitor-energy-and-force":261,"\u002Felectricity-and-magnetism\u002Fcapacitance\u002Fdielectric-polarization-and-breakdown":262,"\u002Felectricity-and-magnetism\u002Fdirect-current-circuits\u002Fcurrent-and-resistance":232,"\u002Felectricity-and-magnetism\u002Fdirect-current-circuits\u002Fkirchhoff-network-analysis":97,"\u002Felectricity-and-magnetism\u002Fdirect-current-circuits\u002Frc-transients":263,"\u002Felectricity-and-magnetism\u002Fmagnetic-field\u002Fmagnetic-trajectories":223,"\u002Felectricity-and-magnetism\u002Fmagnetic-field\u002Fhall-effect":264,"\u002Felectricity-and-magnetism\u002Fmagnetic-field\u002Fmagnetic-force-on-conductors":265,"\u002Felectricity-and-magnetism\u002Fmagnetic-field\u002Fmagnetic-dipoles":266,"\u002Felectricity-and-magnetism\u002Fmagnetic-field\u002Fmass-spectrometry":267,"\u002Felectricity-and-magnetism\u002Fmagnetic-sources\u002Fmoving-charge-fields":268,"\u002Felectricity-and-magnetism\u002Fmagnetic-sources\u002Fbiot-savart-law":269,"\u002Felectricity-and-magnetism\u002Fmagnetic-sources\u002Fcircular-current-loops":270,"\u002Felectricity-and-magnetism\u002Fmagnetic-sources\u002Famperes-law":271,"\u002Felectricity-and-magnetism\u002Fmagnetic-sources\u002Fgauss-law-for-magnetism":272,"\u002Felectricity-and-magnetism\u002Fmagnetic-sources\u002Fmagnetic-materials":197,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Fmagnetic-flux":273,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Ffaradays-law":274,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Flenzs-law":275,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Fmotional-emf":276,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Feddy-currents":277,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Fself-inductance":278,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Fmagnetic-energy":279,"\u002Felectricity-and-magnetism\u002Felectromagnetic-induction\u002Frl-circuits":280,"\u002Felectricity-and-magnetism\u002Falternating-current\u002Fac-fundamentals":215,"\u002Felectricity-and-magnetism\u002Falternating-current\u002Freactance":214,"\u002Felectricity-and-magnetism\u002Falternating-current\u002Frlc-resonance":281,"\u002Felectricity-and-magnetism\u002Falternating-current\u002Fac-power":282,"\u002Felectricity-and-magnetism\u002Falternating-current\u002Ftransformers":283,"\u002Felectricity-and-magnetism\u002Fmaxwell-electromagnetic-waves\u002Fdisplacement-current":284,"\u002Felectricity-and-magnetism\u002Fmaxwell-electromagnetic-waves\u002Felectromagnetic-waves":285,"\u002Felectricity-and-magnetism\u002Fmaxwell-electromagnetic-waves\u002Felectromagnetic-momentum":286,"\u002Felectricity-and-magnetism\u002Fmaxwell-electromagnetic-waves\u002Fdipole-radiation":287,"\u002Felectricity-and-magnetism\u002Fmaxwell-electromagnetic-waves\u002Fpolarization":288,"\u002Felectricity-and-magnetism\u002Foptics\u002Freflection-and-refraction":289,"\u002Felectricity-and-magnetism\u002Foptics\u002Fthin-lenses":241,"\u002Felectricity-and-magnetism\u002Foptics\u002Fspherical-mirrors":239,"\u002Felectricity-and-magnetism":290,"\u002Flinear-algebra\u002Flinear-systems\u002Fsystems-and-echelon-forms":291,"\u002Flinear-algebra\u002Flinear-systems\u002Fvector-and-matrix-equations":292,"\u002Flinear-algebra\u002Flinear-systems\u002Fsolution-sets-and-applications":293,"\u002Flinear-algebra\u002Flinear-systems\u002Flinear-independence":294,"\u002Flinear-algebra\u002Flinear-systems\u002Flinear-transformations":295,"\u002Flinear-algebra\u002Fmatrix-algebra\u002Fmatrix-operations":296,"\u002Flinear-algebra\u002Fmatrix-algebra\u002Fmatrix-inverse-and-invertibility":297,"\u002Flinear-algebra\u002Fmatrix-algebra\u002Fpartitioned-matrices-and-lu":298,"\u002Flinear-algebra\u002Fmatrix-algebra\u002Fsubspaces-dimension-rank":299,"\u002Flinear-algebra\u002Fmatrix-algebra\u002Fapplications-leontief-and-graphics":149,"\u002Flinear-algebra\u002Fdeterminants\u002Fdeterminants-and-cofactors":300,"\u002Flinear-algebra\u002Fdeterminants\u002Fproperties-of-determinants":301,"\u002Flinear-algebra\u002Fdeterminants\u002Fcramer-volume-and-area":153,"\u002Flinear-algebra\u002Fvector-spaces\u002Fvector-spaces-and-subspaces":302,"\u002Flinear-algebra\u002Fvector-spaces\u002Fnull-and-column-spaces":303,"\u002Flinear-algebra\u002Fvector-spaces\u002Fbases-and-independent-sets":304,"\u002Flinear-algebra\u002Fvector-spaces\u002Fcoordinate-systems":305,"\u002Flinear-algebra\u002Fvector-spaces\u002Fdimension-and-rank":306,"\u002Flinear-algebra\u002Fvector-spaces\u002Fchange-of-basis":307,"\u002Flinear-algebra\u002Fvector-spaces\u002Fdifference-equations-and-markov":308,"\u002Flinear-algebra\u002Feigenvalues\u002Feigenvectors-and-eigenvalues":309,"\u002Flinear-algebra\u002Feigenvalues\u002Fthe-characteristic-equation":310,"\u002Flinear-algebra\u002Feigenvalues\u002Fdiagonalization":311,"\u002Flinear-algebra\u002Feigenvalues\u002Feigenvectors-and-linear-transformations":312,"\u002Flinear-algebra\u002Feigenvalues\u002Fcomplex-eigenvalues":313,"\u002Flinear-algebra\u002Feigenvalues\u002Fdynamical-systems":314,"\u002Flinear-algebra\u002Feigenvalues\u002Fpower-method":315,"\u002Flinear-algebra\u002Forthogonality-least-squares\u002Finner-product-length-orthogonality":316,"\u002Flinear-algebra\u002Forthogonality-least-squares\u002Forthogonal-sets-and-projections":317,"\u002Flinear-algebra\u002Forthogonality-least-squares\u002Fgram-schmidt-and-qr":318,"\u002Flinear-algebra\u002Forthogonality-least-squares\u002Fleast-squares-problems":319,"\u002Flinear-algebra\u002Forthogonality-least-squares\u002Fleast-squares-applications":320,"\u002Flinear-algebra\u002Forthogonality-least-squares\u002Finner-product-spaces":321,"\u002Flinear-algebra\u002Fsymmetric-quadratic-svd\u002Fdiagonalizing-symmetric-matrices":188,"\u002Flinear-algebra\u002Fsymmetric-quadratic-svd\u002Fquadratic-forms":322,"\u002Flinear-algebra\u002Fsymmetric-quadratic-svd\u002Fconstrained-optimization":323,"\u002Flinear-algebra\u002Fsymmetric-quadratic-svd\u002Fsingular-value-decomposition":324,"\u002Flinear-algebra\u002Fsymmetric-quadratic-svd\u002Fsvd-applications-pca-imaging":325,"\u002Flinear-algebra\u002Fnumerical-linear-algebra\u002Fnumerical-thinking-and-matrix-computation":326,"\u002Flinear-algebra\u002Fnumerical-linear-algebra\u002Flu-and-cholesky":327,"\u002Flinear-algebra\u002Fnumerical-linear-algebra\u002Fconditioning-and-floating-point":328,"\u002Flinear-algebra\u002Fnumerical-linear-algebra\u002Fstability-and-error-analysis":329,"\u002Flinear-algebra\u002Fnumerical-linear-algebra\u002Fqr-and-numerical-least-squares":330,"\u002Flinear-algebra\u002Fnumerical-linear-algebra\u002Fnumerical-eigenvalues-and-svd":331,"\u002Flinear-algebra\u002Fgeometry-of-vector-spaces\u002Faffine-combinations":332,"\u002Flinear-algebra\u002Fgeometry-of-vector-spaces\u002Faffine-independence-and-barycentric-coordinates":333,"\u002Flinear-algebra\u002Fgeometry-of-vector-spaces\u002Fconvex-combinations-and-convex-sets":334,"\u002Flinear-algebra\u002Fgeometry-of-vector-spaces\u002Fhyperplanes-and-polytopes":335,"\u002Flinear-algebra\u002Fgeometry-of-vector-spaces\u002Fcurves-and-surfaces":336,"\u002Flinear-algebra":337,"\u002Ftheory-of-computation":56,"\u002Fcomputer-architecture\u002Ffoundations\u002Fbits-bytes-and-words":338,"\u002Fcomputer-architecture\u002Ffoundations\u002Finteger-representation":339,"\u002Fcomputer-architecture\u002Ffoundations\u002Finteger-arithmetic":340,"\u002Fcomputer-architecture\u002Ffoundations\u002Ffloating-point":341,"\u002Fcomputer-architecture\u002Ffoundations\u002Fboolean-algebra-and-bit-manipulation":342,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Fthe-machines-view":343,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Fdata-movement":344,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Farithmetic-and-logic":345,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Fcontrol-flow":346,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Fprocedures":347,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Farrays-structs-and-alignment":348,"\u002Fcomputer-architecture\u002Fmachine-level-x86-64\u002Fmemory-layout-and-buffer-overflows":349,"\u002Fcomputer-architecture\u002Finstruction-set-architecture\u002Fwhat-an-isa-is":350,"\u002Fcomputer-architecture\u002Finstruction-set-architecture\u002Finstruction-formats-and-operands":351,"\u002Fcomputer-architecture\u002Finstruction-set-architecture\u002Faddressing-modes":352,"\u002Fcomputer-architecture\u002Finstruction-set-architecture\u002Fthe-y86-64-instruction-set":353,"\u002Fcomputer-architecture\u002Finstruction-set-architecture\u002Fy86-64-programming":354,"\u002Fcomputer-architecture\u002Fdigital-logic\u002Ftransistors-gates-and-boolean-functions":355,"\u002Fcomputer-architecture\u002Fdigital-logic\u002Fcombinational-logic-and-hcl":356,"\u002Fcomputer-architecture\u002Fdigital-logic\u002Fmultiplexers-decoders-and-the-alu":357,"\u002Fcomputer-architecture\u002Fdigital-logic\u002Fmemory-elements-latches-flip-flops-and-clocking":358,"\u002Fcomputer-architecture\u002Fdigital-logic\u002Fregister-files-and-random-access-memory":359,"\u002Fcomputer-architecture\u002Fprocessor-design\u002Fthe-fetch-decode-execute-cycle":360,"\u002Fcomputer-architecture\u002Fprocessor-design\u002Fthe-seq-stages":361,"\u002Fcomputer-architecture\u002Fprocessor-design\u002Fcontrol-logic-and-sequencing":362,"\u002Fcomputer-architecture\u002Fprocessor-design\u002Fassembling-seq":363,"\u002Fcomputer-architecture\u002Fprocessor-design\u002Ftracing-a-program":364,"\u002Fcomputer-architecture\u002Fpipelining\u002Fpipelining-principles":365,"\u002Fcomputer-architecture\u002Fpipelining\u002Ffrom-seq-to-pipe":366,"\u002Fcomputer-architecture\u002Fpipelining\u002Fdata-hazards-stalling-and-forwarding":367,"\u002Fcomputer-architecture\u002Fpipelining\u002Fcontrol-hazards-and-branch-prediction":368,"\u002Fcomputer-architecture\u002Fpipelining\u002Fthe-complete-pipe-processor":369,"\u002Fcomputer-architecture\u002Fmemory-hierarchy\u002Fstorage-technologies-and-the-latency-gap":370,"\u002Fcomputer-architecture\u002Fmemory-hierarchy\u002Flocality":371,"\u002Fcomputer-architecture\u002Fmemory-hierarchy\u002Fcache-memories-direct-mapped":372,"\u002Fcomputer-architecture\u002Fmemory-hierarchy\u002Fset-associative-and-write-policies":373,"\u002Fcomputer-architecture\u002Fmemory-hierarchy\u002Fcache-performance-and-cache-friendly-code":374,"\u002Fcomputer-architecture\u002Fvirtual-memory\u002Faddress-spaces-and-translation":375,"\u002Fcomputer-architecture\u002Fvirtual-memory\u002Fpage-tables-and-page-faults":376,"\u002Fcomputer-architecture\u002Fvirtual-memory\u002Fthe-tlb-and-multi-level-page-tables":377,"\u002Fcomputer-architecture\u002Fexceptions-and-io\u002Fexceptional-control-flow":378,"\u002Fcomputer-architecture\u002Fexceptions-and-io\u002Finterrupts-and-the-kernel":379,"\u002Fcomputer-architecture\u002Fmultithreading-and-multicore\u002Fprocesses-threads-and-parallelism":380,"\u002Fcomputer-architecture\u002Fmultithreading-and-multicore\u002Fhardware-multithreading":381,"\u002Fcomputer-architecture\u002Fmultithreading-and-multicore\u002Fcache-coherence":382,"\u002Fcomputer-architecture\u002Fmultithreading-and-multicore\u002Fmemory-consistency-and-synchronization":383,"\u002Fcomputer-architecture\u002Fmultithreading-and-multicore\u002Fmulticore-organization":384,"\u002Fcomputer-architecture\u002Fcapstone\u002Fthe-whole-machine":385,"\u002Fcomputer-architecture\u002Fcapstone\u002Fassembling-a-complete-cpu":386,"\u002Fcomputer-architecture":56,"\u002Fdifferential-equations\u002Ffoundations\u002Fmodels-and-direction-fields":387,"\u002Fdifferential-equations\u002Ffoundations\u002Fclassification-and-terminology":388,"\u002Fdifferential-equations\u002Ffirst-order\u002Flinear-first-order-integrating-factors":389,"\u002Fdifferential-equations\u002Ffirst-order\u002Fseparable-and-exact":153,"\u002Fdifferential-equations\u002Ffirst-order\u002Fmodeling-first-order":390,"\u002Fdifferential-equations\u002Ffirst-order\u002Fautonomous-and-population-dynamics":152,"\u002Fdifferential-equations\u002Ffirst-order\u002Fexistence-uniqueness-euler":159,"\u002Fdifferential-equations\u002Ffirst-order\u002Ffirst-order-difference-equations":391,"\u002Fdifferential-equations\u002Fsecond-order-linear\u002Fhomogeneous-constant-coefficients":392,"\u002Fdifferential-equations\u002Fsecond-order-linear\u002Fcomplex-and-repeated-roots":193,"\u002Fdifferential-equations\u002Fsecond-order-linear\u002Fnonhomogeneous-undetermined-coefficients":393,"\u002Fdifferential-equations\u002Fsecond-order-linear\u002Fvariation-of-parameters":394,"\u002Fdifferential-equations\u002Fsecond-order-linear\u002Fmechanical-electrical-vibrations":395,"\u002Fdifferential-equations\u002Fsecond-order-linear\u002Fhigher-order-linear":396,"\u002Fdifferential-equations\u002Fseries-solutions\u002Fpower-series-ordinary-points":397,"\u002Fdifferential-equations\u002Fseries-solutions\u002Fregular-singular-frobenius":398,"\u002Fdifferential-equations\u002Fseries-solutions\u002Fbessel-and-special-functions":399,"\u002Fdifferential-equations\u002Flaplace\u002Flaplace-definition-ivps":400,"\u002Fdifferential-equations\u002Flaplace\u002Fstep-impulse-convolution":401,"\u002Fdifferential-equations\u002Fsystems\u002Fmatrices-eigenvalues-review":402,"\u002Fdifferential-equations\u002Fsystems\u002Fconstant-coefficient-systems-phase-portraits":403,"\u002Fdifferential-equations\u002Fsystems\u002Frepeated-eigenvalues-fundamental-matrices":404,"\u002Fdifferential-equations\u002Fnumerical\u002Feuler-and-runge-kutta":399,"\u002Fdifferential-equations\u002Fnumerical\u002Fmultistep-systems-stability":405,"\u002Fdifferential-equations\u002Fnonlinear\u002Fphase-plane-autonomous-stability":406,"\u002Fdifferential-equations\u002Fnonlinear\u002Flocally-linear-and-liapunov":407,"\u002Fdifferential-equations\u002Fnonlinear\u002Fcompeting-species-predator-prey-limit-cycles":408,"\u002Fdifferential-equations\u002Fpdes-fourier-bvp\u002Ffourier-series":409,"\u002Fdifferential-equations\u002Fpdes-fourier-bvp\u002Fheat-wave-laplace-equations":410,"\u002Fdifferential-equations\u002Fpdes-fourier-bvp\u002Fsturm-liouville":411,"\u002Fdifferential-equations\u002Fhistory-variations\u002Fcalculus-of-variations":412,"\u002Fdifferential-equations\u002Fhistory-variations\u002Fhistorical-notes":413,"\u002Fdifferential-equations":414,"\u002Frelativity\u002Ffoundations\u002Fspecial-relativity-postulates":415,"\u002Frelativity\u002Ffoundations\u002Florentz-transformation-spacetime":416,"\u002Frelativity\u002Ffoundations\u002Ftime-dilation-length-contraction":417,"\u002Frelativity\u002Ffoundations\u002Frelativistic-momentum-energy":418,"\u002Frelativity\u002Ffoundations\u002Fgeneral-relativity":298,"\u002Frelativity\u002Fspacetime-and-the-lorentz-group\u002Fminkowski-spacetime-and-the-interval":419,"\u002Frelativity\u002Fspacetime-and-the-lorentz-group\u002Ffour-vectors-and-index-notation":420,"\u002Frelativity\u002Fspacetime-and-the-lorentz-group\u002Fthe-lorentz-group-and-rapidity":421,"\u002Frelativity\u002Fspacetime-and-the-lorentz-group\u002Fdoppler-aberration-and-appearance":422,"\u002Frelativity\u002Frelativistic-dynamics\u002Ffour-momentum-force-and-accelerated-motion":423,"\u002Frelativity\u002Frelativistic-dynamics\u002Fparticle-decays-and-two-body-kinematics":424,"\u002Frelativity\u002Frelativistic-dynamics\u002Fcollisions-thresholds-and-the-cm-frame":168,"\u002Frelativity\u002Frelativistic-dynamics\u002Fmandelstam-variables-and-invariants":425,"\u002Frelativity\u002Fcovariant-electrodynamics\u002Ffour-current-and-the-four-potential":426,"\u002Frelativity\u002Fcovariant-electrodynamics\u002Fthe-electromagnetic-field-tensor":427,"\u002Frelativity\u002Fcovariant-electrodynamics\u002Ftransformation-of-electric-and-magnetic-fields":428,"\u002Frelativity\u002Fcovariant-electrodynamics\u002Fcovariant-maxwell-and-the-stress-energy-tensor":429,"\u002Frelativity\u002Fcurved-spacetime\u002Fthe-equivalence-principle-formalized":430,"\u002Frelativity\u002Fcurved-spacetime\u002Fmanifolds-vectors-and-the-metric":431,"\u002Frelativity\u002Fcurved-spacetime\u002Fcovariant-derivative-and-christoffel-symbols":432,"\u002Frelativity\u002Fcurved-spacetime\u002Fgeodesics-and-the-geodesic-equation":433,"\u002Frelativity\u002Fcurved-spacetime\u002Fcurvature-riemann-and-geodesic-deviation":434,"\u002Frelativity\u002Fcurved-spacetime\u002Fthe-einstein-field-equations":389,"\u002Frelativity\u002Fthe-schwarzschild-solution\u002Fthe-schwarzschild-metric":435,"\u002Frelativity\u002Fthe-schwarzschild-solution\u002Fgeodesics-and-orbits-in-schwarzschild":436,"\u002Frelativity\u002Fthe-schwarzschild-solution\u002Flight-bending-and-null-geodesics":437,"\u002Frelativity\u002Ftests-of-general-relativity\u002Fperihelion-precession-of-mercury":438,"\u002Frelativity\u002Ftests-of-general-relativity\u002Fdeflection-of-light-and-gravitational-lensing":439,"\u002Frelativity\u002Ftests-of-general-relativity\u002Fgravitational-redshift-and-shapiro-delay":319,"\u002Frelativity\u002Ftests-of-general-relativity\u002Frelativity-in-technology-gps":440,"\u002Frelativity\u002Fblack-holes\u002Fhorizons-and-coordinate-singularities":441,"\u002Frelativity\u002Fblack-holes\u002Frotating-and-charged-black-holes":329,"\u002Frelativity\u002Fblack-holes\u002Fblack-hole-thermodynamics":442,"\u002Frelativity\u002Fgravitational-waves\u002Flinearized-gravity-and-wave-solutions":443,"\u002Frelativity\u002Fgravitational-waves\u002Fgeneration-and-the-quadrupole-formula":444,"\u002Frelativity\u002Fgravitational-waves\u002Fdetection-ligo-and-the-first-events":445,"\u002Frelativity\u002Fcosmological-bridge\u002Fthe-cosmological-principle-and-flrw-metric":446,"\u002Frelativity\u002Fcosmological-bridge\u002Ffriedmann-equations-and-cosmic-dynamics":447,"\u002Frelativity":448,"\u002Fphysical-computing":56,"\u002Fquantum-mechanics\u002Fold-quantum-theory\u002Fblackbody-radiation-and-the-planck-quantum":449,"\u002Fquantum-mechanics\u002Fold-quantum-theory\u002Fthe-photoelectric-effect-and-the-photon":428,"\u002Fquantum-mechanics\u002Fold-quantum-theory\u002Fx-rays-and-the-compton-effect":450,"\u002Fquantum-mechanics\u002Fold-quantum-theory\u002Fthe-old-quantum-theory-bohr-and-sommerfeld":451,"\u002Fquantum-mechanics\u002Fmatter-waves\u002Fde-broglie-waves-and-electron-diffraction":452,"\u002Fquantum-mechanics\u002Fmatter-waves\u002Fwave-packets-and-the-probability-interpretation":453,"\u002Fquantum-mechanics\u002Fmatter-waves\u002Fthe-uncertainty-principle":454,"\u002Fquantum-mechanics\u002Fwave-mechanics-1d\u002Fthe-schrodinger-equation-in-one-dimension":455,"\u002Fquantum-mechanics\u002Fwave-mechanics-1d\u002Fthe-free-particle-and-wave-packet-dynamics":456,"\u002Fquantum-mechanics\u002Fwave-mechanics-1d\u002Fparticle-in-infinite-and-finite-square-wells":404,"\u002Fquantum-mechanics\u002Fwave-mechanics-1d\u002Foperators-expectation-values-and-the-harmonic-oscillator":329,"\u002Fquantum-mechanics\u002Fwave-mechanics-1d\u002Fthe-dirac-delta-potential":457,"\u002Fquantum-mechanics\u002Fwave-mechanics-1d\u002Fbarrier-penetration-and-quantum-tunneling":458,"\u002Fquantum-mechanics\u002Fformalism\u002Fhilbert-space-and-dirac-notation":459,"\u002Fquantum-mechanics\u002Fformalism\u002Fobservables-hermitian-operators-and-eigenvalues":460,"\u002Fquantum-mechanics\u002Fformalism\u002Fthe-postulates-and-quantum-measurement":456,"\u002Fquantum-mechanics\u002Fformalism\u002Fposition-momentum-and-continuous-spectra":461,"\u002Fquantum-mechanics\u002Fformalism\u002Fcommutators-and-the-generalized-uncertainty-principle":436,"\u002Fquantum-mechanics\u002Fformalism\u002Ftime-evolution-schrodinger-and-heisenberg-pictures":190,"\u002Fquantum-mechanics\u002Foscillator-and-symmetry\u002Fladder-operators-and-the-number-states":462,"\u002Fquantum-mechanics\u002Foscillator-and-symmetry\u002Fcoherent-and-squeezed-states":463,"\u002Fquantum-mechanics\u002Foscillator-and-symmetry\u002Fsymmetries-generators-and-conservation-laws":170,"\u002Fquantum-mechanics\u002Foscillator-and-symmetry\u002Fparity-time-reversal-and-discrete-symmetries":464,"\u002Fquantum-mechanics\u002Fangular-momentum\u002Forbital-angular-momentum-and-spherical-harmonics":465,"\u002Fquantum-mechanics\u002Fangular-momentum\u002Fthe-angular-momentum-algebra":466,"\u002Fquantum-mechanics\u002Fangular-momentum\u002Faddition-of-angular-momenta-and-clebsch-gordan":467,"\u002Fquantum-mechanics\u002Fcentral-potentials\u002Fthe-schrodinger-equation-in-three-dimensions":468,"\u002Fquantum-mechanics\u002Fcentral-potentials\u002Fthe-hydrogen-atom":469,"\u002Fquantum-mechanics\u002Fcentral-potentials\u002Fthe-isotropic-oscillator-and-hidden-symmetry":470,"\u002Fquantum-mechanics\u002Fspin\u002Fspin-half-pauli-matrices-and-stern-gerlach":471,"\u002Fquantum-mechanics\u002Fspin\u002Fspin-in-a-magnetic-field-precession-and-resonance":472,"\u002Fquantum-mechanics\u002Fspin\u002Ftwo-level-systems-and-the-bloch-sphere":428,"\u002Fquantum-mechanics\u002Fidentical-particles\u002Fidentical-particles-and-exchange-symmetry":473,"\u002Fquantum-mechanics\u002Fidentical-particles\u002Fthe-pauli-principle-atoms-and-the-periodic-table":474,"\u002Fquantum-mechanics\u002Fapproximation-methods\u002Ftime-independent-perturbation-theory":475,"\u002Fquantum-mechanics\u002Fapproximation-methods\u002Ffine-structure-and-the-real-hydrogen-atom":462,"\u002Fquantum-mechanics\u002Fapproximation-methods\u002Fthe-zeeman-and-stark-effects":151,"\u002Fquantum-mechanics\u002Fapproximation-methods\u002Fthe-variational-method":476,"\u002Fquantum-mechanics\u002Fapproximation-methods\u002Fthe-wkb-approximation":477,"\u002Fquantum-mechanics":478,"\u002Freal-analysis\u002Ffoundations\u002Fsets-logic-functions":412,"\u002Freal-analysis\u002Ffoundations\u002Fordered-fields-completeness":479,"\u002Freal-analysis\u002Ffoundations\u002Fabsolute-value-bounds":480,"\u002Freal-analysis\u002Ffoundations\u002Fintervals-uncountability":302,"\u002Freal-analysis\u002Fsequences-series\u002Fsequences-limits":481,"\u002Freal-analysis\u002Fsequences-series\u002Flimit-laws-monotone":188,"\u002Freal-analysis\u002Fsequences-series\u002Flimsup-bolzano-weierstrass":482,"\u002Freal-analysis\u002Fsequences-series\u002Fcauchy-completeness":483,"\u002Freal-analysis\u002Fsequences-series\u002Fseries-convergence":325,"\u002Freal-analysis\u002Fsequences-series\u002Fabsolute-conditional-rearrangement":434,"\u002Freal-analysis\u002Fmetric-spaces\u002Fmetric-spaces-norms":484,"\u002Freal-analysis\u002Fmetric-spaces\u002Fopen-closed-sets":485,"\u002Freal-analysis\u002Fmetric-spaces\u002Fconvergence-completeness":486,"\u002Freal-analysis\u002Fmetric-spaces\u002Fcompactness":487,"\u002Freal-analysis\u002Fmetric-spaces\u002Fconnectedness":488,"\u002Freal-analysis\u002Fcontinuity\u002Flimits-of-functions":461,"\u002Freal-analysis\u002Fcontinuity\u002Fcontinuous-functions":489,"\u002Freal-analysis\u002Fcontinuity\u002Fevt-ivt":295,"\u002Freal-analysis\u002Fcontinuity\u002Funiform-continuity":490,"\u002Freal-analysis\u002Fcontinuity\u002Fcontinuity-metric-spaces":491,"\u002Freal-analysis\u002Fcontinuity\u002Flimits-infinity-monotone":149,"\u002Freal-analysis\u002Fdifferentiation\u002Fthe-derivative":492,"\u002Freal-analysis\u002Fdifferentiation\u002Fmean-value-theorem":493,"\u002Freal-analysis\u002Fdifferentiation\u002Ftaylors-theorem":450,"\u002Freal-analysis\u002Fdifferentiation\u002Finverse-function-1d":178,"\u002Freal-analysis\u002Friemann-integration\u002Fdarboux-integral":329,"\u002Freal-analysis\u002Friemann-integration\u002Fintegrability-classes":494,"\u002Freal-analysis\u002Friemann-integration\u002Fproperties-of-the-integral":495,"\u002Freal-analysis\u002Friemann-integration\u002Ffundamental-theorem":314,"\u002Freal-analysis\u002Friemann-integration\u002Flog-exp-improper":438,"\u002Freal-analysis\u002Ffunction-sequences\u002Fpointwise-uniform-convergence":496,"\u002Freal-analysis\u002Ffunction-sequences\u002Finterchange-of-limits":497,"\u002Freal-analysis\u002Ffunction-sequences\u002Fpower-series-weierstrass":498,"\u002Freal-analysis\u002Ffunction-sequences\u002Fpicard-ode":334,"\u002Freal-analysis\u002Fseveral-variables\u002Fdifferentiability-rn":499,"\u002Freal-analysis\u002Fseveral-variables\u002Fgradient-chain-rule":500,"\u002Freal-analysis\u002Fseveral-variables\u002Fhigher-derivatives-taylor-extrema":501,"\u002Freal-analysis\u002Fseveral-variables\u002Finverse-implicit-theorems":501,"\u002Freal-analysis\u002Fseveral-variables\u002Fmultiple-integrals":502,"\u002Freal-analysis":503,"\u002Fabstract-algebra\u002Ffoundations\u002Fsets-functions-relations":504,"\u002Fabstract-algebra\u002Ffoundations\u002Fintegers-and-modular-arithmetic":505,"\u002Fabstract-algebra\u002Fgroups-and-symmetry\u002Fgroup-axioms-and-first-examples":506,"\u002Fabstract-algebra\u002Fgroups-and-symmetry\u002Fdihedral-and-symmetric-groups":507,"\u002Fabstract-algebra\u002Fgroups-and-symmetry\u002Fmatrix-and-quaternion-groups":508,"\u002Fabstract-algebra\u002Fgroups-and-symmetry\u002Fhomomorphisms-and-group-actions":509,"\u002Fabstract-algebra\u002Fsubgroups-and-quotients\u002Fsubgroups-and-substructures":510,"\u002Fabstract-algebra\u002Fsubgroups-and-quotients\u002Fcyclic-groups":511,"\u002Fabstract-algebra\u002Fsubgroups-and-quotients\u002Fgeneration-and-subgroup-lattices":512,"\u002Fabstract-algebra\u002Fsubgroups-and-quotients\u002Fcosets-lagrange-and-normal-subgroups":513,"\u002Fabstract-algebra\u002Fsubgroups-and-quotients\u002Fisomorphism-theorems":477,"\u002Fabstract-algebra\u002Fsubgroups-and-quotients\u002Fcomposition-series-and-the-alternating-group":514,"\u002Fabstract-algebra\u002Fgroup-actions-and-sylow\u002Factions-and-cayleys-theorem":506,"\u002Fabstract-algebra\u002Fgroup-actions-and-sylow\u002Fconjugation-and-the-class-equation":423,"\u002Fabstract-algebra\u002Fgroup-actions-and-sylow\u002Fsylow-theorems":515,"\u002Fabstract-algebra\u002Fgroup-actions-and-sylow\u002Fautomorphisms-and-simple-groups":516,"\u002Fabstract-algebra\u002Fproducts-and-group-structure\u002Fdirect-products-and-finite-abelian-groups":517,"\u002Fabstract-algebra\u002Fproducts-and-group-structure\u002Fsemidirect-products":518,"\u002Fabstract-algebra\u002Fproducts-and-group-structure\u002Fnilpotent-and-solvable-groups":519,"\u002Fabstract-algebra\u002Fproducts-and-group-structure\u002Fclassifying-small-groups":520,"\u002Fabstract-algebra\u002Fring-theory\u002Frings-definitions-and-examples":521,"\u002Fabstract-algebra\u002Fring-theory\u002Fideals-quotients-and-homomorphisms":522,"\u002Fabstract-algebra\u002Fring-theory\u002Ffractions-and-the-chinese-remainder-theorem":516,"\u002Fabstract-algebra\u002Ffactorization-and-polynomials\u002Feuclidean-domains-pids-ufds":523,"\u002Fabstract-algebra\u002Ffactorization-and-polynomials\u002Fpolynomial-rings-over-fields":492,"\u002Fabstract-algebra\u002Ffactorization-and-polynomials\u002Fgauss-lemma-and-unique-factorization":524,"\u002Fabstract-algebra\u002Ffactorization-and-polynomials\u002Firreducibility-criteria-and-groebner":525,"\u002Fabstract-algebra\u002Fmodule-theory\u002Fintroduction-to-modules":526,"\u002Fabstract-algebra\u002Fmodule-theory\u002Ffree-modules-and-direct-sums":527,"\u002Fabstract-algebra\u002Fmodule-theory\u002Ftensor-products-and-exact-sequences":528,"\u002Fabstract-algebra\u002Fmodule-theory\u002Fvector-spaces-and-linear-maps":529,"\u002Fabstract-algebra\u002Fmodules-over-pids\u002Fstructure-theorem-over-pids":530,"\u002Fabstract-algebra\u002Fmodules-over-pids\u002Frational-canonical-form":531,"\u002Fabstract-algebra\u002Fmodules-over-pids\u002Fjordan-canonical-form":532,"\u002Fabstract-algebra\u002Ffield-theory\u002Ffield-extensions-and-algebraic-elements":533,"\u002Fabstract-algebra\u002Ffield-theory\u002Fstraightedge-and-compass-constructions":167,"\u002Fabstract-algebra\u002Ffield-theory\u002Fsplitting-fields-and-algebraic-closure":534,"\u002Fabstract-algebra\u002Ffield-theory\u002Fseparable-and-cyclotomic-extensions":535,"\u002Fabstract-algebra\u002Fgalois-theory\u002Fthe-galois-correspondence":408,"\u002Fabstract-algebra\u002Fgalois-theory\u002Ffinite-fields":536,"\u002Fabstract-algebra\u002Fgalois-theory\u002Fcyclotomic-and-abelian-extensions":537,"\u002Fabstract-algebra\u002Fgalois-theory\u002Fgalois-groups-of-polynomials":465,"\u002Fabstract-algebra\u002Fgalois-theory\u002Fsolvability-by-radicals-and-the-quintic":537,"\u002Fabstract-algebra\u002Fcapstone\u002Fcommutative-algebra-and-algebraic-geometry":538,"\u002Fabstract-algebra\u002Fcapstone\u002Frepresentation-and-character-theory":539,"\u002Fabstract-algebra":540,"\u002Fatomic-physics\u002Fearly-models-and-old-quantum-theory\u002Fatomic-spectra-rutherford":541,"\u002Fatomic-physics\u002Fearly-models-and-old-quantum-theory\u002Fbohr-model-hydrogen":542,"\u002Fatomic-physics\u002Fearly-models-and-old-quantum-theory\u002Fx-ray-spectra-franck-hertz":543,"\u002Fatomic-physics\u002Fearly-models-and-old-quantum-theory\u002Fbohr-sommerfeld-old-quantum-theory":544,"\u002Fatomic-physics\u002Fearly-models-and-old-quantum-theory\u002Fold-quantum-theory-limits-wkb":545,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Fschrodinger-3d-hydrogen":457,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Fhydrogen-wave-functions":546,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Fradial-equation-in-full":547,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Fsymmetry-degeneracy-runge-lenz":548,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Fexpectation-values-virial":549,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Fquantum-defects-alkali-spectra":550,"\u002Fatomic-physics\u002Fquantum-hydrogen-atom\u002Frydberg-atoms":551,"\u002Fatomic-physics\u002Ffine-structure-and-the-dirac-atom\u002Frelativistic-kinetic-correction":552,"\u002Fatomic-physics\u002Ffine-structure-and-the-dirac-atom\u002Fspin-orbit-thomas-precession":292,"\u002Fatomic-physics\u002Ffine-structure-and-the-dirac-atom\u002Fdarwin-term-fine-structure-formula":420,"\u002Fatomic-physics\u002Ffine-structure-and-the-dirac-atom\u002Fdirac-equation-hydrogen":167,"\u002Fatomic-physics\u002Fqed-corrections-and-hyperfine-structure\u002Flamb-shift-qed":553,"\u002Fatomic-physics\u002Fqed-corrections-and-hyperfine-structure\u002Fhyperfine-structure-21cm":151,"\u002Fatomic-physics\u002Fqed-corrections-and-hyperfine-structure\u002Fnuclear-effects-isotope-shift":554,"\u002Fatomic-physics\u002Fmany-electron-atoms\u002Fperiodic-table-atomic-spectra":555,"\u002Fatomic-physics\u002Fmany-electron-atoms\u002Fcentral-field-self-consistent":189,"\u002Fatomic-physics\u002Fmany-electron-atoms\u002Fidentical-particles-hartree-fock":482,"\u002Fatomic-physics\u002Fmany-electron-atoms\u002Fhelium-two-electron-atom":556,"\u002Fatomic-physics\u002Fmany-electron-atoms\u002Fls-jj-coupling-term-symbols":557,"\u002Fatomic-physics\u002Fmany-electron-atoms\u002Fhund-rules-ground-terms":558,"\u002Fatomic-physics\u002Fatoms-in-external-fields\u002Fzeeman-effect":559,"\u002Fatomic-physics\u002Fatoms-in-external-fields\u002Fpaschen-back-intermediate":560,"\u002Fatomic-physics\u002Fatoms-in-external-fields\u002Fstark-effect-polarizability":561,"\u002Fatomic-physics\u002Fradiative-transitions-and-line-shapes\u002Ftime-dependent-perturbation-golden-rule":562,"\u002Fatomic-physics\u002Fradiative-transitions-and-line-shapes\u002Fdipole-approximation-einstein-coefficients":563,"\u002Fatomic-physics\u002Fradiative-transitions-and-line-shapes\u002Fselection-rules-forbidden-transitions":564,"\u002Fatomic-physics\u002Fradiative-transitions-and-line-shapes\u002Flifetimes-and-line-shapes":565,"\u002Fatomic-physics\u002Flasers-and-spectroscopy\u002Flaser-principles":566,"\u002Fatomic-physics\u002Flasers-and-spectroscopy\u002Fspectroscopy-techniques":567,"\u002Fatomic-physics\u002Flasers-and-spectroscopy\u002Fline-catalog-nist-asd":568,"\u002Fatomic-physics\u002Fmodern-atomic-physics\u002Flaser-cooling-doppler":569,"\u002Fatomic-physics\u002Fmodern-atomic-physics\u002Fsub-doppler-trapping":513,"\u002Fatomic-physics\u002Fmodern-atomic-physics\u002Fbose-einstein-condensation":570,"\u002Fatomic-physics\u002Fmodern-atomic-physics\u002Foptical-clocks-precision":571,"\u002Fatomic-physics":572,"\u002Fdatabases":56,"\u002Fcategory-theory\u002Ffoundations\u002Fwhat-is-a-category":573,"\u002Fcategory-theory\u002Ffoundations\u002Fexamples-of-categories":574,"\u002Fcategory-theory\u002Ffoundations\u002Fspecial-morphisms":575,"\u002Fcategory-theory\u002Ffoundations\u002Ffunctors":504,"\u002Fcategory-theory\u002Ffoundations\u002Fnatural-transformations":576,"\u002Fcategory-theory\u002Ffoundations\u002Fsize-and-set-theory":577,"\u002Fcategory-theory\u002Funiversal-properties\u002Funiversal-properties":578,"\u002Fcategory-theory\u002Funiversal-properties\u002Fproducts-and-coproducts":579,"\u002Fcategory-theory\u002Funiversal-properties\u002Fconstructions-on-categories":580,"\u002Fcategory-theory\u002Frepresentables-yoneda\u002Frepresentable-functors":581,"\u002Fcategory-theory\u002Frepresentables-yoneda\u002Fyoneda-lemma":582,"\u002Fcategory-theory\u002Frepresentables-yoneda\u002Fyoneda-consequences":583,"\u002Fcategory-theory\u002Flimits-colimits\u002Flimits":584,"\u002Fcategory-theory\u002Flimits-colimits\u002Fproducts-equalizers-pullbacks":585,"\u002Fcategory-theory\u002Flimits-colimits\u002Fcolimits":586,"\u002Fcategory-theory\u002Flimits-colimits\u002Fcomputing-limits":587,"\u002Fcategory-theory\u002Flimits-colimits\u002Flimits-and-functors":588,"\u002Fcategory-theory\u002Fadjunctions\u002Fadjunctions":589,"\u002Fcategory-theory\u002Fadjunctions\u002Funits-and-counits":590,"\u002Fcategory-theory\u002Fadjunctions\u002Fadjunctions-via-universal-arrows":591,"\u002Fcategory-theory\u002Fadjunctions\u002Ffree-forgetful-adjunctions":592,"\u002Fcategory-theory\u002Fadjoints-limits\u002Flimits-via-adjoints":593,"\u002Fcategory-theory\u002Fadjoints-limits\u002Fpresheaf-limits-colimits":594,"\u002Fcategory-theory\u002Fadjoints-limits\u002Fadjoints-preserve-limits":595,"\u002Fcategory-theory\u002Fadjoints-limits\u002Fadjoint-functor-theorem":586,"\u002Fcategory-theory\u002Fmonads-algebras\u002Fmonads":596,"\u002Fcategory-theory\u002Fmonads-algebras\u002Falgebras-eilenberg-moore":597,"\u002Fcategory-theory\u002Fmonads-algebras\u002Fkleisli-and-programming":598,"\u002Fcategory-theory\u002Fmonads-algebras\u002Falgebras-for-endofunctors":599,"\u002Fcategory-theory\u002Fcartesian-closed-lambda\u002Fcartesian-closed-categories":600,"\u002Fcategory-theory\u002Fcartesian-closed-lambda\u002Flambda-calculus-correspondence":545,"\u002Fcategory-theory\u002Fcartesian-closed-lambda\u002Ffixed-points-and-recursion":601,"\u002Fcategory-theory":602,"\u002Fdeep-learning\u002Fmathematical-background\u002Flinear-algebra-for-deep-learning":603,"\u002Fdeep-learning\u002Fmathematical-background\u002Fprobability-and-information-theory":604,"\u002Fdeep-learning\u002Fmathematical-background\u002Fnumerical-computation":605,"\u002Fdeep-learning\u002Fmathematical-background\u002Fcalculus":606,"\u002Fdeep-learning\u002Ffoundations\u002Fwhat-is-deep-learning":607,"\u002Fdeep-learning\u002Ffoundations\u002Fmachine-learning-refresher":608,"\u002Fdeep-learning\u002Ffoundations\u002Flinear-models-and-the-perceptron":567,"\u002Fdeep-learning\u002Fneural-networks\u002Fthe-multilayer-perceptron":609,"\u002Fdeep-learning\u002Fneural-networks\u002Factivation-functions":610,"\u002Fdeep-learning\u002Fneural-networks\u002Funiversal-approximation":611,"\u002Fdeep-learning\u002Fneural-networks\u002Fbackpropagation":612,"\u002Fdeep-learning\u002Fneural-networks\u002Floss-functions-and-output-units":613,"\u002Fdeep-learning\u002Foptimization\u002Fgradient-descent-and-sgd":614,"\u002Fdeep-learning\u002Foptimization\u002Fmomentum-and-adaptive-methods":615,"\u002Fdeep-learning\u002Foptimization\u002Finitialization":616,"\u002Fdeep-learning\u002Foptimization\u002Fthe-optimization-landscape":617,"\u002Fdeep-learning\u002Foptimization\u002Fsecond-order-and-approximate-methods":618,"\u002Fdeep-learning\u002Fregularization\u002Fregularization-overview":619,"\u002Fdeep-learning\u002Fregularization\u002Fdropout-and-data-augmentation":620,"\u002Fdeep-learning\u002Fregularization\u002Fearly-stopping-and-parameter-sharing":621,"\u002Fdeep-learning\u002Fregularization\u002Fnormalization":622,"\u002Fdeep-learning\u002Farchitectures\u002Fconvolutional-networks":623,"\u002Fdeep-learning\u002Farchitectures\u002Fcnn-architectures":624,"\u002Fdeep-learning\u002Farchitectures\u002Frecurrent-networks":625,"\u002Fdeep-learning\u002Farchitectures\u002Flstm-and-gru":626,"\u002Fdeep-learning\u002Farchitectures\u002Fattention-and-transformers":627,"\u002Fdeep-learning\u002Farchitectures\u002Fthe-transformer-architecture":628,"\u002Fdeep-learning\u002Farchitectures\u002Ftransformers-in-practice":629,"\u002Fdeep-learning\u002Farchitectures\u002Fgraph-neural-networks":630,"\u002Fdeep-learning\u002Farchitectures\u002Fstate-space-models":631,"\u002Fdeep-learning\u002Ftheory\u002Fgeneralization-theory":632,"\u002Fdeep-learning\u002Ftheory\u002Fadversarial-robustness":633,"\u002Fdeep-learning\u002Ftheory\u002Fadversarial-defenses":634,"\u002Fdeep-learning\u002Ftheory\u002Fbayesian-and-ensemble-methods":635,"\u002Fdeep-learning\u002Ftheory\u002Fdeep-equilibrium-models":579,"\u002Fdeep-learning\u002Fgenerative-models\u002Flinear-factor-models":636,"\u002Fdeep-learning\u002Fgenerative-models\u002Fautoencoders":637,"\u002Fdeep-learning\u002Fgenerative-models\u002Fvariational-autoencoders":638,"\u002Fdeep-learning\u002Fgenerative-models\u002Fgenerative-adversarial-networks":639,"\u002Fdeep-learning\u002Fgenerative-models\u002Fautoregressive-and-normalizing-flows":640,"\u002Fdeep-learning\u002Fgenerative-models\u002Fenergy-based-and-boltzmann-machines":641,"\u002Fdeep-learning\u002Fgenerative-models\u002Fdiffusion-and-score-based-models":642,"\u002Fdeep-learning\u002Fprobabilistic-methods\u002Fstructured-probabilistic-models":132,"\u002Fdeep-learning\u002Fprobabilistic-methods\u002Fmonte-carlo-and-mcmc":643,"\u002Fdeep-learning\u002Fprobabilistic-methods\u002Fapproximate-inference":644,"\u002Fdeep-learning\u002Fpractical\u002Fpractical-methodology":367,"\u002Fdeep-learning\u002Fpractical\u002Fhyperparameters-and-debugging":645,"\u002Fdeep-learning\u002Fpractical\u002Frepresentation-learning":646,"\u002Fdeep-learning\u002Fpractical\u002Ftransfer-learning":647,"\u002Fdeep-learning\u002Fpractical\u002Fapplications":648,"\u002Fdeep-learning\u002Fpractical\u002Fmodel-compression-and-distillation":649,"\u002Fdeep-learning\u002Fpractical\u002Fmeta-learning-and-few-shot":650,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Flarge-language-models":651,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fscaling-inference-and-alignment":652,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fseq2seq-pretraining-and-bart":653,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Ftext-to-text-transfer-and-conditional-generation":654,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fspeech-and-audio-models":655,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fself-supervised-speech-and-synthesis":656,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fai-agents":340,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fagent-memory-retrieval-and-orchestration":657,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fmixture-of-experts":658,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Fmultimodal-models":659,"\u002Fdeep-learning\u002Flarge-models-and-agents\u002Ffusion-and-vision-language-models":660,"\u002Fdeep-learning\u002Freinforcement-learning\u002Ffoundations-of-reinforcement-learning":377,"\u002Fdeep-learning\u002Freinforcement-learning\u002Fmodel-free-prediction-and-control":661,"\u002Fdeep-learning\u002Freinforcement-learning\u002Fdeep-q-networks":662,"\u002Fdeep-learning\u002Freinforcement-learning\u002Fpolicy-gradients-and-actor-critic":663,"\u002Fdeep-learning\u002Freinforcement-learning\u002Frl-from-human-feedback":664,"\u002Fdeep-learning":56,"\u002Fstatistical-mechanics\u002Fthermodynamics\u002Fequilibrium-state-variables-zeroth-law":665,"\u002Fstatistical-mechanics\u002Fthermodynamics\u002Ffirst-law-heat-and-work":462,"\u002Fstatistical-mechanics\u002Fthermodynamics\u002Fsecond-law-entropy-and-the-carnot-bound":666,"\u002Fstatistical-mechanics\u002Fthermodynamics\u002Fthermodynamic-potentials-and-maxwell-relations":667,"\u002Fstatistical-mechanics\u002Fthermodynamics\u002Fstability-response-functions-and-the-third-law":668,"\u002Fstatistical-mechanics\u002Ffoundations\u002Fclassical-statistics-and-equipartition":669,"\u002Fstatistical-mechanics\u002Ffoundations\u002Fphase-space-and-liouvilles-theorem":670,"\u002Fstatistical-mechanics\u002Ffoundations\u002Fensembles-and-the-equal-probability-postulate":671,"\u002Fstatistical-mechanics\u002Ffoundations\u002Fstatistical-entropy-boltzmann-and-gibbs":672,"\u002Fstatistical-mechanics\u002Fmicrocanonical\u002Fmicrocanonical-ensemble-and-entropy":673,"\u002Fstatistical-mechanics\u002Fmicrocanonical\u002Fequilibrium-conditions-temperature-pressure-chemical-potential":509,"\u002Fstatistical-mechanics\u002Fmicrocanonical\u002Fideal-gas-phase-space-and-the-sackur-tetrode-entropy":674,"\u002Fstatistical-mechanics\u002Fmicrocanonical\u002Ftwo-state-systems-paramagnets-and-negative-temperature":675,"\u002Fstatistical-mechanics\u002Fcanonical\u002Fcanonical-ensemble-and-the-boltzmann-distribution":676,"\u002Fstatistical-mechanics\u002Fcanonical\u002Fpartition-function-and-the-helmholtz-free-energy":396,"\u002Fstatistical-mechanics\u002Fcanonical\u002Fenergy-fluctuations-and-ensemble-equivalence":190,"\u002Fstatistical-mechanics\u002Fcanonical\u002Fthe-einstein-solid-and-harmonic-systems":677,"\u002Fstatistical-mechanics\u002Fcanonical\u002Fparamagnetism-and-the-schottky-anomaly":678,"\u002Fstatistical-mechanics\u002Fclassical-gas\u002Fideal-gas-partition-function-and-the-gibbs-paradox":679,"\u002Fstatistical-mechanics\u002Fclassical-gas\u002Fequipartition-and-the-virial-theorem":324,"\u002Fstatistical-mechanics\u002Fclassical-gas\u002Fmolecular-gases-rotation-and-vibration":680,"\u002Fstatistical-mechanics\u002Fgrand-canonical\u002Fgrand-canonical-ensemble-and-the-grand-partition-function":681,"\u002Fstatistical-mechanics\u002Fgrand-canonical\u002Fchemical-potential-fugacity-and-number-fluctuations":401,"\u002Fstatistical-mechanics\u002Fgrand-canonical\u002Fensemble-summary-and-the-thermodynamic-web":329,"\u002Fstatistical-mechanics\u002Fquantum-statistics\u002Fquantum-statistics-bose-einstein-and-fermi-dirac":682,"\u002Fstatistical-mechanics\u002Fquantum-statistics\u002Fderiving-the-quantum-distributions":298,"\u002Fstatistical-mechanics\u002Fquantum-statistics\u002Fthe-classical-limit-and-quantum-concentration":168,"\u002Fstatistical-mechanics\u002Fquantum-statistics\u002Fideal-quantum-gases-general-framework":184,"\u002Fstatistical-mechanics\u002Fbose-systems\u002Fbose-einstein-condensation-and-the-fermion-gas":683,"\u002Fstatistical-mechanics\u002Fbose-systems\u002Fthe-photon-gas-and-plancks-radiation-law":684,"\u002Fstatistical-mechanics\u002Fbose-systems\u002Fblackbody-thermodynamics-and-radiation-pressure":685,"\u002Fstatistical-mechanics\u002Fbose-systems\u002Fphonons-and-the-debye-model":316,"\u002Fstatistical-mechanics\u002Fbose-systems\u002Fbose-einstein-condensation-derived":686,"\u002Fstatistical-mechanics\u002Fbose-systems\u002Fthermodynamics-of-the-bose-gas-and-superfluidity":178,"\u002Fstatistical-mechanics\u002Ffermi-gas\u002Fthe-ideal-fermi-gas-at-zero-temperature":687,"\u002Fstatistical-mechanics\u002Ffermi-gas\u002Fsommerfeld-expansion-and-electrons-in-metals":688,"\u002Fstatistical-mechanics\u002Ffermi-gas\u002Fwhite-dwarfs-and-the-chandrasekhar-limit":689,"\u002Fstatistical-mechanics\u002Ffermi-gas\u002Fneutron-stars-and-nuclear-matter":690,"\u002Fstatistical-mechanics\u002Finteractions\u002Fthe-cluster-expansion-and-virial-coefficients":511,"\u002Fstatistical-mechanics\u002Finteractions\u002Fthe-van-der-waals-gas-and-liquid-gas-coexistence":691,"\u002Fstatistical-mechanics\u002Finteractions\u002Fquantum-gases-with-interactions-and-exchange":692,"\u002Fstatistical-mechanics\u002Fphase-transitions\u002Fphases-coexistence-and-classification":693,"\u002Fstatistical-mechanics\u002Fphase-transitions\u002Fthe-ising-model-and-exact-solutions":694,"\u002Fstatistical-mechanics\u002Fphase-transitions\u002Fmean-field-theory-and-the-weiss-model":457,"\u002Fstatistical-mechanics\u002Fphase-transitions\u002Fcritical-exponents-and-landau-theory":695,"\u002Fstatistical-mechanics\u002Fphase-transitions\u002Fthe-renormalization-group-idea":435,"\u002Fstatistical-mechanics\u002Ffluctuations\u002Fthermodynamic-fluctuations-and-response":696,"\u002Fstatistical-mechanics\u002Ffluctuations\u002Fbrownian-motion-and-the-langevin-equation":297,"\u002Fstatistical-mechanics\u002Ffluctuations\u002Flinear-response-and-the-fluctuation-dissipation-theorem":697,"\u002Fstatistical-mechanics":698,"\u002Fcondensed-matter\u002Fmolecules-and-bonding\u002Fbonding-mechanisms":699,"\u002Fcondensed-matter\u002Fmolecules-and-bonding\u002Fmolecular-orbitals-and-h2-plus":163,"\u002Fcondensed-matter\u002Fmolecules-and-bonding\u002Fhydrogen-molecule-and-exchange":422,"\u002Fcondensed-matter\u002Fmolecules-and-bonding\u002Fvan-der-waals-forces":700,"\u002Fcondensed-matter\u002Fmolecular-spectra\u002Frotational-vibrational-spectra":701,"\u002Fcondensed-matter\u002Fmolecular-spectra\u002Fanharmonicity-and-rovibrational-structure":702,"\u002Fcondensed-matter\u002Fmolecular-spectra\u002Framan-and-electronic-bands":703,"\u002Fcondensed-matter\u002Fmolecular-spectra\u002Flasers-and-masers":704,"\u002Fcondensed-matter\u002Fcrystal-structure\u002Fstructure-of-solids":705,"\u002Fcondensed-matter\u002Fcrystal-structure\u002Fbravais-lattices-and-crystal-systems":296,"\u002Fcondensed-matter\u002Fcrystal-structure\u002Freciprocal-lattice-and-brillouin-zones":706,"\u002Fcondensed-matter\u002Fcrystal-structure\u002Fdiffraction-and-structure-factors":707,"\u002Fcondensed-matter\u002Flattice-dynamics\u002Fphonon-dispersion":708,"\u002Fcondensed-matter\u002Flattice-dynamics\u002Fphonons-quantization-and-dos":709,"\u002Fcondensed-matter\u002Flattice-dynamics\u002Fdebye-einstein-heat-capacity":438,"\u002Fcondensed-matter\u002Flattice-dynamics\u002Fanharmonicity-and-thermal-transport":710,"\u002Fcondensed-matter\u002Ffree-electron-fermi-gas\u002Ffree-electron-gas-and-conduction":711,"\u002Fcondensed-matter\u002Ffree-electron-fermi-gas\u002Fsommerfeld-model-and-heat-capacity":712,"\u002Fcondensed-matter\u002Ffree-electron-fermi-gas\u002Ftransport-and-the-hall-effect":713,"\u002Fcondensed-matter\u002Ffree-electron-fermi-gas\u002Fscreening-and-plasmons":714,"\u002Fcondensed-matter\u002Fband-theory\u002Fblochs-theorem-and-energy-bands":493,"\u002Fcondensed-matter\u002Fband-theory\u002Fnearly-free-electron-model":437,"\u002Fcondensed-matter\u002Fband-theory\u002Ftight-binding-method":715,"\u002Fcondensed-matter\u002Fband-theory\u002Ffermi-surfaces-and-semiclassical-dynamics":716,"\u002Fcondensed-matter\u002Fsemiconductors\u002Fsemiconductor-bands-and-junctions":717,"\u002Fcondensed-matter\u002Fsemiconductors\u002Fintrinsic-and-extrinsic-semiconductors":718,"\u002Fcondensed-matter\u002Fsemiconductors\u002Fcarrier-transport-and-recombination":428,"\u002Fcondensed-matter\u002Fsemiconductors\u002Fthe-pn-junction":719,"\u002Fcondensed-matter\u002Fsemiconductors\u002Ftransistors-and-optoelectronics":720,"\u002Fcondensed-matter\u002Fdielectrics-and-ferroelectrics\u002Fdielectrics-and-polarization":665,"\u002Fcondensed-matter\u002Fdielectrics-and-ferroelectrics\u002Fferroelectrics-and-piezoelectrics":565,"\u002Fcondensed-matter\u002Fmagnetism\u002Fdiamagnetism-and-paramagnetism":293,"\u002Fcondensed-matter\u002Fmagnetism\u002Fexchange-and-ferromagnetism":721,"\u002Fcondensed-matter\u002Fmagnetism\u002Fantiferromagnetism-and-domains":159,"\u002Fcondensed-matter\u002Fmagnetism\u002Fspin-waves-and-magnons":722,"\u002Fcondensed-matter\u002Fsuperconductivity\u002Fsuperconductivity-phenomenology":723,"\u002Fcondensed-matter\u002Fsuperconductivity\u002Flondon-theory-and-the-meissner-effect":304,"\u002Fcondensed-matter\u002Fsuperconductivity\u002Fginzburg-landau-theory":724,"\u002Fcondensed-matter\u002Fsuperconductivity\u002Fbcs-theory":558,"\u002Fcondensed-matter\u002Fsuperconductivity\u002Fjosephson-and-high-tc":725,"\u002Fcondensed-matter\u002Fnanostructures\u002Fquantum-wells-wires-and-dots":151,"\u002Fcondensed-matter\u002Fnanostructures\u002Finteger-quantum-hall-effect":726,"\u002Fcondensed-matter\u002Fnanostructures\u002Ffractional-quantum-hall-and-topology":162,"\u002Fcondensed-matter\u002Fnanostructures\u002Fgraphene-and-dirac-materials":727,"\u002Fcondensed-matter":478,"\u002Flogic\u002Ffoundations\u002Flogic-as-a-mathematical-model":728,"\u002Flogic\u002Fsentential-logic\u002Fformal-languages-and-well-formed-formulas":729,"\u002Flogic\u002Fsentential-logic\u002Ftruth-assignments-and-tautologies":730,"\u002Flogic\u002Fsentential-logic\u002Funique-readability-and-parsing":731,"\u002Flogic\u002Fsentential-logic\u002Finduction-and-recursion":176,"\u002Flogic\u002Fsentential-logic\u002Fexpressive-completeness-and-normal-forms":732,"\u002Flogic\u002Fsentential-logic\u002Fboolean-circuits":733,"\u002Flogic\u002Fsentential-logic\u002Fcompactness-and-effectiveness":176,"\u002Flogic\u002Ffirst-order-languages\u002Ffirst-order-languages":734,"\u002Flogic\u002Ffirst-order-languages\u002Fstructures-truth-and-satisfaction":588,"\u002Flogic\u002Ffirst-order-languages\u002Fdefinability-and-elementary-equivalence":735,"\u002Flogic\u002Ffirst-order-languages\u002Fterms-substitution-and-parsing":736,"\u002Flogic\u002Fdeductive-calculus\u002Fa-deductive-calculus":737,"\u002Flogic\u002Fdeductive-calculus\u002Fdeduction-theorem-and-derived-rules":735,"\u002Flogic\u002Fdeductive-calculus\u002Fsoundness":738,"\u002Flogic\u002Fdeductive-calculus\u002Fcompleteness-and-consistency":739,"\u002Flogic\u002Fmodels-and-theories\u002Fcompactness-and-lowenheim-skolem":740,"\u002Flogic\u002Fmodels-and-theories\u002Ftheories-elementary-classes-and-categoricity":741,"\u002Flogic\u002Fmodels-and-theories\u002Finterpretations-between-theories":742,"\u002Flogic\u002Fmodels-and-theories\u002Fnonstandard-analysis":743,"\u002Flogic\u002Farithmetic-and-definability\u002Fdefinability-in-arithmetic":744,"\u002Flogic\u002Farithmetic-and-definability\u002Fnatural-numbers-with-successor":745,"\u002Flogic\u002Farithmetic-and-definability\u002Fpresburger-and-reducts":666,"\u002Flogic\u002Farithmetic-and-definability\u002Fa-subtheory-and-representability":746,"\u002Flogic\u002Fincompleteness\u002Farithmetization-of-syntax":739,"\u002Flogic\u002Fincompleteness\u002Fincompleteness-and-undecidability":747,"\u002Flogic\u002Fincompleteness\u002Fsecond-incompleteness-theorem":748,"\u002Flogic\u002Fcomputability-and-representability\u002Frecursive-functions":381,"\u002Flogic\u002Fcomputability-and-representability\u002Frepresenting-exponentiation":749,"\u002Flogic\u002Fsecond-order-logic\u002Fsecond-order-languages":566,"\u002Flogic\u002Fsecond-order-logic\u002Fskolem-functions-and-many-sorted-logic":750,"\u002Flogic\u002Fsecond-order-logic\u002Fgeneral-structures":751,"\u002Flogic":752,"\u002Freinforcement-learning\u002Ffoundations\u002Fwhat-is-reinforcement-learning":753,"\u002Freinforcement-learning\u002Ffoundations\u002Fa-brief-history-of-rl":754,"\u002Freinforcement-learning\u002Ffoundations\u002Fmulti-armed-bandits":367,"\u002Freinforcement-learning\u002Ffoundations\u002Fbandit-exploration-algorithms":755,"\u002Freinforcement-learning\u002Ffoundations\u002Fmarkov-decision-processes":756,"\u002Freinforcement-learning\u002Ffoundations\u002Fvalue-functions-and-optimality":757,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fdynamic-programming":758,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fdp-async-and-gpi":748,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fmonte-carlo-methods":759,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fmonte-carlo-off-policy":760,"\u002Freinforcement-learning\u002Ftabular-methods\u002Ftemporal-difference-learning":761,"\u002Freinforcement-learning\u002Ftabular-methods\u002Ftd-control-sarsa-and-q-learning":659,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fn-step-bootstrapping":762,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fn-step-off-policy-methods":763,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fplanning-and-learning":764,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fplanning-focusing-and-decision-time":765,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fdecision-time-planning":766,"\u002Freinforcement-learning\u002Ftabular-methods\u002Fmonte-carlo-tree-search":384,"\u002Freinforcement-learning\u002Fapproximation\u002Fon-policy-prediction":767,"\u002Freinforcement-learning\u002Fapproximation\u002Ffeature-construction-and-nonlinear":768,"\u002Freinforcement-learning\u002Fapproximation\u002Fon-policy-control":769,"\u002Freinforcement-learning\u002Fapproximation\u002Faverage-reward-control":770,"\u002Freinforcement-learning\u002Fapproximation\u002Foff-policy-and-the-deadly-triad":575,"\u002Freinforcement-learning\u002Fapproximation\u002Fbellman-error-and-gradient-td":771,"\u002Freinforcement-learning\u002Fapproximation\u002Feligibility-traces":747,"\u002Freinforcement-learning\u002Fapproximation\u002Ftrue-online-and-sarsa-lambda":772,"\u002Freinforcement-learning\u002Fapproximation\u002Fpolicy-gradient-methods":773,"\u002Freinforcement-learning\u002Fapproximation\u002Factor-critic-and-continuous-actions":774,"\u002Freinforcement-learning\u002Fapproximation\u002Fleast-squares-and-memory-based-methods":395,"\u002Freinforcement-learning\u002Fapproximation\u002Fmemory-and-kernel-methods":775,"\u002Freinforcement-learning\u002Fapproximation\u002Foff-policy-eligibility-traces":525,"\u002Freinforcement-learning\u002Fapproximation\u002Fstable-off-policy-traces":776,"\u002Freinforcement-learning\u002Fdeep-rl\u002Fdeep-q-networks":777,"\u002Freinforcement-learning\u002Fdeep-rl\u002Fdqn-improvements":357,"\u002Freinforcement-learning\u002Fdeep-rl\u002Factor-critic-and-ppo":778,"\u002Freinforcement-learning\u002Fdeep-rl\u002Fppo-and-continuous-control":779,"\u002Freinforcement-learning\u002Fdeep-rl\u002Fcase-studies":780,"\u002Freinforcement-learning\u002Fdeep-rl\u002Frl-beyond-games":781,"\u002Freinforcement-learning\u002Fdeep-rl\u002Ffrontiers":782,"\u002Freinforcement-learning\u002Fdeep-rl\u002Freward-design-and-open-problems":635,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fdistributional-and-rainbow":783,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fdistributional-and-rainbow-part-2":784,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fcontinuous-control":785,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fcontinuous-control-part-2":670,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmodel-based-rl":786,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmodel-based-rl-part-2":787,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fexploration":788,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fexploration-part-2":384,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Foffline-rl":451,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Foffline-rl-part-2":789,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fimitation-and-inverse-rl":790,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fimitation-and-inverse-rl-part-2":791,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmulti-agent-rl":792,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmulti-agent-rl-part-2":793,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fhierarchical-rl":794,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fhierarchical-rl-part-2":795,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Frlhf-and-language-models":796,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fpartial-observability-pomdps":797,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fpartial-observability-pomdps-part-2":798,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fsafe-and-constrained-rl":799,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fsafe-and-constrained-rl-part-2":800,"\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmeta-rl-and-generalization":801,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fpsychology-of-reinforcement":802,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Finstrumental-conditioning-and-control":803,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fdopamine-and-td-error":804,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fdopamine-in-the-brain":805,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fanimal-learning-and-cognition":806,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fcognitive-maps-and-planning":807,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fneuroscience-of-reinforcement":808,"\u002Freinforcement-learning\u002Fminds-and-brains\u002Fseveral-learning-systems":809,"\u002Freinforcement-learning":56,"\u002Fartificial-intelligence\u002Ffoundations\u002Fwhat-is-ai":810,"\u002Fartificial-intelligence\u002Ffoundations\u002Ffoundations-of-ai":811,"\u002Fartificial-intelligence\u002Ffoundations\u002Fintelligent-agents":812,"\u002Fartificial-intelligence\u002Ffoundations\u002Fagent-architectures":813,"\u002Fartificial-intelligence\u002Fsearch\u002Funinformed-search":814,"\u002Fartificial-intelligence\u002Fsearch\u002Fsearch-strategies-compared":815,"\u002Fartificial-intelligence\u002Fsearch\u002Finformed-search":816,"\u002Fartificial-intelligence\u002Fsearch\u002Fheuristic-functions":817,"\u002Fartificial-intelligence\u002Fsearch\u002Flocal-search":818,"\u002Fartificial-intelligence\u002Fsearch\u002Fpopulation-and-continuous-search":819,"\u002Fartificial-intelligence\u002Fsearch\u002Fadversarial-search":820,"\u002Fartificial-intelligence\u002Fsearch\u002Fgames-of-chance-and-imperfect-information":821,"\u002Fartificial-intelligence\u002Fsearch\u002Fconstraint-satisfaction":822,"\u002Fartificial-intelligence\u002Fsearch\u002Fcsp-search-and-structure":664,"\u002Fartificial-intelligence\u002Fsearch\u002Fsearch-under-uncertainty":519,"\u002Fartificial-intelligence\u002Fsearch\u002Fbelief-state-and-online-search":823,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fpropositional-logic":824,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fpropositional-inference":825,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Ffirst-order-logic":826,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Ffirst-order-logic-in-use":827,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Finference-and-resolution":828,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Ffirst-order-resolution":646,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fclassical-planning":829,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fplanning-graphs-and-graphplan":830,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fplanning-in-the-real-world":831,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fplanning-under-uncertainty":832,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Fknowledge-representation":833,"\u002Fartificial-intelligence\u002Flogic-and-planning\u002Freasoning-systems-and-defaults":834,"\u002Fartificial-intelligence\u002Funcertainty\u002Fprobability-and-bayes":835,"\u002Fartificial-intelligence\u002Funcertainty\u002Fbayes-rule-and-naive-bayes":836,"\u002Fartificial-intelligence\u002Funcertainty\u002Fbayesian-networks":837,"\u002Fartificial-intelligence\u002Funcertainty\u002Finference-in-bayesian-networks":838,"\u002Fartificial-intelligence\u002Funcertainty\u002Freasoning-over-time":839,"\u002Fartificial-intelligence\u002Funcertainty\u002Ftracking-and-data-association":840,"\u002Fartificial-intelligence\u002Funcertainty\u002Fmaking-decisions":578,"\u002Fartificial-intelligence\u002Funcertainty\u002Fmarkov-decision-processes":830,"\u002Fartificial-intelligence\u002Funcertainty\u002Fdecision-networks-and-game-theory":841,"\u002Fartificial-intelligence\u002Funcertainty\u002Fgame-theory-and-mechanism-design":119,"\u002Fartificial-intelligence\u002Flearning\u002Flearning-from-examples":842,"\u002Fartificial-intelligence\u002Flearning\u002Ftheory-and-model-families":843,"\u002Fartificial-intelligence\u002Flearning\u002Fprobabilistic-learning":844,"\u002Fartificial-intelligence\u002Flearning\u002Fexpectation-maximization":845,"\u002Fartificial-intelligence\u002Flearning\u002Freinforcement-learning":846,"\u002Fartificial-intelligence\u002Flearning\u002Fgeneralization-and-policy-search":627,"\u002Fartificial-intelligence\u002Flearning\u002Fknowledge-in-learning":847,"\u002Fartificial-intelligence\u002Flearning\u002Fknowledge-based-learning-methods":848,"\u002Fartificial-intelligence\u002Ffrontiers\u002Fvision-and-perception":849,"\u002Fartificial-intelligence\u002Ffrontiers\u002Freconstructing-the-3d-world":850,"\u002Fartificial-intelligence\u002Ffrontiers\u002Frobotics":851,"\u002Fartificial-intelligence\u002Ffrontiers\u002Frobot-planning-and-control":852,"\u002Fartificial-intelligence\u002Ffrontiers\u002Fnatural-language-in-ai":853,"\u002Fartificial-intelligence\u002Ffrontiers\u002Fnlp-grammar-translation-and-speech":854,"\u002Fartificial-intelligence\u002Ffrontiers\u002Fphilosophy-and-future":855,"\u002Fartificial-intelligence\u002Ffrontiers\u002Fai-ethics-and-future":856,"\u002Fartificial-intelligence":56,"\u002Fnuclear-physics\u002Fnuclear-properties\u002Fnuclear-constituents-nuclide-chart":594,"\u002Fnuclear-physics\u002Fnuclear-properties\u002Fnuclear-size-charge-distributions":857,"\u002Fnuclear-physics\u002Fnuclear-properties\u002Fnuclear-masses-binding-energy":155,"\u002Fnuclear-physics\u002Fnuclear-properties\u002Fsemi-empirical-mass-formula":153,"\u002Fnuclear-physics\u002Fnuclear-properties\u002Fnuclear-moments-multipoles":688,"\u002Fnuclear-physics\u002Fnuclear-force-deuteron\u002Fnuclear-force-shell-overview":858,"\u002Fnuclear-physics\u002Fnuclear-force-deuteron\u002Fthe-deuteron":181,"\u002Fnuclear-physics\u002Fnuclear-force-deuteron\u002Fnucleon-nucleon-scattering":490,"\u002Fnuclear-physics\u002Fnuclear-force-deuteron\u002Fmeson-theory-isospin":859,"\u002Fnuclear-physics\u002Fnuclear-models\u002Ffermi-gas-model":860,"\u002Fnuclear-physics\u002Fnuclear-models\u002Fliquid-drop-collective-coordinates":861,"\u002Fnuclear-physics\u002Fnuclear-models\u002Fshell-model-single-particle":550,"\u002Fnuclear-physics\u002Fnuclear-models\u002Fcollective-model-rotations-vibrations":862,"\u002Fnuclear-physics\u002Fradioactive-decay\u002Fdecay-law-modes":863,"\u002Fnuclear-physics\u002Fradioactive-decay\u002Fdecay-kinetics-equilibrium":864,"\u002Fnuclear-physics\u002Falpha-decay\u002Falpha-decay-gamow-theory":749,"\u002Fnuclear-physics\u002Falpha-decay\u002Falpha-fine-structure-hindrance":865,"\u002Fnuclear-physics\u002Fbeta-decay\u002Fbeta-decay-energetics-neutrino":866,"\u002Fnuclear-physics\u002Fbeta-decay\u002Ffermi-theory-beta-decay":165,"\u002Fnuclear-physics\u002Fbeta-decay\u002Fweak-interaction-parity-violation":393,"\u002Fnuclear-physics\u002Fbeta-decay\u002Fdouble-beta-decay-neutrino-mass":867,"\u002Fnuclear-physics\u002Fgamma-decay\u002Fgamma-multipole-radiation":494,"\u002Fnuclear-physics\u002Fgamma-decay\u002Finternal-conversion-isomers":868,"\u002Fnuclear-physics\u002Fgamma-decay\u002Fangular-correlations-mossbauer":869,"\u002Fnuclear-physics\u002Fnuclear-reactions\u002Freaction-kinematics-cross-sections":389,"\u002Fnuclear-physics\u002Fnuclear-reactions\u002Fcompound-nucleus-resonances":477,"\u002Fnuclear-physics\u002Fnuclear-reactions\u002Fdirect-reactions-optical-model":870,"\u002Fnuclear-physics\u002Ffission\u002Ffission-barrier-dynamics":871,"\u002Fnuclear-physics\u002Ffission\u002Fchain-reactions-reactor-physics":872,"\u002Fnuclear-physics\u002Ffusion-nucleosynthesis\u002Ffusion-reactions-confinement":160,"\u002Fnuclear-physics\u002Ffusion-nucleosynthesis\u002Fstellar-nucleosynthesis":534,"\u002Fnuclear-physics\u002Ffusion-nucleosynthesis\u002Fbig-bang-nucleosynthesis":396,"\u002Fnuclear-physics\u002Fradiation-matter-applications\u002Fcharged-particle-stopping-power":873,"\u002Fnuclear-physics\u002Fradiation-matter-applications\u002Fphoton-neutron-interactions":444,"\u002Fnuclear-physics\u002Fradiation-matter-applications\u002Fradiation-detectors":508,"\u002Fnuclear-physics\u002Fradiation-matter-applications\u002Fdosimetry-radiation-biology":874,"\u002Fnuclear-physics\u002Fradiation-matter-applications\u002Fnuclear-applications-dating-medicine":875,"\u002Fnuclear-physics":876,"\u002Fnatural-language-processing\u002Ffoundations\u002Fwhat-is-nlp":877,"\u002Fnatural-language-processing\u002Ffoundations\u002Fregex-and-text-normalization":878,"\u002Fnatural-language-processing\u002Ffoundations\u002Fminimum-edit-distance":515,"\u002Fnatural-language-processing\u002Ffoundations\u002Fn-gram-language-models":879,"\u002Fnatural-language-processing\u002Ffoundations\u002Fsmoothing-and-backoff":880,"\u002Fnatural-language-processing\u002Fclassification\u002Fnaive-bayes-and-sentiment":881,"\u002Fnatural-language-processing\u002Fclassification\u002Fevaluating-classifiers":363,"\u002Fnatural-language-processing\u002Fclassification\u002Flogistic-regression":882,"\u002Fnatural-language-processing\u002Fclassification\u002Fsentiment-and-affect-lexicons":883,"\u002Fnatural-language-processing\u002Fsemantics\u002Fvector-semantics-and-embeddings":660,"\u002Fnatural-language-processing\u002Fsemantics\u002Fstatic-word-embeddings":884,"\u002Fnatural-language-processing\u002Fsemantics\u002Fneural-language-models":829,"\u002Fnatural-language-processing\u002Fsequences\u002Fsequence-labeling":885,"\u002Fnatural-language-processing\u002Fsequences\u002Fcrfs-and-neural-taggers":886,"\u002Fnatural-language-processing\u002Fsequences\u002Frnns-and-lstms":887,"\u002Fnatural-language-processing\u002Ftransformers\u002Ftransformers-and-attention":888,"\u002Fnatural-language-processing\u002Ftransformers\u002Fthe-transformer-architecture":889,"\u002Fnatural-language-processing\u002Ftransformers\u002Flarge-language-models":890,"\u002Fnatural-language-processing\u002Ftransformers\u002Fllm-pretraining-and-scaling":891,"\u002Fnatural-language-processing\u002Ftransformers\u002Ffine-tuning-and-prompting":341,"\u002Fnatural-language-processing\u002Ftransformers\u002Fprompting-and-alignment":892,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fconstituency-parsing":893,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fcky-scoring-and-evaluation":835,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fdependency-parsing":894,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fgraph-based-and-neural-dependency-parsing":895,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fword-senses-and-wsd":896,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fwsd-in-practice-and-induction":897,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fsemantic-roles-and-information-extraction":898,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Frelations-events-and-templates":899,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fcoreference-and-discourse":900,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fcoherence-and-discourse-structure":901,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Flogical-semantics":759,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fcompositional-semantics-and-description-logics":902,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fsemantic-parsing":903,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fneural-semantic-parsing":904,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Finformation-extraction":905,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Ftimes-events-and-templates":906,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fdiscourse-coherence":907,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fentity-based-and-global-coherence":908,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Fconstituency-grammars":909,"\u002Fnatural-language-processing\u002Flinguistic-structure\u002Ftreebanks-and-lexicalized-grammars":910,"\u002Fnatural-language-processing\u002Fapplications\u002Fmachine-translation":911,"\u002Fnatural-language-processing\u002Fapplications\u002Fmachine-translation-decoding-and-evaluation":912,"\u002Fnatural-language-processing\u002Fapplications\u002Fquestion-answering":913,"\u002Fnatural-language-processing\u002Fapplications\u002Fquestion-answering-knowledge-and-llms":613,"\u002Fnatural-language-processing\u002Fapplications\u002Fdialogue-and-chatbots":780,"\u002Fnatural-language-processing\u002Fapplications\u002Fdialogue-systems-and-assistants":369,"\u002Fnatural-language-processing\u002Fapplications\u002Ftext-summarization":914,"\u002Fnatural-language-processing\u002Fapplications\u002Fabstractive-summarization-and-evaluation":915,"\u002Fnatural-language-processing\u002Fspeech\u002Fphonetics":916,"\u002Fnatural-language-processing\u002Fspeech\u002Facoustic-phonetics":917,"\u002Fnatural-language-processing\u002Fspeech\u002Fautomatic-speech-recognition":627,"\u002Fnatural-language-processing\u002Fspeech\u002Fasr-evaluation-and-applications":918,"\u002Fnatural-language-processing":56,"\u002Fparticle-physics\u002Ffoundations\u002Fhistorical-overview-particle-zoo":919,"\u002Fparticle-physics\u002Ffoundations\u002Fparticle-physics-basic-concepts":535,"\u002Fparticle-physics\u002Ffoundations\u002Ffundamental-interactions-force-carriers":920,"\u002Fparticle-physics\u002Funits-kinematics\u002Fnatural-units-and-scales":680,"\u002Fparticle-physics\u002Funits-kinematics\u002Ffour-vectors-invariant-mass":921,"\u002Fparticle-physics\u002Funits-kinematics\u002Fdecay-scattering-kinematics-mandelstam":922,"\u002Fparticle-physics\u002Funits-kinematics\u002Fcross-sections-golden-rule":870,"\u002Fparticle-physics\u002Fsymmetries\u002Fconservation-laws-symmetries":923,"\u002Fparticle-physics\u002Fsymmetries\u002Fdiscrete-symmetries-cpt":924,"\u002Fparticle-physics\u002Fsymmetries\u002Fparity-violation-weak":485,"\u002Fparticle-physics\u002Fsymmetries\u002Fsu2-su3-flavor-symmetry":188,"\u002Fparticle-physics\u002Fquark-model\u002Feightfold-way-su3":925,"\u002Fparticle-physics\u002Fquark-model\u002Fmeson-spectroscopy":439,"\u002Fparticle-physics\u002Fquark-model\u002Fbaryon-spectroscopy":926,"\u002Fparticle-physics\u002Fquark-model\u002Fcolor-confinement-exotics":472,"\u002Fparticle-physics\u002Frelativistic-wave-equations\u002Fklein-gordon-equation":927,"\u002Fparticle-physics\u002Frelativistic-wave-equations\u002Fdirac-equation-spinors":726,"\u002Fparticle-physics\u002Frelativistic-wave-equations\u002Fantiparticles-hole-theory":928,"\u002Fparticle-physics\u002Fqed\u002Ffeynman-rules-qed":929,"\u002Fparticle-physics\u002Fqed\u002Fqed-tree-processes":695,"\u002Fparticle-physics\u002Fqed\u002Frenormalization-running-coupling":930,"\u002Fparticle-physics\u002Fqed\u002Felectron-g-2":159,"\u002Fparticle-physics\u002Fweak-interaction\u002Fva-structure-weak":931,"\u002Fparticle-physics\u002Fweak-interaction\u002Fw-z-bosons-decays":932,"\u002Fparticle-physics\u002Fweak-interaction\u002Fckm-matrix":933,"\u002Fparticle-physics\u002Fweak-interaction\u002Fcp-violation-kaons-b-mesons":476,"\u002Fparticle-physics\u002Fqcd\u002Fcolor-su3-gluons":692,"\u002Fparticle-physics\u002Fqcd\u002Fasymptotic-freedom-confinement":934,"\u002Fparticle-physics\u002Fqcd\u002Fdeep-inelastic-scattering-partons":935,"\u002Fparticle-physics\u002Fqcd\u002Fjets-hadronization":551,"\u002Fparticle-physics\u002Felectroweak-higgs\u002Felectroweak-su2-u1":936,"\u002Fparticle-physics\u002Felectroweak-higgs\u002Fspontaneous-symmetry-breaking":596,"\u002Fparticle-physics\u002Felectroweak-higgs\u002Fhiggs-mechanism":937,"\u002Fparticle-physics\u002Felectroweak-higgs\u002Fhiggs-boson-discovery":938,"\u002Fparticle-physics\u002Felectroweak-higgs\u002Fstandard-model":485,"\u002Fparticle-physics\u002Fneutrinos\u002Fneutrino-oscillations":518,"\u002Fparticle-physics\u002Fneutrinos\u002Fneutrino-mass-pmns":939,"\u002Fparticle-physics\u002Fneutrinos\u002Fdirac-majorana-experiments":940,"\u002Fparticle-physics\u002Fexperiment\u002Faccelerators-luminosity":941,"\u002Fparticle-physics\u002Fexperiment\u002Fdetectors-subsystems":574,"\u002Fparticle-physics\u002Fexperiment\u002Fhow-discoveries-are-made":942,"\u002Fparticle-physics\u002Fbeyond-standard-model\u002Fbeyond-standard-model":316,"\u002Fparticle-physics\u002Fbeyond-standard-model\u002Fgrand-unified-theories":943,"\u002Fparticle-physics\u002Fbeyond-standard-model\u002Fsupersymmetry":418,"\u002Fparticle-physics\u002Fbeyond-standard-model\u002Fhierarchy-problem-naturalness":944,"\u002Fparticle-physics\u002Fbeyond-standard-model\u002Fdark-matter-candidates":945,"\u002Fparticle-physics\u002Fbeyond-standard-model\u002Fmatter-antimatter-open-questions":742,"\u002Fparticle-physics":946,"\u002Fastrophysics-cosmology\u002Forientation\u002Fthe-sun-and-stars":536,"\u002Fastrophysics-cosmology\u002Forientation\u002Fstellar-death-final-states":690,"\u002Fastrophysics-cosmology\u002Forientation\u002Fgalaxies-and-cosmology":947,"\u002Fastrophysics-cosmology\u002Fobservational-foundations\u002Fmagnitudes-fluxes-and-the-distance-modulus":475,"\u002Fastrophysics-cosmology\u002Fobservational-foundations\u002Fstellar-spectra-and-spectral-classification":948,"\u002Fastrophysics-cosmology\u002Fobservational-foundations\u002Ftelescopes-and-detectors-across-the-spectrum":535,"\u002Fastrophysics-cosmology\u002Fobservational-foundations\u002Fthe-cosmic-distance-ladder":447,"\u002Fastrophysics-cosmology\u002Fradiation-and-matter\u002Fblackbody-radiation-and-specific-intensity":949,"\u002Fastrophysics-cosmology\u002Fradiation-and-matter\u002Fradiative-transfer-and-the-transfer-equation":950,"\u002Fastrophysics-cosmology\u002Fradiation-and-matter\u002Fspectral-line-formation-and-broadening":951,"\u002Fastrophysics-cosmology\u002Fradiation-and-matter\u002Fopacity-and-the-rosseland-mean":952,"\u002Fastrophysics-cosmology\u002Fstellar-structure\u002Fhydrostatic-equilibrium-and-the-virial-theorem":953,"\u002Fastrophysics-cosmology\u002Fstellar-structure\u002Fthe-equations-of-stellar-structure":741,"\u002Fastrophysics-cosmology\u002Fstellar-structure\u002Fthe-equation-of-state-and-polytropes":672,"\u002Fastrophysics-cosmology\u002Fstellar-structure\u002Fthe-standard-solar-model":954,"\u002Fastrophysics-cosmology\u002Fnuclear-astrophysics\u002Fthermonuclear-reaction-rates-and-the-gamow-peak":955,"\u002Fastrophysics-cosmology\u002Fnuclear-astrophysics\u002Fhydrogen-burning-pp-chains-and-cno":956,"\u002Fastrophysics-cosmology\u002Fnuclear-astrophysics\u002Fhelium-burning-and-the-triple-alpha-process":957,"\u002Fastrophysics-cosmology\u002Fnuclear-astrophysics\u002Fadvanced-burning-and-neutron-capture-nucleosynthesis":869,"\u002Fastrophysics-cosmology\u002Fism-and-star-formation\u002Fphases-of-the-interstellar-medium":958,"\u002Fastrophysics-cosmology\u002Fism-and-star-formation\u002Fmolecular-clouds-and-gravitational-collapse":431,"\u002Fastrophysics-cosmology\u002Fism-and-star-formation\u002Fprotostars-and-the-pre-main-sequence":494,"\u002Fastrophysics-cosmology\u002Fstellar-evolution\u002Fthe-main-sequence-and-its-structure":476,"\u002Fastrophysics-cosmology\u002Fstellar-evolution\u002Fpost-main-sequence-low-mass-evolution":959,"\u002Fastrophysics-cosmology\u002Fstellar-evolution\u002Fthe-evolution-of-massive-stars":560,"\u002Fastrophysics-cosmology\u002Fstellar-evolution\u002Fstellar-pulsation-and-the-instability-strip":177,"\u002Fastrophysics-cosmology\u002Fstellar-death-and-compact-remnants\u002Fwhite-dwarfs-and-the-chandrasekhar-limit":960,"\u002Fastrophysics-cosmology\u002Fstellar-death-and-compact-remnants\u002Fcore-collapse-supernovae":961,"\u002Fastrophysics-cosmology\u002Fstellar-death-and-compact-remnants\u002Fthermonuclear-supernovae-type-ia":691,"\u002Fastrophysics-cosmology\u002Fstellar-death-and-compact-remnants\u002Fneutron-stars-and-pulsars":962,"\u002Fastrophysics-cosmology\u002Fstellar-death-and-compact-remnants\u002Fblack-holes-schwarzschild-and-kerr":963,"\u002Fastrophysics-cosmology\u002Fbinaries-and-gravitational-waves\u002Fbinary-systems-and-mass-transfer":964,"\u002Fastrophysics-cosmology\u002Fbinaries-and-gravitational-waves\u002Faccreting-compact-objects":965,"\u002Fastrophysics-cosmology\u002Fbinaries-and-gravitational-waves\u002Fgravitational-waves-from-inspiraling-binaries":158,"\u002Fastrophysics-cosmology\u002Fbinaries-and-gravitational-waves\u002Fmultimessenger-astronomy-and-gamma-ray-bursts":966,"\u002Fastrophysics-cosmology\u002Fgalaxies\u002Fthe-milky-way":967,"\u002Fastrophysics-cosmology\u002Fgalaxies\u002Fgalaxy-morphology-and-classification":418,"\u002Fastrophysics-cosmology\u002Fgalaxies\u002Fgalaxy-rotation-curves-and-dark-matter":968,"\u002Fastrophysics-cosmology\u002Fgalaxies\u002Factive-galactic-nuclei-and-supermassive-black-holes":969,"\u002Fastrophysics-cosmology\u002Fgalaxies\u002Fgalaxy-clusters-and-large-scale-structure":684,"\u002Fastrophysics-cosmology\u002Fcosmology-expansion-and-dynamics\u002Fthe-expanding-universe-and-hubbles-law":339,"\u002Fastrophysics-cosmology\u002Fcosmology-expansion-and-dynamics\u002Fthe-frw-metric-and-cosmological-redshift":970,"\u002Fastrophysics-cosmology\u002Fcosmology-expansion-and-dynamics\u002Fthe-friedmann-equations-and-cosmic-dynamics":971,"\u002Fastrophysics-cosmology\u002Fcosmology-expansion-and-dynamics\u002Fcosmological-models-and-distances":596,"\u002Fastrophysics-cosmology\u002Fcosmology-expansion-and-dynamics\u002Fdark-energy-and-the-accelerating-universe":582,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Fthe-thermal-history-of-the-universe":594,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Fbig-bang-nucleosynthesis":691,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Frecombination-and-the-cosmic-microwave-background":972,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Fcmb-anisotropies-and-cosmological-parameters":157,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Fcosmic-inflation":348,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Fstructure-formation-and-the-growth-of-perturbations":973,"\u002Fastrophysics-cosmology\u002Fthe-hot-big-bang\u002Fdark-matter-dark-energy-and-open-questions":435,"\u002Fastrophysics-cosmology":540,"\u002Fcolophon":974,"\u002F":56},4250,4808,3626,2682,4109,4786,3878,3875,3751,3415,4067,3153,3000,4042,5461,5808,3961,3749,4327,5067,4246,4655,4154,5436,2640,4003,3601,2158,4331,4189,2273,3252,4633,4964,4172,3131,5524,3160,4031,2309,4207,3226,2648,4842,5340,3307,5701,4977,4039,2615,3472,4460,3848,4075,4400,3382,3010,3602,3737,3740,3707,3922,5191,4043,3804,4542,4214,5062,2850,4361,3443,3627,4044,3766,4140,3860,4006,5199,4334,5234,3651,5509,5680,153,1375,1073,1093,1125,1146,1014,1132,876,1541,1189,1173,984,1402,1301,950,1268,1063,1107,1408,1161,925,1012,866,964,1090,1142,1085,1020,1207,973,980,728,764,1225,1329,796,929,801,878,774,1044,1488,1175,1130,890,814,870,154,4073,5140,4961,5127,4870,5382,5195,4955,5369,4501,5576,3824,4132,4289,4307,4570,3403,5084,5105,5201,5116,5341,5175,5368,5188,5211,5499,5155,4981,5125,5415,5255,5304,5130,5167,5552,5164,5094,5239,5036,5190,5004,5099,5035,5159,5088,5026,4937,5023,5264,5244,133,5114,5078,5043,5312,5170,5342,5139,5151,5049,5212,5013,5068,5079,5102,5121,5081,5029,5379,5854,5110,2139,3798,5055,5364,4984,4935,4895,4972,5289,5112,5156,4987,5031,5025,5149,5302,5042,5002,4979,4922,4960,5279,126,1877,1180,1129,907,958,1112,1300,1053,1250,1181,1241,1234,966,1050,734,1190,484,1082,926,733,761,571,607,798,804,952,977,731,784,645,771,1017,742,1004,1000,1562,1254,1288,1101,1011,1486,1061,856,992,1169,988,137,2037,1782,2384,2254,2123,2332,1643,1714,2089,1751,1367,1660,2511,1998,1892,1854,1791,2438,2487,1917,2375,2525,2266,1845,2275,1810,1631,2310,2166,2233,2113,2505,2347,2672,2112,2473,2592,2380,3013,2513,3256,3218,2194,2173,2205,2326,2081,3342,3152,1799,1670,1027,960,1095,1291,986,897,1209,1055,1817,1801,1593,1465,1196,1464,1201,1230,1435,1684,1461,1926,1500,1409,1284,1774,1869,162,1487,1122,1188,1001,1351,982,1005,979,1325,1046,943,1279,824,1008,989,1798,1277,1025,987,1043,1211,1074,981,939,1002,739,1139,1108,1013,1070,978,1458,1317,157,1357,1077,2355,1116,1037,1178,1637,1314,1109,1056,1702,1474,1071,1158,832,993,1404,1024,1068,1339,1106,1264,1248,913,1848,1328,1633,1224,1143,135,1378,959,1028,998,911,1527,1203,1266,1483,1165,990,938,965,1257,1418,1099,942,1352,956,1035,1398,1003,1094,1292,138,1721,1827,1449,1354,1148,1184,1285,1281,1213,1290,1271,1252,1274,1778,1591,1503,1437,1571,1584,1957,1117,1781,1648,1342,1667,1510,1965,1607,1365,1849,1259,1303,1356,1238,2208,1564,173,1671,1286,1227,1638,1529,668,1078,918,709,865,880,940,1534,1015,874,922,841,794,1194,822,1105,1658,1359,1296,1438,1921,1844,1570,1429,1324,1400,140,1787,1558,1654,1492,1747,2224,2002,2009,1323,1349,1785,1573,1722,1829,1353,1548,1552,1583,1624,1585,1245,1364,1514,1343,1397,1355,2211,1481,1770,160,2388,2293,2256,2552,2569,2478,2039,2496,2578,2814,2519,2461,2587,2492,2714,3278,2654,3050,2447,2849,2238,2369,2061,2214,2602,2563,2186,2985,2749,3364,2038,2282,2409,2126,2573,2206,2176,2268,2182,2402,2705,2633,2414,2213,2801,3313,3410,3195,1952,2017,1509,2537,2645,2027,2415,2838,2356,1906,3184,2950,2807,2954,1683,1316,1034,1138,1763,1822,1705,1246,1701,1097,1104,1187,1032,1083,1228,916,1489,1033,1652,997,692,837,1023,888,864,1089,1231,1214,1675,1156,1075,1520,1309,139,1205,1051,735,1123,1072,915,567,768,825,1253,983,1007,762,1058,861,862,971,1208,1149,1145,1029,1084,927,810,838,857,807,936,949,2321,1622,1069,1113,1057,854,1958,1528,1618,2049,1432,1679,1796,1685,1346,1275,1476,1505,1610,2018,1599,1215,1838,1909,132,3902,2215,2240,3266,3208,3073,2454,2969,2451,1875,2728,1884,2371,2516,2842,1690,1904,2346,3146,1386,2607,1966,2668,1665,2885,1606,2577,3074,2869,2403,2433,2082,1939,1587,2460,2747,2032,2642,1619,3123,1993,2090,2339,3829,1737,2622,2340,2322,3828,4409,2305,3411,2510,4527,3030,3569,3043,2457,1946,2277,2044,2909,1693,1945,2093,2399,2115,2898,2742,2242,3895,3378,3376,2769,2223,3062,3262,2651,2949,2768,3128,2423,1977,2087,2866,3388,2830,2210,2489,2884,3945,2099,2713,3402,1692,2931,4195,3989,3206,4391,3004,3704,3494,2902,999,881,901,919,748,869,1018,1045,1049,1333,954,1092,1019,976,1771,1480,1396,953,1026,161,3533,2495,1818,3007,2595,3427,3537,2216,1895,2304,3396,1739,2073,1962,2203,1767,2666,2264,2276,2852,1807,3735,1560,4144,1669,1676,1972,2418,3291,1525,2040,2766,2337,2220,2800,3001,2078,1759,2836,1896,2026,1758,1543,1047,896,946,1060,1384,1482,815,1414,1322,1440,1240,1468,1098,1133,847,1009,1381,1052,1191,1258,1370,1712,1441,1199,957,1079,150,1262,1417,1368,1219,1136,1064,1463,1636,1059,931,1115,1736,1174,1376,1363,1411,1247,1746,1313,1299,1617,1102,1076,1495,1265,1193,1263,80,[976,1015,1086,1163,1208,1330],{"module":977,"moduleNumber":978,"slug":979,"lessons":980},"Foundations",1,"foundations",[981,986,991,997,1003,1009],{"title":982,"path":983,"lessonNumber":978,"topics":984,"summary":985},"What Is Reinforcement Learning?","\u002Freinforcement-learning\u002Ffoundations\u002Fwhat-is-reinforcement-learning",[977],"Reinforcement learning is learning what to do — how to map situations to actions — so as to maximize a numerical reward signal, discovered by trial and error rather than told. We set up the agent–environment loop, separate it from supervised and unsupervised learning, name the four elements (policy, reward, value, and an optional model), and train a tic-tac-toe player with a temporal-difference value update.\n",{"title":987,"path":988,"lessonNumber":12,"topics":989,"summary":990},"A Brief History of Reinforcement Learning","\u002Freinforcement-learning\u002Ffoundations\u002Fa-brief-history-of-rl",[977],"The origins of reinforcement learning. Three threads — trial-and-error learning from animal psychology, optimal control and dynamic programming, and temporal-difference learning — ran independently for decades and merged around 1989 into the modern field. Replacing the lookup table with a neural network then produced deep reinforcement learning: DQN, AlphaGo, AlphaZero, MuZero, and RLHF.\n",{"title":992,"path":993,"lessonNumber":994,"topics":995,"summary":996},"Multi-Armed Bandits","\u002Freinforcement-learning\u002Ffoundations\u002Fmulti-armed-bandits",3,[977],"A bandit is reinforcement learning stripped to a single decision, repeated: no state, no consequences, only the tension between exploiting the arm that looks best and exploring the ones that might be better. We build up the whole toolkit — sample-average value estimates, the incremental update rule, ε-greedy, optimistic initialization, UCB, and gradient bandits — and use it to study exploration in isolation, the one problem that carries over to the full setting.\n",{"title":998,"path":999,"lessonNumber":1000,"topics":1001,"summary":1002},"Bandit Exploration Algorithms","\u002Freinforcement-learning\u002Ffoundations\u002Fbandit-exploration-algorithms",4,[977],"Better ways to explore than picking at random. Upper-confidence-bound selection explores by optimism about what it hasn't measured; gradient bandits learn action preferences by stochastic gradient ascent on reward. We then add context to get the contextual bandit, the bridge to full RL, and measure everything by regret — where UCB1 and Thompson sampling reach the logarithmic optimum that fixed-ε greedy cannot.\n",{"title":1004,"path":1005,"lessonNumber":1006,"topics":1007,"summary":1008},"Markov Decision Processes","\u002Freinforcement-learning\u002Ffoundations\u002Fmarkov-decision-processes",5,[977],"A Markov decision process is the formal interface between an agent and its environment: at each step the agent reads a state, chooses an action, and receives a reward and a next state. We fix that loop, the dynamics function that governs it, and the Markov property that makes the state sufficient; then turn goals into a scalar reward and rewards into a discounted return, with one notation that covers both episodic and continuing tasks.\n",{"title":1010,"path":1011,"lessonNumber":1012,"topics":1013,"summary":1014},"Value Functions and Optimality","\u002Freinforcement-learning\u002Ffoundations\u002Fvalue-functions-and-optimality",6,[977],"A value function scores how good a state (or state–action pair) is under a policy: the expected return from there onward. Its defining property is the Bellman equation, a self-consistency condition linking a state's value to its successors' values, which we derive from the return and the dynamics. Pushing the same idea to the best-achievable value gives the Bellman optimality equations, whose solution yields an optimal policy — and whose intractability is what the rest of the course is about.\n",{"module":1016,"moduleNumber":12,"slug":1017,"lessons":1018},"Tabular Solution Methods","tabular-methods",[1019,1025,1030,1035,1040,1045,1050,1056,1062,1068,1074,1080],{"title":1020,"path":1021,"lessonNumber":978,"topics":1022,"summary":1024},"Dynamic Programming","\u002Freinforcement-learning\u002Ftabular-methods\u002Fdynamic-programming",[1023],"Tabular Methods","Dynamic programming computes optimal policies when a perfect model of the MDP is given, by turning the Bellman equations into assignment statements. We build up iterative policy evaluation (the expected update), the policy improvement theorem, and the two classic algorithms that alternate them — policy iteration and value iteration — worked on the gridworld, a two-state MDP, Jack's car rental, and the gambler's problem.\n",{"title":1026,"path":1027,"lessonNumber":12,"topics":1028,"summary":1029},"Dynamic Programming: Asynchronous DP and Generalized Policy Iteration","\u002Freinforcement-learning\u002Ftabular-methods\u002Fdp-async-and-gpi",[1023],"Policy and value iteration both sweep the entire state set on every pass, which is impossible once the state space is huge. This lesson loosens the schedule: asynchronous DP updates states in any order, generalized policy iteration names the alternation of evaluation and improvement that underlies nearly every RL method, and a look at efficiency and the curse of dimensionality places DP among the alternatives. We close past Sutton & Barto with prioritized sweeping, neuro-dynamic programming, value-iteration networks, and MuZero.\n",{"title":1031,"path":1032,"lessonNumber":994,"topics":1033,"summary":1034},"Monte Carlo Methods","\u002Freinforcement-learning\u002Ftabular-methods\u002Fmonte-carlo-methods",[1023],"Monte Carlo methods learn value functions and optimal policies from complete sampled episodes, with no model of the environment: they simply average the returns that actually followed each state. We build prediction (first-visit and every-visit averaging), see why estimating action values forces the exploration question, and answer it two ways on-policy — exploring starts and epsilon-soft control. Throughout, Monte Carlo samples one whole trajectory to termination and never bootstraps.\n",{"title":1036,"path":1037,"lessonNumber":1000,"topics":1038,"summary":1039},"Monte Carlo Methods: Off-Policy Learning","\u002Freinforcement-learning\u002Ftabular-methods\u002Fmonte-carlo-off-policy",[1023],"On-policy Monte Carlo can only reach the best exploring policy, not the true optimum. Off-policy methods remove that ceiling by learning about a greedy target policy from data generated by a soft behavior policy, corrected with importance sampling. We derive the importance-sampling ratio, weigh ordinary against weighted estimators on real numbers, give the incremental off-policy algorithm, sharpen it with discounting-aware sampling, and close by placing Monte Carlo on the model\u002Fbootstrap map beside DP and temporal-difference learning.\n",{"title":1041,"path":1042,"lessonNumber":1006,"topics":1043,"summary":1044},"Temporal-Difference Learning","\u002Freinforcement-learning\u002Ftabular-methods\u002Ftemporal-difference-learning",[1023],"Temporal-difference learning is the one idea most central to reinforcement learning: learn a value directly from experience, like Monte Carlo, but update each guess toward the next guess before the episode ends, like dynamic programming. We derive the TD(0) prediction rule and its reward-prediction error, contrast its one-step backup with MC and DP, work the driving-home and random-walk examples, and show the batch-updating optimality that makes TD approximate the certainty-equivalence estimate.\n",{"title":1046,"path":1047,"lessonNumber":1012,"topics":1048,"summary":1049},"TD Control: Sarsa, Q-learning, and Double Learning","\u002Freinforcement-learning\u002Ftabular-methods\u002Ftd-control-sarsa-and-q-learning",[1023],"With TD prediction in hand, control follows the generalized-policy-iteration pattern with TD as the evaluation step. We build Sarsa (on-policy), Q-learning (off-policy, targeting the optimal policy), and Expected Sarsa that spans the two, then confront the maximization bias every max-based method inherits and fix it with Double Q-learning. We close past Sutton & Barto, following each one-step tabular update into its deep-RL descendant — DQN, Double DQN, and Rainbow.\n",{"title":1051,"path":1052,"lessonNumber":1053,"topics":1054,"summary":1055},"n-Step Bootstrapping","\u002Freinforcement-learning\u002Ftabular-methods\u002Fn-step-bootstrapping",7,[1023],"Monte Carlo waits for the full return; one-step TD bootstraps after a single reward. Between them lies a whole spectrum, indexed by one integer n: look ahead n real rewards, then bootstrap from the value n steps out. The n-step return unifies the previous two lessons, and — on the random walk — an intermediate n beats both extremes. We build the n-step return, the n-step TD update, the backup-diagram spectrum, and n-step Sarsa for control.\n",{"title":1057,"path":1058,"lessonNumber":1059,"topics":1060,"summary":1061},"n-Step Bootstrapping: Off-Policy Methods","\u002Freinforcement-learning\u002Ftabular-methods\u002Fn-step-off-policy-methods",8,[1023],"Taking the n-step family off-policy raises the same importance-sampling questions Monte Carlo did, now over a window of exactly n actions. We reweight n-step returns by the policy ratio, watch the ratio product inflate variance on real numbers, then build the tree-backup algorithm that learns off-policy with no ratios at all — and finally n-step Q(sigma), one algorithm whose per-step switch recovers Sarsa, tree backup, and Expected Sarsa as special cases.\n",{"title":1063,"path":1064,"lessonNumber":1065,"topics":1066,"summary":1067},"Planning and Learning","\u002Freinforcement-learning\u002Ftabular-methods\u002Fplanning-and-learning",9,[1023],"Planning and learning are the same operation run on two kinds of experience. A model turns states and actions into simulated transitions; planning backs up values over that simulated experience exactly as learning backs them up over real experience. We build the Dyna architecture that interleaves acting, model-learning, direct RL, and planning in one loop, trace a single Dyna-Q step by hand, and patch the architecture for when the model goes stale.\n",{"title":1069,"path":1070,"lessonNumber":1071,"topics":1072,"summary":1073},"Planning: Focusing Updates and Decision-Time Search","\u002Freinforcement-learning\u002Ftabular-methods\u002Fplanning-focusing-and-decision-time",10,[1023],"Dyna plans by replaying remembered transitions, but sampling them uniformly wastes most of the effort. This lesson sharpens planning: prioritized sweeping works backward from states whose value just changed, expected versus sample updates weigh thoroughness against cost, and trajectory sampling and real-time DP focus updates on the states the policy actually visits. We trace Dyna forward to model-based deep RL, then turn to decision-time planning — heuristic search, rollouts, and Monte Carlo Tree Search.\n",{"title":1075,"path":1076,"lessonNumber":1077,"topics":1078,"summary":1079},"Decision-Time Planning","\u002Freinforcement-learning\u002Ftabular-methods\u002Fdecision-time-planning",11,[1023],"Planning need not build a global policy. Decision-time planning runs a fresh lookahead every time a state arrives and returns just one action, then throws the work away. We start from real-time dynamic programming — asynchronous value iteration on the states the agent actually visits — then move through heuristic search and rollout algorithms, each a one-step policy improvement applied on the fly to the current state.\n",{"title":1081,"path":1082,"lessonNumber":1083,"topics":1084,"summary":1085},"Monte Carlo Tree Search","\u002Freinforcement-learning\u002Ftabular-methods\u002Fmonte-carlo-tree-search",12,[1023],"Monte Carlo Tree Search is a rollout algorithm with memory: it accumulates value estimates across simulations and steers later ones toward promising branches. We work through the four steps — selection, expansion, simulation, backup — the UCT selection rule computed on real numbers, the asymmetric growing tree, and the full pseudocode. We close past Sutton & Barto with the lineage from UCT to AlphaGo, AlphaZero, and MuZero, where a learned network stands in for the leaf value and the rollout.\n",{"module":1087,"moduleNumber":994,"slug":1088,"lessons":1089},"Approximate Solution Methods","approximation",[1090,1096,1101,1106,1111,1116,1121,1126,1131,1136,1141,1146,1151,1157],{"title":1091,"path":1092,"lessonNumber":978,"topics":1093,"summary":1095},"On-Policy Prediction with Approximation","\u002Freinforcement-learning\u002Fapproximation\u002Fon-policy-prediction",[1094],"Approximation","Every tabular method so far stored one number per state, which fails once the state space is large or continuous. We replace the table with a parameterized value function $\\hat v(s,\\mathbf{w})$, define the mean squared value error it should minimize under the on-policy distribution, and derive stochastic- and semi-gradient learning rules — the semi-gradient TD(0) update that bootstraps and so is not a true gradient. Linear methods make the analysis clean and give the TD fixed point; feature construction (polynomials, Fourier basis, coarse and tile coding, RBFs) supplies the vectors $\\mathbf{x}(s)$, and neural networks are the nonlinear bridge to deep RL.\n",{"title":1097,"path":1098,"lessonNumber":12,"topics":1099,"summary":1100},"Feature Construction and Nonlinear Approximation","\u002Freinforcement-learning\u002Fapproximation\u002Ffeature-construction-and-nonlinear",[1094],"Linear methods are only as good as the feature vectors $\\mathbf{x}(s)$ fed to them, and this lesson builds those vectors. Polynomials and the Fourier basis turn a state's coordinates into smooth global features; coarse coding, tile coding, and radial basis functions cover a continuous space with overlapping local receptive fields whose size sets the reach of generalization. Then we stop designing features by hand: a neural network learns the representation itself by gradient descent, trading the convergence guarantees of the linear case for expressiveness — the bridge to deep reinforcement learning.\n",{"title":1102,"path":1103,"lessonNumber":994,"topics":1104,"summary":1105},"On-Policy Control with Approximation","\u002Freinforcement-learning\u002Fapproximation\u002Fon-policy-control",[1094],"Prediction learned a value function from features; control learns to act. We carry semi-gradient methods over to action values $\\hat q(s,a,\\mathbf{w})$, giving episodic semi-gradient Sarsa and its n-step form, and solve Mountain Car by descending a cost-to-go surface. In the continuing case, function approximation makes discounting unable to affect which policy is best, so we replace it with the average-reward setting — the differential return, differential value functions, and differential semi-gradient Sarsa.\n",{"title":1107,"path":1108,"lessonNumber":1000,"topics":1109,"summary":1110},"Average-Reward Control for Continuing Tasks","\u002Freinforcement-learning\u002Fapproximation\u002Faverage-reward-control",[1094],"With function approximation, discounting has no effect on a continuing task: averaged over the on-policy distribution, the discounted objective equals the average reward times a policy-independent constant, so $\\gamma$ cannot change which policy is best. This lesson replaces discounting with the average-reward setting — the long-run reward rate $r(\\pi)$, the differential return that measures each state's transient advantage over that rate, differential value functions and TD error, and differential semi-gradient Sarsa, the control method for continuing tasks that never invokes a discount factor.\n",{"title":1112,"path":1113,"lessonNumber":1006,"topics":1114,"summary":1115},"Off-Policy Methods and the Deadly Triad","\u002Freinforcement-learning\u002Fapproximation\u002Foff-policy-and-the-deadly-triad",[1094],"Off-policy learning with function approximation is where the convergence guarantees of reinforcement learning fail. We extend the tabular off-policy updates to semi-gradient form with per-step importance sampling, show Baird's counterexample driving the weights to infinity, and identify the cause: the deadly triad of function approximation, bootstrapping, and off-policy training — any two are safe, all three can diverge. The divergence is not caused by sampling noise: a fully synchronous dynamic-programming update blows up just the same, which is what makes the triad a structural hazard rather than a fluke.\n",{"title":1117,"path":1118,"lessonNumber":1012,"topics":1119,"summary":1120},"Value-Function Geometry and Gradient-TD Methods","\u002Freinforcement-learning\u002Fapproximation\u002Fbellman-error-and-gradient-td",[1094],"Why does the deadly triad diverge, and how do you stop it? This lesson develops the geometry that explains the failure: value functions as vectors, the projection operator onto the representable subspace, and the split between the Bellman error, the value error, and the projected Bellman error: the three objectives have different minimizers. The projected Bellman error is the learnable one, and Gradient-TD methods (GTD2, TDC) do true stochastic gradient descent on it, staying stable even off-policy at $O(d)$ cost. Emphatic TD reweights states instead, and a survey of variance-reduction techniques closes the gap between stability and usable learning.\n",{"title":1122,"path":1123,"lessonNumber":1053,"topics":1124,"summary":1125},"Eligibility Traces","\u002Freinforcement-learning\u002Fapproximation\u002Feligibility-traces",[1094],"n-step methods unify TD and Monte Carlo by storing the last n feature vectors; eligibility traces do the same job with a single short-term memory vector. The λ-return averages every n-step return under a geometric weighting; the forward view looks ahead to that average, and the backward view produces nearly the same updates online through a decaying trace vector. We build the λ-return, TD(λ) with its trace, the two ways λ recovers TD(0) and Monte Carlo, a note on the exact equivalence of true online TD(λ), and Sarsa(λ) for control.\n",{"title":1127,"path":1128,"lessonNumber":1059,"topics":1129,"summary":1130},"True Online TD(λ) and Sarsa(λ)","\u002Freinforcement-learning\u002Fapproximation\u002Ftrue-online-and-sarsa-lambda",[1094],"Plain TD(λ) makes the forward and backward views nearly agree; this lesson closes the gap. True online TD(λ) uses a dutch trace and a small correction term to produce exactly the same weight sequence as the online λ-return algorithm, at the same memory and only a constant factor more compute — the sharpest statement of the forward\u002Fbackward duality. The whole apparatus then lifts to control unchanged: Sarsa(λ) threads a single delayed reward back along an entire trajectory in one sweep, and the λ-weighting reappears in modern deep RL as generalized advantage estimation.\n",{"title":1132,"path":1133,"lessonNumber":1065,"topics":1134,"summary":1135},"Policy Gradient Methods","\u002Freinforcement-learning\u002Fapproximation\u002Fpolicy-gradient-methods",[1094],"Every method so far learned values and read a policy off them. Policy gradient methods drop the intermediary: parameterize the policy directly and climb the performance gradient. We build the softmax-in-preferences parameterization, prove the policy gradient theorem that makes the gradient computable without the unknown state distribution, and derive REINFORCE and its variance-cutting state-value baseline — the launch point for the bootstrapping actor-critic that follows.\n",{"title":1137,"path":1138,"lessonNumber":1071,"topics":1139,"summary":1140},"Actor-Critic Methods and Continuous Actions","\u002Freinforcement-learning\u002Fapproximation\u002Factor-critic-and-continuous-actions",[1094],"REINFORCE with a baseline learns a value function but never bootstraps; this lesson adds the bootstrapping critic that completes the actor-critic architecture. The critic scores each transition into a single TD error that steers both the actor's policy step and its own value step, trading a little bias for much lower variance and fully online, continuing-task learning. The policy gradient theorem carries over unchanged to the average-reward setting, a Gaussian policy handles real-valued actions with self-tuning exploration, and the natural policy gradient leads straight to TRPO, PPO, and the deep actor-critic methods that train today's agents.\n",{"title":1142,"path":1143,"lessonNumber":1077,"topics":1144,"summary":1145},"Least-Squares TD","\u002Freinforcement-learning\u002Fapproximation\u002Fleast-squares-and-memory-based-methods",[1094],"Semi-gradient TD spends one cheap step per example and needs many examples; this lesson makes the opposite tradeoff. Least-Squares TD (LSTD) accumulates the matrices $\\mathbf{A}$ and $\\mathbf{b}$ and solves the TD fixed point $\\mathbf{w} = \\mathbf{A}^{-1}\\mathbf{b}$ directly, using the Sherman-Morrison identity to maintain the inverse in $O(d^2)$ — the most data-efficient linear TD method, at a quadratic cost. We work a solve by hand, weigh the quadratic cost against semi-gradient TD's cheap steps, and note that LSTD never forgets — a problem in control, where least-squares policy iteration is the natural extension.\n",{"title":1147,"path":1148,"lessonNumber":1083,"topics":1149,"summary":1150},"Memory-Based and Kernel Methods","\u002Freinforcement-learning\u002Fapproximation\u002Fmemory-and-kernel-methods",[1094],"Least-squares TD spent more compute to extract more from each example; this lesson drops the parametric form entirely. Memory-based methods store training examples untouched and answer a query locally at retrieval time — nearest neighbor, weighted average, locally weighted regression — so accuracy grows with the data and effort concentrates where the agent actually goes. Kernel-based methods weight stored examples by a similarity kernel $k(s,s')$, and every linear method turns out to be a kernel method. Interest and emphasis, finally, make the on-policy weighting itself a design choice, aiming scarce approximation capacity at the states that matter.\n",{"title":1152,"path":1153,"lessonNumber":1154,"topics":1155,"summary":1156},"Off-Policy Eligibility Traces","\u002Freinforcement-learning\u002Fapproximation\u002Foff-policy-eligibility-traces",13,[1094],"Eligibility traces meet off-policy learning and function approximation — the corner where stability gets hard. We first let the bootstrapping and discounting parameters vary with state, so a single generalized return covers episodic and continuing tasks and folds termination into the discount. Then we fold the per-decision importance ratio into the trace with a control-variate correction, and build Watkins's Q(λ) and its importance-sampling-free successor Tree-Backup(λ) — all correct in expectation, but still semi-gradient, so the deadly triad and its fixes wait for the next lesson.\n",{"title":1158,"path":1159,"lessonNumber":1160,"topics":1161,"summary":1162},"Stable Off-Policy Methods with Traces","\u002Freinforcement-learning\u002Fapproximation\u002Fstable-off-policy-traces",14,[1094],"Off-policy traces get the expected target right, but with $\\lambda \u003C 1$ they bootstrap, so off-policy plus bootstrapping plus function approximation is the deadly triad and the weights can diverge. This lesson carries the two one-step fixes to traces: GTD(λ) and GQ(λ) add a second weight vector and a gradient correction for true gradient descent on the projected Bellman error, while Emphatic TD(λ) reweights updates through a followon trace and interest to recover the on-policy stability. It closes with the implementation reality that traces are cheap because they are sparse, and with Retrace and V-trace — the clipped-ratio descendants that make off-policy traces work at deep-RL scale.\n",{"module":1164,"moduleNumber":1000,"slug":1165,"lessons":1166},"Deep Reinforcement Learning","deep-rl",[1167,1173,1178,1183,1188,1193,1198,1203],{"title":1168,"path":1169,"lessonNumber":978,"topics":1170,"summary":1172},"Deep Q-Networks","\u002Freinforcement-learning\u002Fdeep-rl\u002Fdeep-q-networks",[1171],"Deep RL","Deep Q-networks replace the linear value function with a neural network $Q(s,a;\\mathbf{w})$ and confront the fact that a nonlinear approximator, off-policy bootstrapping, and correlated online data — the deadly triad — make naive Q-learning diverge. DQN counters this empirically with two stabilizers: an experience replay buffer that decorrelates and reuses samples, and a periodically-frozen target network that fixes the bootstrap target. We derive the DQN loss and gradient, walk through the Atari convolutional architecture and its results, and then add the three refinements that define modern value-based deep RL — Double DQN, dueling networks, and prioritized experience replay.\n",{"title":1174,"path":1175,"lessonNumber":12,"topics":1176,"summary":1177},"DQN Improvements: Double, Dueling, and Prioritized Replay","\u002Freinforcement-learning\u002Fdeep-rl\u002Fdqn-improvements",[1171],"Three refinements that turn plain DQN into the standard modern value-based agent, each touching a different part of the system. Double DQN fixes the maximization bias in the target by splitting action selection from evaluation; dueling networks restructure the network around a state value and per-action advantages; prioritized replay changes which transitions are learned from. We close with Rainbow, which combines them, and the distributional view that predicts the whole return distribution rather than its mean.\n",{"title":1179,"path":1180,"lessonNumber":994,"topics":1181,"summary":1182},"Actor–Critic and GAE","\u002Freinforcement-learning\u002Fdeep-rl\u002Factor-critic-and-ppo",[1171],"Make the actor and the critic deep networks and the policy-gradient architecture becomes modern deep RL. We build the neural actor-critic, the advantage estimate that replaces the raw return, and Generalized Advantage Estimation as a λ-blend of n-step advantages, then the parallel-worker methods A3C and A2C that decorrelate on-policy data. The step-size constraints — trust regions, PPO, and the continuous-control family — follow in the next lesson.\n",{"title":1184,"path":1185,"lessonNumber":1000,"topics":1186,"summary":1187},"PPO and Continuous Control","\u002Freinforcement-learning\u002Fdeep-rl\u002Fppo-and-continuous-control",[1171],"Keeping the policy-gradient step from destroying the policy, and the algorithms that result. Trust-region optimization bounds each update by a KL constraint; PPO keeps that goal but replaces the second-order machinery with a first-order clip on the probability ratio, which is why it is the modern default and the optimizer inside RLHF. We then tour the off-policy continuous-control family — DDPG, TD3, and SAC — and where actor-critic went at scale, from OpenAI Five to language-model alignment.\n",{"title":1189,"path":1190,"lessonNumber":1006,"topics":1191,"summary":1192},"Case Studies: Learning to Play","\u002Freinforcement-learning\u002Fdeep-rl\u002Fcase-studies",[1171],"The game-playing systems that turned reinforcement learning from a theory into a track record: Samuel's checkers player, TD-Gammon, Watson's Daily-Double wagering, a reinforcement-learning memory controller, DQN, and AlphaGo through AlphaGo Zero. Read as a set they draw one line — a value function, learned by self-play or interaction, refined by search, carried by a deep network — that runs from a 1959 checkers program to superhuman Go.\n",{"title":1194,"path":1195,"lessonNumber":1012,"topics":1196,"summary":1197},"Reinforcement Learning Beyond Games","\u002Freinforcement-learning\u002Fdeep-rl\u002Frl-beyond-games",[1171],"The same value-and-reward machinery, pointed at problems with no opponent. Web personalization as a contextual bandit and then a full MDP for life-time value; thermal soaring, where a glider learns to climb on turbulent air and reward design does most of the work; and the industrial-scale systems that carried the same design past Sutton & Barto — AlphaStar, OpenAI Five, GT Sophy, and RLHF, where the reward itself is learned from human preference.\n",{"title":1199,"path":1200,"lessonNumber":1053,"topics":1201,"summary":1202},"Frontiers: Beyond the Standard MDP","\u002Freinforcement-learning\u002Fdeep-rl\u002Ffrontiers",[1171],"The standard MDP fixes three things — state, reward, and single-step actions — and this lesson loosens two of them. We generalize the value function into a general value function that predicts any signal, and use those predictions as auxiliary tasks that shape representations; we extend actions in time with the options framework; and we treat state as a construction the agent builds from a stream of observations. Reward design and the open problems follow in the next lesson.\n",{"title":1204,"path":1205,"lessonNumber":1059,"topics":1206,"summary":1207},"Reward Design and Open Problems","\u002Freinforcement-learning\u002Fdeep-rl\u002Freward-design-and-open-problems",[1171],"How to design a reward signal that encodes the intended goal — sparse reward, shaping, and reward hacking — and the problems the whole tabular, approximate, and deep arc leaves unsolved. We close with how the frontiers were pushed after Sutton & Barto: auxiliary tasks, learned options, intrinsic-motivation bonuses, learned world models, and offline RL, then the two concerns of reward hacking and safety that any real-world agent must address.\n",{"module":1209,"moduleNumber":1006,"slug":1210,"lessons":1211},"Modern Deep Reinforcement Learning","modern-deep-rl",[1212,1217,1222,1227,1232,1237,1242,1247,1252,1257,1262,1267,1272,1277,1282,1288,1294,1300,1306,1312,1318,1324],{"title":1213,"path":1214,"lessonNumber":978,"topics":1215,"summary":1216},"Sharpening DQN: Improvements and the Distributional Idea","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fdistributional-and-rainbow",[1171],"In the years after the 2015 DQN paper, a stream of focused improvements each fixed one weakness of the baseline without disturbing its frame. This lesson recaps five that keep the scalar $Q$-value — Double DQN, multi-step returns, dueling networks, prioritized replay, and NoisyNets, each changing a different slot of the same Q-learning loop — then develops the sixth, distributional RL, which changes the objective itself: learn the whole return distribution $Z(s,a)$. We build the distributional Bellman equation and the C51 categorical algorithm, projection step and all, worked end to end on real numbers. A companion lesson takes up QR-DQN, Rainbow, and the modern distributional line.\n",{"title":1218,"path":1219,"lessonNumber":12,"topics":1220,"summary":1221},"Distributional RL and Rainbow","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fdistributional-and-rainbow-part-2",[1171],"A companion to the DQN improvements lesson. C51 fixed the return atoms and learned their probabilities; QR-DQN does the reverse — fix the probabilities, learn the values — which removes the projection and trains with a quantile loss. We cover why the distribution helps even when you act on the mean, then assemble Rainbow: all six improvements in one Q-learning loop, with the component ablation that shows each one's real weight. The distributional line then runs on through IQN, FQF, and Agent57, the first agent to beat the human baseline on all 57 Atari games.\n",{"title":1223,"path":1224,"lessonNumber":994,"topics":1225,"summary":1226},"Continuous Control: DDPG and TD3","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fcontinuous-control",[1171],"When actions are real-valued, the $\\arg\\max_a Q(s,a)$ in Q-learning becomes an optimization problem on every step. This lesson builds the off-policy actor-critic family that sidesteps it: the deterministic policy gradient and DDPG, which replaces the max with a learned actor, and the three fixes of TD3 that counter the value overestimation DDPG inherits. A companion lesson takes up SAC's maximum-entropy objective and the methods built on this template.\n",{"title":1228,"path":1229,"lessonNumber":1000,"topics":1230,"summary":1231},"Continuous Control: SAC and Beyond","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fcontinuous-control-part-2",[1171],"A companion to the DDPG and TD3 lesson. Where those actors are deterministic and explore with bolted-on noise, soft actor-critic (SAC) changes the objective itself: maximize return plus the entropy of the policy, so exploration becomes intrinsic and the agent stays robust. We develop the maximum-entropy objective, the reparameterized squashed-Gaussian actor, and automatic temperature tuning, then survey the methods built on this off-policy template — distributional critics (D4PG), critic ensembles (REDQ), and control from pixels (DrQ, RAD).\n",{"title":1233,"path":1234,"lessonNumber":1006,"topics":1235,"summary":1236},"Model-Based Deep RL: Sample Efficiency and PETS","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmodel-based-rl",[1171],"A model turns experience into imagined planning. This lesson makes the sample-efficiency case for learning a dynamics model, works through why a learned model's errors compound over the planning horizon, and builds the most direct model-based method: PETS plans online with a probabilistic ensemble under model-predictive control, distrusting the model exactly where its members disagree. A companion lesson takes up latent world models (Dreamer) and MuZero.\n",{"title":1238,"path":1239,"lessonNumber":1012,"topics":1240,"summary":1241},"Model-Based Deep RL: World Models, Dreamer, and MuZero","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmodel-based-rl-part-2",[1171],"A companion to the PETS lesson. PETS plans in the environment's native state space; these methods change what the model represents. World Models and Dreamer learn a compact latent state and do almost all their learning by imagining inside it, with value gradients flowing through the differentiable dynamics. MuZero predicts neither states nor pixels — only the reward, value, and policy that MCTS reads — and plans with search against that learned model, AlphaZero without the rules. We close with MBPO, TD-MPC, and EfficientZero.\n",{"title":1243,"path":1244,"lessonNumber":1053,"topics":1245,"summary":1246},"Exploration in Deep RL: Novelty as Reward","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fexploration",[1171],"When the state space is enormous and reward is rare, ε-greedy amounts to a random walk that almost never reaches the first reward. This lesson scales the bandit's exploration ideas up to deep RL through the dominant approach — manufacture a reward for novelty and let the agent chase it: optimism and pseudo-counts from density models, and intrinsic motivation and curiosity (the Intrinsic Curiosity Module and Random Network Distillation). A companion lesson takes up posterior sampling, Go-Explore, and the modern methods.\n",{"title":1248,"path":1249,"lessonNumber":1059,"topics":1250,"summary":1251},"Exploration in Deep RL: Posterior Sampling and Go-Explore","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fexploration-part-2",[1171],"A companion to the novelty-as-reward lesson. Pseudo-counts and curiosity reward the unfamiliar after the agent stumbles into it; this lesson covers two ideas that go further. Bootstrapped DQN keeps an ensemble that approximates a posterior over value functions and explores by committing to one sampled hypothesis per episode — the deep, directed exploration ε-greedy cannot manage. Go-Explore remembers and returns to the frontier, defeating detachment and derailment to solve Montezuma's Revenge. We close with episodic memory (Never Give Up), Agent57, and model-based exploration.\n",{"title":1253,"path":1254,"lessonNumber":1065,"topics":1255,"summary":1256},"Offline RL: The Problem and Value-Based Fixes","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Foffline-rl",[1171],"Offline reinforcement learning learns a policy from a fixed logged dataset with no further environment interaction — off-policy learning pushed to the extreme, and it breaks for the extreme version of the same reason. Bootstrapping queries the value function at out-of-distribution actions the data never covers, those errors are optimistic, and with no online feedback to correct them they compound through the Bellman backup. This lesson sets up the failure and off-policy evaluation, then builds the first two families of pessimistic fixes: policy constraint (BCQ) and conservative value estimation (CQL). A companion lesson takes up implicit methods, model-based offline RL, and Decision Transformer.\n",{"title":1258,"path":1259,"lessonNumber":1071,"topics":1260,"summary":1261},"Offline RL: Implicit Methods, Sequence Models, and Beyond","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Foffline-rl-part-2",[1171],"A companion to the offline-RL problem lesson. Policy constraint and conservative value estimation both still query a learned value function; implicit methods (IQL) avoid querying it off the data at all, using an in-sample expectile backup. We then build pessimism into a learned model (MOPO, COMBO) and drop bootstrapping entirely with Decision Transformer's return-conditioned sequence modeling, closing with offline-to-online fine-tuning, diffusion planners, and the offline view of RLHF. The one rule throughout: without online correction, be pessimistic about what you cannot verify.\n",{"title":1263,"path":1264,"lessonNumber":1077,"topics":1265,"summary":1266},"Imitation Learning: Cloning, DAgger, and Inverse RL","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fimitation-and-inverse-rl",[1171],"When a reward is hard to specify but an expert is easy to watch, learn from demonstrations instead. Behavioral cloning treats control as supervised learning of the expert's state-to-action map, and fails through compounding error: small mistakes carry the agent off the expert's distribution, where it was never trained. DAgger fixes the mismatch by querying the expert on the learner's own states. Inverse RL instead recovers the reward the expert seems to optimize — an ill-posed problem that maximum-entropy IRL disambiguates. A companion lesson casts imitation as adversarial occupancy matching (GAIL, AIRL).\n",{"title":1268,"path":1269,"lessonNumber":1083,"topics":1270,"summary":1271},"Imitation as Adversarial Matching: GAIL and AIRL","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fimitation-and-inverse-rl-part-2",[1171],"A companion to the imitation-learning lesson. If the point of recovering a reward is only to re-run RL and match the expert, you can skip the reward and match the behavior directly. GAIL casts imitation as a GAN — a discriminator separating expert from learner state-action pairs supplies the reward a policy-gradient method optimizes — matching occupancy measures without ever naming a reward. AIRL reads a transferable reward back out of the discriminator. We compare all four methods and close with reward models in RLHF, scaled cloning, and diffusion policies.\n",{"title":1273,"path":1274,"lessonNumber":1154,"topics":1275,"summary":1276},"Multi-Agent RL: Markov Games and Centralized Training","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmulti-agent-rl",[1171],"With more than one learning agent in an environment, each agent's world becomes non-stationary because the others are changing too. This lesson builds the Markov-game generalization of the MDP, diagnoses non-stationarity as the central obstacle, shows why the naive baselines fail, and develops the dominant fix — centralized training with decentralized execution (MADDPG, VDN, QMIX). A companion lesson takes up self-play, the landmark game-playing systems, and the equilibrium concepts that define what \"solved\" means.\n",{"title":1278,"path":1279,"lessonNumber":1160,"topics":1280,"summary":1281},"Multi-Agent RL: Self-Play and Solution Concepts","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmulti-agent-rl-part-2",[1171],"A companion to the Markov-games lesson. In the purely competitive setting, an agent can generate its own training curriculum by playing against copies of itself — self-play, the method behind AlphaGo, OpenAI Five, and AlphaStar. We develop why self-play produces an ever-improving opponent, the systems it built, and then the equilibrium solution concepts (Nash, correlated, coarse-correlated) that define what \"solved\" means once there is an opponent, closing with PSRO, MAPPO, and the language-model-agent frontier.\n",{"title":1283,"path":1284,"lessonNumber":1285,"topics":1286,"summary":1287},"Hierarchical RL: Options and the Option-Critic","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fhierarchical-rl",15,[1171],"Flat RL cannot explore a long horizon: reaching reward through hundreds of primitive actions is exponentially unlikely, and every credit-assignment update crawls one step at a time. Hierarchy breaks one hard long-horizon problem into many short ones. This lesson develops temporal abstraction — the options framework and its semi-Markov view, and learning options end to end with the option-critic. A companion lesson takes up goal-conditioned manager\u002Fworker hierarchies (FeUdal Networks and HIRO), hindsight relabeling, and unsupervised skill discovery.\n",{"title":1289,"path":1290,"lessonNumber":1291,"topics":1292,"summary":1293},"Hierarchical RL: Goal-Conditioned Hierarchies and Skills","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fhierarchical-rl-part-2",16,[1171],"A companion to the options lesson. Options package a behavior; goal-conditioned hierarchies instead give the top level an explicit language of goals — a manager proposes a target state or a latent direction, and a worker is rewarded for reaching it (FeUdal Networks, HIRO). We develop that architecture, the hindsight relabeling that lets it learn from sparse reward, and unsupervised skill discovery (DIAYN) that learns a repertoire of behaviors with no reward at all. The shared idea throughout: shorten the horizon by inserting a level that decides less often.\n",{"title":1295,"path":1296,"lessonNumber":1297,"topics":1298,"summary":1299},"RLHF and Language Models","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Frlhf-and-language-models",17,[1171],"A language model trained to predict the next token is fluent but not helpful, honest, or harmless — the objective it was optimized for is not the objective we want. RLHF closes that gap by turning the one thing humans do reliably, comparing two outputs, into a reward. We build the three-stage pipeline: supervised fine-tuning, a Bradley-Terry reward model fit to preference pairs, then PPO against that reward with a KL penalty keeping it near the reference policy. We then cover reward hacking and why the KL penalty matters, Direct Preference Optimization, which folds the reward model into a single classification loss, and the RLAIF and verifiable-reward variants. This pipeline is what makes the largest models usable as assistants.\n",{"title":1301,"path":1302,"lessonNumber":1303,"topics":1304,"summary":1305},"Partial Observability: POMDPs and the Belief State","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fpartial-observability-pomdps",18,[1171],"Drop the assumption that the agent sees the state. It sees an observation, a partial and noisy function of a hidden state, and one observation is no longer a Markov signal. This lesson builds the POMDP tuple, shows that the belief state — the posterior over hidden states — is a sufficient statistic that turns a POMDP back into an MDP over beliefs, and works the Bayes-filter belief update step by step. A companion lesson explains why exact planning is intractable and develops the deep-RL answer of recurrent, history-based policies.\n",{"title":1307,"path":1308,"lessonNumber":1309,"topics":1310,"summary":1311},"Partial Observability: Planning and Recurrent Policies","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fpartial-observability-pomdps-part-2",19,[1171],"A companion to the belief-state lesson. In principle a POMDP reduces to an MDP over beliefs; in practice two obstacles block that. Exact planning over the belief simplex is intractable — the value function is piecewise-linear-and-convex with a number of pieces that can explode — and computing the belief needs a model the agent rarely has. This lesson develops the intractability, the point-based approximations that address it, and the deep-RL answer: make the policy a function of history with a recurrent network (DRQN, R2D2), with frame-stacking, attention, and world-model latents as learned beliefs.\n",{"title":1313,"path":1314,"lessonNumber":1315,"topics":1316,"summary":1317},"Safe and Constrained RL: The CMDP and Policy Methods","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fsafe-and-constrained-rl",20,[1171],"Maximizing a scalar reward is not the same as behaving well: a capable optimizer will find and exploit any gap between the reward and what its designer actually meant, a failure called specification gaming or reward hacking. The remedy is to add explicit cost constraints — the constrained MDP — maximizing return subject to an expected-cost budget. This lesson builds the core toolkit: the CMDP itself, Lagrangian primal-dual methods that learn a multiplier on the constraint (RCPO), and constrained policy optimization (CPO) with its trust-region cost bound. A companion lesson covers risk-sensitivity, safe exploration, and the alignment framing.\n",{"title":1319,"path":1320,"lessonNumber":1321,"topics":1322,"summary":1323},"Safe RL: Risk, Safe Exploration, and Alignment","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fsafe-and-constrained-rl-part-2",21,[1171],"A companion to the constrained-MDP lesson. Constraining the mean cost is not enough: a policy safe on average can be catastrophic in the tail, and a policy safe at convergence can violate its limits wildly while learning. This lesson optimizes the tail with risk-sensitive objectives (CVaR), then makes exploration itself safe with shields, Lyapunov methods, and safety layers that project unsafe actions onto the feasible set — closing with benchmarks, safe RLHF, robustness, and the alignment framing that ties safety back to the problem of incompletely specified reward.\n",{"title":1325,"path":1326,"lessonNumber":1327,"topics":1328,"summary":1329},"Meta-RL and Generalization","\u002Freinforcement-learning\u002Fmodern-deep-rl\u002Fmeta-rl-and-generalization",22,[1171],"An agent that masters one task often fails on the next; it has overfit to a single environment. This lesson treats fast adaptation as a meta-problem over a distribution of tasks: meta-train so that a few episodes at meta-test time suffice. We cover the two families — optimization-based (MAML learns an initialization) and context-based (RL-squared and PEARL infer a latent task) — the exploration cost of adaptation, and the parallel problem of generalization: why deep RL memorizes environments and what fixes it (domain randomization, procedural generation, augmentation, regularization). It closes on foundation models and sequence-model agents as the generalist endpoint.\n",{"module":1331,"moduleNumber":1012,"slug":1332,"lessons":1333},"Reinforcement Learning in Minds and Brains","minds-and-brains",[1334,1340,1345,1350,1355,1360,1365,1370],{"title":1335,"path":1336,"lessonNumber":978,"topics":1337,"summary":1339},"The Psychology of Reinforcement","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fpsychology-of-reinforcement",[1338],"Minds and Brains","Reinforcement learning is both an engineering method and a theory of how animals learn. The prediction\u002Fcontrol split of the algorithms mirrors the psychologist's split between classical and instrumental conditioning. We trace the correspondence: the Rescorla–Wagner model as a prediction-error rule that explains blocking, its real-time TD extension, Thorndike's Law of Effect behind trial-and-error control, and the habitual\u002Fgoal-directed distinction that maps onto model-free versus model-based learning.\n",{"title":1341,"path":1342,"lessonNumber":12,"topics":1343,"summary":1344},"The Psychology of Reinforcement: Instrumental Control","\u002Freinforcement-learning\u002Fminds-and-brains\u002Finstrumental-conditioning-and-control",[1338],"Classical conditioning was prediction; instrumental conditioning is control. Thorndike's Law of Effect is trial-and-error control — selection plus association, search plus memory — and Skinner's shaping and schedules are reward engineering. The habitual\u002Fgoal-directed distinction maps onto model-free versus model-based control, dissociated by outcome devaluation and arbitrated by uncertainty. Delayed reinforcement is the credit-assignment problem, and the stimulus traces and secondary reinforcers of animal-learning theory are eligibility traces and value functions.\n",{"title":1346,"path":1347,"lessonNumber":994,"topics":1348,"summary":1349},"Dopamine and the TD Error","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fdopamine-and-td-error",[1338],"The TD error was invented as an algorithm; a decade later it turned out to closely describe the firing of the brain's dopamine neurons. We follow Schultz's experiments — dopamine fires at an unpredicted reward, shifts to the earliest predictive cue, and dips below baseline when a predicted reward is withheld — and match each result to the TD error term by term. We then read the basal ganglia as a neural actor–critic with dopamine as its shared training signal, and close on addiction as a hijacking of that signal.\n",{"title":1351,"path":1352,"lessonNumber":1000,"topics":1353,"summary":1354},"Dopamine in the Brain: The Neural Actor–Critic","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fdopamine-in-the-brain",[1338],"If phasic dopamine is a TD error, where does it go and what does it change? We follow the axons into the basal ganglia, read the corticostriatal synapse as the place where state, action, and error meet, and map the ventral and dorsal striatum onto the critic and the actor of an actor–critic. Addiction becomes a broken cancellation in the same learning signal, and distributional dopamine extends the scalar RPE into a population code.\n",{"title":1356,"path":1357,"lessonNumber":1006,"topics":1358,"summary":1359},"Animal Learning and Cognition","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fanimal-learning-and-cognition",[1338],"Three classic associative phenomena turn out to be reinforcement-learning mechanisms seen in behavior. Blocking says learning is driven by prediction error, not co-occurrence, and reduces to least-squares regression fitting a collinear feature. Higher-order conditioning and conditioned reinforcement make a value estimate a secondary reinforcer — bootstrapping in an animal. Delayed reinforcement is the credit-assignment problem, and the stimulus traces and goal gradients of Pavlov and Hull are eligibility traces and TD-learned value functions.\n",{"title":1361,"path":1362,"lessonNumber":1012,"topics":1363,"summary":1364},"Cognitive Maps and Model-Based Learning","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fcognitive-maps-and-planning",[1338],"Tolman's rats learned the layout of a maze with no reward, then used it the moment food appeared — latent learning, a cognitive map, and the behavioral face of model-based reinforcement learning. The map is learned by system identification (stimulus–stimulus associations), which fills in whether or not reward is present, and queried by planning, which re-solves a route from a single changed reward. The successor representation sits between cache and model, and hippocampal predictive maps and scaled-up world models carry the same idea into brain and machine.\n",{"title":1366,"path":1367,"lessonNumber":1053,"topics":1368,"summary":1369},"The Neuroscience of Reinforcement","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fneuroscience-of-reinforcement",[1338],"The dopamine story is one contact point between reinforcement learning and the brain; this lesson fills in the surrounding neuroscience so the mapping stands on its own. We build a working primer of neurons, synapses, and neuromodulation; separate four signals that casual usage conflates — reward, reinforcement, value, and prediction error; and read the actor and critic as corticostriatal synapses updated by two- and three-factor rules, grounded in spike-timing-dependent and reward-modulated plasticity.\n",{"title":1371,"path":1372,"lessonNumber":1059,"topics":1373,"summary":1374},"The Brain's Several Learning Systems","\u002Freinforcement-learning\u002Fminds-and-brains\u002Fseveral-learning-systems",[1338],"The actor's three-factor rule has an ancestor in Klopf's hedonistic neuron — a single cell as a reinforcement-seeking agent — and a bacterium's run-and-twiddle shows the Law of Effect with no synapses at all. Teams of such neurons implement policy gradient collectively, the broadcast reward replacing backpropagation. And the brain is not only model-free: outcome devaluation, prefrontal value coding, and hippocampal forward sweeps localize a model-based system. The recurring conclusion is that the brain is several interacting learning systems, not one algorithm.\n",1785117679012]