<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://shawnxu0913.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://shawnxu0913.github.io/" rel="alternate" type="text/html" /><updated>2026-10-08T23:16:28+00:00</updated><id>https://shawnxu0913.github.io/feed.xml</id><title type="html">Shawn Xu</title><subtitle>Notes on building game-playing agents.</subtitle><author><name>Shawn Xu</name></author><entry><title type="html">Milestone 1: Solving Slay the Spire 2 at Ascension 0</title><link href="https://shawnxu0913.github.io/2026/10/07/solving-sts2-ascension-0.html" rel="alternate" type="text/html" title="Milestone 1: Solving Slay the Spire 2 at Ascension 0" /><published>2026-10-07T00:00:00+00:00</published><updated>2026-10-07T00:00:00+00:00</updated><id>https://shawnxu0913.github.io/2026/10/07/solving-sts2-ascension-0</id><content type="html" xml:base="https://shawnxu0913.github.io/2026/10/07/solving-sts2-ascension-0.html"><![CDATA[<p>This is a write up of my personal project: Building an AI to play Slay the Spire 2. Building autonomous AI agents to play games has always been a hobby of mine, and in the past I’ve made AIs to play chess, poker, scrabble, RTS video games, and various other projects. Another motivation is that I wanted to gauge how well frontier LLM agents can (largely) navigate an open machine learning problem such as this one (my day job involves training LLM agents at one of the frontier labs). I’ve used Claude and Gemini extensively in this project as research partners and coders. However, note that I am not interested in using an LLM <em>itself</em> to act as the game agent, even though they will eventually get there. Instead, this project aims at using LLMs to build a <em>small</em> game AI.</p>

<p>I started this project in March 2026. When I started, “solving” STS2 still seemed out of reach beyond my wildest imagination. Sts1 has been out for 7 years and many AI attempts have been made, but none was convincing. I was inspired to post this after seeing <a href="https://www.reddit.com/r/slaythespire/comments/1vxfsf4/jorbs_actually_built_a_fight_solver_for_sts_2/">Jorbs’ recent post</a>.</p>

<p><em>Headline first</em>: The autonomous agent now wins about 87% of Ironclad runs at Ascension 0, up from about 10% in August 2026. I consider A0 effectively solved and am declaring Milestone 1 complete.</p>

<p><em>The runs</em>: I’ve published every run from a fresh batch of 4,960 random seeds, losses included, with each action and the game state after it recorded: <a href="https://github.com/shawnxu0913/sts2ai-runs/releases/tag/v1.0">4,251 wins out of 4,960 (85.7%)</a>. The format is described in the <a href="https://github.com/shawnxu0913/sts2ai-runs">repository’s README</a>.</p>

<p><em>Important caveat</em>: This AI does “cheat” via “save scumming”. The game engine is identical to the real game, including all deterministic rollouts, so the agent knows exactly what cards will be drawn the next turn, etc. Technically, humans have access to this information as well, via save scumming. This was an early decision – I wanted to make the problem easier first. Future milestones will aim at removing this extra information.</p>

<h2 id="the-result">The result</h2>

<figure>
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 760 370" font-family="-apple-system, Segoe UI, Helvetica, Arial, sans-serif" font-size="12" role="img" style="width:100%;height:auto;display:block" aria-label="Win rate rose from 10% to 87% in six weeks">
<rect width="760" height="370" fill="white" />
<text x="20" y="26" font-size="16" font-weight="600" fill="#1f2328">Win rate rose from 10% to 87% in six weeks</text>
<text x="20" y="46" fill="#59636e">Promoted agent's win rate in each A/B test (Ironclad, Ascension 0). Each test used its own seed pool, so small dips are noise.</text>
<line x1="70" x2="650" y1="330.0" y2="330.0" stroke="#e6e6e6" />
<text x="60" y="334.0" text-anchor="end" fill="#59636e">0%</text>
<line x1="70" x2="650" y1="267.5" y2="267.5" stroke="#e6e6e6" />
<text x="60" y="271.5" text-anchor="end" fill="#59636e">25%</text>
<line x1="70" x2="650" y1="205.0" y2="205.0" stroke="#e6e6e6" />
<text x="60" y="209.0" text-anchor="end" fill="#59636e">50%</text>
<line x1="70" x2="650" y1="142.5" y2="142.5" stroke="#e6e6e6" />
<text x="60" y="146.5" text-anchor="end" fill="#59636e">75%</text>
<line x1="70" x2="650" y1="80.0" y2="80.0" stroke="#e6e6e6" />
<text x="60" y="84.0" text-anchor="end" fill="#59636e">100%</text>
<text x="124.6341463414634" y="352" text-anchor="middle" fill="#59636e">Sep 1</text>
<text x="315.8536585365854" y="352" text-anchor="middle" fill="#59636e">Sep 15</text>
<text x="534.390243902439" y="352" text-anchor="middle" fill="#59636e">Oct 1</text>
<path d="M70.0 306.0 L83.7 296.0 L90.5 290.5 L97.3 283.8 L247.6 269.0 L329.5 260.8 L411.5 246.0 L425.1 198.8 L438.8 205.8 L493.4 213.5 L507.1 164.0 L520.7 155.0 L575.4 135.8 L616.3 124.0 L624.5 112.5" fill="none" stroke="#b9b9b0" stroke-width="2" />
<g data-tip="Aug 28: Baseline before survival-based route planning — 9.6% win rate" style="cursor:pointer"><circle cx="70.0" cy="306.0" r="5" fill="#b9b9b0" /><circle cx="70.0" cy="306.0" r="9" fill="transparent" /></g>
<text x="60.0" y="310.0" text-anchor="end" font-weight="600" fill="#59636e">Start 10%</text>
<g data-tip="Aug 29: Per-floor replanning + engine-forked route enumeration — 13.6% win rate" style="cursor:pointer"><circle cx="83.7" cy="296.0" r="3" fill="#b9b9b0" /><circle cx="83.7" cy="296.0" r="9" fill="transparent" /></g>
<g data-tip="Aug 29: Joint deck-value scorer — 15.8% win rate" style="cursor:pointer"><circle cx="90.5" cy="290.5" r="3" fill="#b9b9b0" /><circle cx="90.5" cy="290.5" r="9" fill="transparent" /></g>
<g data-tip="Aug 30: Future-boss survival in route value — 18.5% win rate" style="cursor:pointer"><circle cx="97.3" cy="283.8" r="3" fill="#b9b9b0" /><circle cx="97.3" cy="283.8" r="9" fill="transparent" /></g>
<g data-tip="Sep 10: Potion-aware leaf evaluation — 24.4% win rate" style="cursor:pointer"><circle cx="247.6" cy="269.0" r="3" fill="#b9b9b0" /><circle cx="247.6" cy="269.0" r="9" fill="transparent" /></g>
<g data-tip="Sep 16: Forkrank card picks — 27.7% win rate" style="cursor:pointer"><circle cx="329.5" cy="260.8" r="3" fill="#b9b9b0" /><circle cx="329.5" cy="260.8" r="9" fill="transparent" /></g>
<g data-tip="Sep 22: Forkrank retrained on its own runs — 33.6% win rate" style="cursor:pointer"><circle cx="411.5" cy="246.0" r="3" fill="#b9b9b0" /><circle cx="411.5" cy="246.0" r="9" fill="transparent" /></g>
<g data-tip="Sep 23: Causal event table — 52.5% win rate" style="cursor:pointer"><circle cx="425.1" cy="198.8" r="5" fill="#2f6fd6" /><circle cx="425.1" cy="198.8" r="9" fill="transparent" /></g>
<text x="415.1" y="186.8" text-anchor="end" font-weight="600" fill="#2f6fd6">Event table 52%</text>
<g data-tip="Sep 24: Planner uses the same pick model — 49.7% win rate" style="cursor:pointer"><circle cx="438.8" cy="205.8" r="3" fill="#b9b9b0" /><circle cx="438.8" cy="205.8" r="9" fill="transparent" /></g>
<g data-tip="Sep 28: Relics over cards in shops — 46.6% win rate" style="cursor:pointer"><circle cx="493.4" cy="213.5" r="3" fill="#b9b9b0" /><circle cx="493.4" cy="213.5" r="9" fill="transparent" /></g>
<g data-tip="Sep 29: Playout gate — 66.4% win rate" style="cursor:pointer"><circle cx="507.1" cy="164.0" r="5" fill="#2f6fd6" /><circle cx="507.1" cy="164.0" r="9" fill="transparent" /></g>
<text x="497.1" y="152.0" text-anchor="end" font-weight="600" fill="#2f6fd6">Playout gate 66%</text>
<g data-tip="Sep 30: Potion and racing fallback rungs — 70.0% win rate" style="cursor:pointer"><circle cx="520.7" cy="155.0" r="3" fill="#b9b9b0" /><circle cx="520.7" cy="155.0" r="9" fill="transparent" /></g>
<g data-tip="Oct 4: Strict defensive potions — 77.7% win rate" style="cursor:pointer"><circle cx="575.4" cy="135.8" r="3" fill="#b9b9b0" /><circle cx="575.4" cy="135.8" r="9" fill="transparent" /></g>
<g data-tip="Oct 7: Turn beam — 82.4% win rate" style="cursor:pointer"><circle cx="616.3" cy="124.0" r="5" fill="#2f6fd6" /><circle cx="616.3" cy="124.0" r="9" fill="transparent" /></g>
<text x="628.3" y="136.0" text-anchor="start" font-weight="600" fill="#2f6fd6">Turn beam 82%</text>
<g data-tip="Oct 7: Turn-beam variants — 87.0% win rate" style="cursor:pointer"><circle cx="624.5" cy="112.5" r="5" fill="#2f6fd6" /><circle cx="624.5" cy="112.5" r="9" fill="transparent" /></g>
<text x="636.5" y="108.5" text-anchor="start" font-weight="600" fill="#2f6fd6">Beam variants 87%</text>
</svg>
  <figcaption>Fleet A/B tests, Aug 28 – Oct 7 2026; promoted arm of each test, about 1,000 seeds per arm. Hover a point to see the change it shipped.</figcaption>
</figure>

<p>Three ideas account for most of the climb: choosing events by their measured effect on winning, simulating each fight before playing it, and the turn beam. Every point is a change that won a paired test against the agent before it.</p>

<h2 id="how-it-works">How it works</h2>

<p>The agent is a stack of separate components: combat is solved by search, and the run-level decisions use learned models plus search over the map. No component is hand-scripted for specific fights.</p>

<p><strong>Combat</strong></p>

<ul>
  <li><strong>Combat search agent.</strong> A 3-turn lookahead search over card plays, targets and card-selection prompts, run inside an exact copy of the game engine. It branches on every meaningful choice and undoes moves with an undo log.</li>
  <li><strong>Learned leaf evaluation.</strong> A <a href="https://en.wikipedia.org/wiki/Gradient_boosting">gradient-boosted</a> model predicts the HP the player will end the fight with, from the board state (hand, piles, powers, enemy intents, potions). Search uses it to score positions three turns out.</li>
  <li><strong>Playout gate.</strong> At the start of each fight, the agent plays the whole fight out in simulation with its default strategy. If that playout dies, it tries a ladder of alternative strategies (deeper search, racing the enemy, drinking potions) and adopts the first one that wins. The winning line is cached and replayed turn by turn.</li>
  <li><strong>Turn beam.</strong> The last rungs of that ladder. Instead of committing to one plan per turn, it keeps the best 40 to 80 end-of-turn positions alive across turns. This finds setup turns and burst windows that look bad in the short term, which the per-turn search prunes away. Three score variants (racing, scaling, wider) cover fights the default misses.</li>
</ul>

<p><strong>Between fights</strong></p>

<ul>
  <li><strong>Act path planner.</strong> A beam search over the act’s map that forks the engine to simulate each route’s fights, rests, shops and events. It scores routes by a learned estimate of surviving the act and the bosses after it, and replans every floor.</li>
  <li><strong>Forkrank (card picks).</strong> A ranking model trained on counterfactuals. For each card reward, the run is forked and continued with each option, and the model learns which pick actually leads to wins. It takes over card picks from floor 17 on.</li>
  <li><strong>Event table.</strong> For each event option, the causal effect of taking it on winning, estimated from forked continuations of the first step only.</li>
  <li><strong>Shop, potions and rest.</strong> Learned shop picks with a rule that buys affordable relics over cards, card removal of the weakest card, and rule-based potion use (for example, defensive potions only when the next hit is lethal).</li>
</ul>

<figure>
  <img src="/assets/img/architecture.svg" alt="Learned models plan the run; search plays each fight" />
  <figcaption>Agent architecture: four run-level components, a four-step fight pipeline, and the shared simulator.</figcaption>
</figure>

<p>Every component queries the same simulator: the planner forks the run to try routes, forkrank forks it to try card picks, and the fight pipeline plays whole fights out before committing.</p>

<h2 id="the-math-behind-it">The math behind it</h2>

<p>Both halves of the agent search a model of the game: combat searches the exact game with a learned value at the leaves, and the run-level planner searches routes scored by learned survival probabilities.</p>

<h3 id="combat-deterministic-search-with-a-learned-leaf-value">Combat: deterministic search with a learned leaf value</h3>

<p>A fight is a sequence of states \(s\) and actions \(a\) (play card \(i\) on target \(j\), drink a potion, end the turn). The simulator gives the next state, \(f(s, a)\), including the enemies’ moves. Because it reproduces the game’s random number generator exactly, the search treats each fight as deterministic. The objective is the HP the player ends the fight with.</p>

<p>The default policy looks three turns ahead and picks the action sequence whose end state scores highest:</p>

\[a^{*} = \arg\max_{\pi \in \Pi_{3}(s_0)} \hat{V}\big(s^{\pi}\big), \qquad
\hat{V}(s) = \begin{cases} \mathrm{HP}(s) &amp; \text{fight won} \\ -\infty &amp; \text{player dead} \\ \mathrm{HP}(s) + g_{\theta}(\phi(s)) &amp; \text{otherwise} \end{cases}\]

<p>Here \(\Pi_3(s_0)\) is the set of action sequences covering the next three turns, \(\phi(s)\) is a feature vector of the board (hand, piles, powers, enemy intents, potions), and \(g_\theta\) is a <a href="https://en.wikipedia.org/wiki/Gradient_boosting">gradient-boosted tree</a> model trained to predict the HP still to be lost, \(\mathrm{HP}_{\text{final}} - \mathrm{HP}(s)\). The search re-plans every turn.</p>

<p><strong>Playout gate.</strong> Before the first turn, the agent plays the entire fight out under an ordered list of strategies \(\sigma_1, \dots, \sigma_K\). Each playout returns whether it won, the final HP and how many turns it survived. The agent adopts the first strategy that wins (or, if none does, the one that survived longest) and replays its line turn by turn:</p>

\[k^{*} = \min\{\, k : \mathrm{won}(\sigma_k) \,\}\]

<p><strong>Turn beam.</strong> The last strategies in the list replace one-plan-per-turn with a beam over end-of-turn states. \(B_t\) holds the \(W\) best states after turn \(t\); each is expanded by every distinct full turn \(\tau\) (a sequence of plays ending in end turn):</p>

\[B_{t+1} = \operatorname{top}_W \big\{\, f(s, \tau) : s \in B_t,\ \tau \in \mathcal{T}(s) \,\big\}, \qquad
h(s) = 3\,\mathrm{HP} - \sum_{e}(\mathrm{HP}_e + \mathrm{Block}_e) + 8\,\mathrm{Str} + 6\,n_{\mathrm{buffs}} + 20\,n_{\mathrm{dead}}\]

<p>The score \(h\) is deliberately simple and needs no learned model. Its variants reweight it: less weight on HP to race the enemy, more weight on Strength and buffs to scale, or a wider beam.</p>

<h3 id="between-fights-planning-against-learned-survival-probabilities">Between fights: planning against learned survival probabilities</h3>

<p>The run-level objective is the probability of winning the run. The act path planner runs a beam search over routes \(p\) through the act’s map (beam width 32) and scores each route roughly as</p>

\[J(p) = \prod_{r \in p} \Pr\big(L_r &lt; \mathrm{HP}_r\big) \cdot \prod_{b \in \text{future bosses}} q_{\psi}\big(b,\ \mathrm{deck}_{\text{end}},\ \mathrm{HP}_{\text{end}}\big)\]

<p>\(L_r\) is the HP lost in room \(r\), with its distribution per encounter estimated from past fights (censored Kaplan–Meier curves). \(q_\psi\) is a learned model of beating a future boss with a given deck and HP. The planner simulates each route’s rooms in a fork of the engine, and replans every floor.</p>

<p><strong>Forkrank.</strong> At a card reward with options \(c_1, \dots, c_m\), the run is forked once per option and played to the end, giving a win label \(y_c\) for each. A scoring model \(r_\omega\) is trained on pairs where the outcomes disagree, with a pairwise logistic (RankNet) loss, and the agent picks the highest-scoring option:</p>

\[\mathcal{L}(\omega) = -\sum_{(i,j)\,:\,y_i &gt; y_j} \log \sigma\big(r_{\omega}(x, c_i) - r_{\omega}(x, c_j)\big)\]

<p><strong>Event table.</strong> For each event and option, the table stores the estimated effect of choosing that option on the chance of winning, measured by forking the run at the event and continuing with each option. Only the event’s first step is used, which avoided push-your-luck chains that looked good on average but lost runs.</p>

<h3 id="future-direction-playing-with-only-what-a-player-can-see">Future direction: playing with only what a player can see</h3>

<p>Everything above searches the exact engine, random number generator included. Inside a search, the agent therefore knows things a human player cannot: the order of the draw pile, which cards a potion or event will generate, and the enemies’ moves beyond the intent shown for the next turn. A natural next step is an agent that plays with only the information a real player has, which also makes the comparison with human play fair.</p>

<p>Formally, a fight becomes a partially observable problem. The player sees an observation \(o\): the hand, the contents (but not the order) of the draw and discard piles, HP, powers, potions, and each enemy’s current intent. The true state \(s\) adds the hidden part, mainly the generator’s state. The agent keeps a belief \(b(s \mid o)\) over the states consistent with what it has seen.</p>

<p>The simplest change is <strong>determinization</strong>: sample \(K\) plausible worlds from the belief (shuffle the draw pile, reseed the generator, keep everything observed), search each one with the existing engine, and pick the action that does best on average:</p>

\[a^{*} = \arg\max_{a} \; \frac{1}{K} \sum_{k=1}^{K} \hat{V}\big(f_k(s_k, a)\big), \qquad s_k \sim b(\,\cdot \mid o)\]

<p>Each component would change accordingly:</p>

<ul>
  <li><strong>Combat search</strong> turns end-of-turn draws into chance nodes. Searching each sampled world separately is cheap to build but can assume it will know the future once it gets there. Searching over information sets, so that one plan has to work across all the worlds (information-set Monte Carlo tree search), fixes that at a higher cost.</li>
  <li><strong>The leaf evaluation</strong> needs no retraining in principle. Its features are mostly observable already; the hidden draw order would be removed.</li>
  <li><strong>The playout gate</strong> can no longer adopt the first strategy that wins one deterministic playout. It would play each strategy out in \(K\) sampled worlds, adopt the one with the highest estimated chance of winning, \(\hat{p}(\sigma) = \tfrac{1}{K}\sum_k \mathrm{won}_k(\sigma)\), and re-plan every turn instead of replaying a cached line.</li>
  <li><strong>The turn beam</strong> would score each end-of-turn state by its average over sampled worlds, so a setup turn survives only if it pays off in most of them.</li>
  <li><strong>Between fights</strong>, the planner and forkrank also fork the true engine and see future card rewards and event outcomes. They would sample those futures too; forkrank’s training labels would come from many continuations per option rather than one.</li>
</ul>

<p>The cost is roughly \(K\) times more compute per decision, and some loss of win rate from no longer knowing the future. Measuring that gap with the same paired A/B tests would tell us how much of the 87% depends on hidden information, and how much an honest agent can recover.</p>

<h3 id="deciding-what-ships">Deciding what ships</h3>

<p>Each change runs against the current agent on the same 1,000 seeds. With \(b\) runs rescued (new wins) and \(c\) runs thrown (new losses), it ships when a continuity-corrected McNemar test is significant and it still comes out ahead with unfinished runs counted as losses:</p>

\[\chi^2 = \frac{\big(\lvert b - c\rvert - 1\big)^2}{b + c}\]

<h2 id="timeline-of-the-big-jumps">Timeline of the big jumps</h2>

<p>The biggest gains came from changing what the agent optimizes or how it chooses, not from searching deeper. Each step below shipped only after a paired A/B test on 1,000 seeds per arm; the lift is the win-rate gain over the agent it replaced.</p>

<table>
  <thead>
    <tr>
      <th>Shipped</th>
      <th>Idea</th>
      <th>Win rate (before → after)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Oct 7</td>
      <td><strong>Turn-beam variants</strong>: racing, scaling and wider beams for fights the default beam misses</td>
      <td>83% → 87%</td>
    </tr>
    <tr>
      <td>Oct 7</td>
      <td><strong>Turn beam</strong>: keep many end-of-turn positions alive, so setup turns survive until they pay off. Distilled from how Claude won fights the agent had lost</td>
      <td>76% → 82%</td>
    </tr>
    <tr>
      <td>Oct 4</td>
      <td><strong>Strict defensive potions</strong>: drink block or healing potions only when they prevent a lethal hit</td>
      <td>74% → 78%</td>
    </tr>
    <tr>
      <td>Sep 29 – Oct 2</td>
      <td><strong>Playout gate</strong>: simulate each whole fight before playing it, switch strategy if the default dies; later added potion and racing fallbacks and a replay cache</td>
      <td>57% → 74%</td>
    </tr>
    <tr>
      <td>Sep 28</td>
      <td><strong>Relics over cards in shops</strong></td>
      <td>42% → 47%</td>
    </tr>
    <tr>
      <td>Sep 24</td>
      <td><strong>Consistent picks</strong>: the map planner uses the same pick model the agent plays with</td>
      <td>45% → 50%</td>
    </tr>
    <tr>
      <td>Sep 23</td>
      <td><strong>Causal event table</strong>: choose event options by their measured effect on winning</td>
      <td>36% → 53%</td>
    </tr>
    <tr>
      <td>Sep 16 – 22</td>
      <td><strong>Forkrank</strong>: learn card picks from forked counterfactual runs, then retrain on the agent’s own runs</td>
      <td>22% → 34%</td>
    </tr>
    <tr>
      <td>Sep 10</td>
      <td><strong>Potion-aware leaf evaluation</strong></td>
      <td>23% → 24%</td>
    </tr>
    <tr>
      <td>Aug 29 – 30</td>
      <td><strong>Survival-based route planning</strong>: replan every floor, score routes by the chance of surviving the act and future bosses</td>
      <td>10% → 19%</td>
    </tr>
  </tbody>
</table>

<p>Each A/B used its own seed pool and the agent kept improving between tests, so the “before” of one row does not exactly match the “after” of the previous one. Earlier work (June to August) built the simulator, the combat search and the first learned models; it moved the agent from dying mid-act to reaching bosses.</p>

<h2 id="what-made-it-possible">What made it possible</h2>

<p>Three pieces of infrastructure carry everything above.</p>

<ul>
  <li><strong>A headless simulator built from the game’s own code.</strong> The game’s logic is compiled directly, with the graphics and audio stubbed out, so the agent searches in the real rules rather than a reimplementation. It runs fully deterministically, which lets the agent fork a run, try every option, and rewind.</li>
  <li><strong>Parity checking against the real game.</strong> Recorded runs from the live game are replayed step by step in the simulator and compared state by state. About 1,300 recorded runs replay cleanly, and a separate check confirms that what the search predicts matches what actually happens when the plan is played.</li>
  <li><strong>Fleet A/B testing.</strong> Every change is tested on a cluster against the current agent on 1,000 fresh seeds per arm, paired by seed. It ships only if it wins significantly and still comes out ahead when unfinished runs are counted as losses. This kept us from shipping ideas that looked good on a handful of seeds, and caught one test that ran on a broken setup.</li>
</ul>

<p>Claude also worked as an offline teacher. It played fights the agent had lost, through a step-by-step interface into the simulator. Its winning lines showed which strategies were missing, and the turn beam turned the most common one into an automatic search.</p>

<h2 id="whats-next">What’s next</h2>

<p>About 13% of runs are still lost, half of them at the final boss. The open problems:</p>

<ul>
  <li><strong>Removing save scumming.</strong> Future milestones will only be published if the agent wins without knowing the random rollouts.</li>
  <li><strong>Entry HP into bosses.</strong> Arriving with 10 more HP turns about a third of the fights we lose into wins. We are testing a gate that picks the line that keeps the most HP in ordinary fights before a boss.</li>
  <li><strong>Teaching the leaf evaluation.</strong> The turn beam’s winning lines can train the evaluation model, so the everyday search finds these lines without the slow fallback.</li>
  <li><strong>Higher Ascensions and other characters.</strong> Milestone 1 covers the Ironclad at Ascension 0. Higher Ascensions add harder enemies and fewer resources, and the other characters need their own card knowledge.</li>
</ul>]]></content><author><name>Shawn Xu</name></author><summary type="html"><![CDATA[This is a write up of my personal project: Building an AI to play Slay the Spire 2. Building autonomous AI agents to play games has always been a hobby of mine, and in the past I’ve made AIs to play chess, poker, scrabble, RTS video games, and various other projects. Another motivation is that I wanted to gauge how well frontier LLM agents can (largely) navigate an open machine learning problem such as this one (my day job involves training LLM agents at one of the frontier labs). I’ve used Claude and Gemini extensively in this project as research partners and coders. However, note that I am not interested in using an LLM itself to act as the game agent, even though they will eventually get there. Instead, this project aims at using LLMs to build a small game AI.]]></summary></entry></feed>