📖 中文版在下方 — scroll down for the Chinese version.
Written September 2026. Numbers are from the bot's game database as of September 26, 2026.
Grandmaster Spider in Microsoft Solitaire Collection is two suits, 104 cards, and a lot of deals where only a handful of lines win. I wanted a bot that plays it on my Android phone, wins every deal it starts, and keeps going for hours without me.
The bot is not allowed anything a human doesn't have: it takes a screenshot, reads the board, plans, drags a card, and checks the result. No memory reading, no game API.
As of September 26 its game database recorded 20 deals finished, 20 won, over 153 attempts and 11,052 moves.
This post is about the problems that mattered, what I tried, and what actually worked — including the things that didn't.
1. The shape of the problem
Spider with two suits: 54 cards dealt into 10 columns (44 of them face down), 50 more in the stock, dealt 10 at a time. A same-suit run from King to Ace is removed; remove 8 runs and you win. At the start you see exactly 10 cards.
Constraints:
- ADB only:
screencap,input tap,input swipe. Everything else is inferred from pixels. - A 2376×1080 landscape phone. Long columns get squeezed, a toolbar slides up over the bottom, and the system navigation bar covers part of the screen.
- Never give up on a deal, and run unattended, game after game.
2. Architecture
adb screencap ──► vision.py ──► brain.py ─────────────► executor.py ──► adb tap / swipe
▲ screen → exact solve drag, verify, │
│ Observation → guessed solve undo, dialogs │
│ → promising line │ │
│ → beam search ▼ │
│ │ ▲ knowledge.py │
│ ▼ │ (card memory) │
│ fastsolver (Rust) record.py │
│ full-game search (games.db) │
└────────────────────────────────────────────────────────────────────────────┘
| Module | Job |
|---|---|
vision.py | screenshot → observation: face-down count and face-up cards per column, stock deals left, toolbar, dialogs |
rules.py / model.py | pure rules engine: legal moves, apply a move (auto-flip, auto-complete), Zobrist hash |
knowledge.py | card memory: ("H", col, depth) and ("D", deals_left, col) → card |
fastsolver/ | Rust full-game solver |
brain.py | exact solve → guessed solve → most promising line → Python beam search; rewind planning |
executor.py | drag, verify, retry / undo / resync, dialogs, Undo All, new-game flow, smart rewind |
record.py | every deal, attempt and move into SQLite |
3. Vision: find the edges, don't trust fixed offsets
As a column grows, the game squeezes the cards together, so no fixed vertical offset survives. What does survive: every card's top edge is a 2–4 px light-grey line across the card, face-down cards are blue below it and face-up cards are white. Each column gets a top-to-bottom edge scan that counts face-down and face-up cards.
A covered card only shows a thin strip, so the rank is template-matched on the top-left glyph and the suit on the small pip at the top right. The templates are position-aligned and built from labelled screenshots.
A single glyph can be wrong, so the board is cross-checked as a whole: it must add up to exactly 104 cards (including completed runs), and no card may appear more often than it exists. Once, one card at 0.748 against a 0.75 threshold halted the bot although every board-level check passed. Now "board consistent, a few cards just under the threshold" is accepted and logged.
The pitfalls were all about things covering the board:
- The navigation bar covers the bottom of the stock. Count the stacked card tops instead of looking at the bottom of the pile.
- A long column hid the toolbar from the detector. The "toolbar is up" check looked at one small patch under the "New" button. A long column hanging over exactly that patch made a raised toolbar invisible — so the bot couldn't raise it, and the column scan ran into the toolbar and read it as one extra card (21 instead of 20), and then the bot tried to undo a perfectly good move. The fix is the median of strips across the whole toolbar width: on 158 old screenshots it is ≤ 64 with the toolbar and ≥ 123 without.
- A very long column can't be read cold. At ~30 px per card the glyphs are half cut off. At start-up the bot now takes the last position the previous run saved (and the position after the move it was about to make), and uses it if the screen agrees.
4. Execution: moving one card reliably
- Too fast is a fling. A fast long swipe is treated as a fling and the game places the cards somewhere else. Finger speed stays under ~2.5 px/ms, never shorter than 200 ms.
- The game swallows taps. A tap that only clears a hint glow does nothing else, and neither does the first tap after the app has been idle. Every button goes through
tap_until: tap, look, tap again. - The toolbar arrow sits under the navigation bar, so it never gets the tap. Tapping empty felt raises the toolbar instead.
- Verify every move. The rules engine predicts the next position; a screenshot is compared against it. On a mismatch: re-read → drag again slowly → undo and resync from the screen.
- Batch certain moves. Moves whose outcome is certain are sent back to back while the source column's geometry is still valid, then checked once.
That lands at about 2.4 s per verified move. The screenshot alone is ~0.8 s.
5. The idea everything rests on: card memory
Undo All restores the deal to its opening with every card in the same place.
So a lost attempt is not wasted. Every face-down card it turned up and every card a stock deal produced is recorded:
("H", col, depth)— the card at a given depth of a column's face-down pile;("D", deals_left, col)— the card a given stock deal put on a given column.
The next attempt starts knowing more. The attempt counter is saved too, so a restarted bot keeps counting. A different opening clears the memory, and if memory ever contradicts the screen, the screen wins.
This one observation turns a hidden-information game into one you can learn deal by deal. It is worth more than any search optimisation in this post.
6. The solver
fastsolver is a parallel best-first search in Rust over the whole game. Known cards play normally; unknown cards are walls — once turned up they just sit there, and no line may depend on what they are.
What made it fast enough:
- A compact search tree. A node is parent + move (16 bytes); full positions are stored only every 8th depth and rebuilt by replay; boards are fixed-size arrays, no heap allocation per move. ~34 bytes per node instead of ~130, and ~3.5× faster.
- A staged portfolio. 12 weight sets with a small budget first (most positions fall in under a second), then 6 and 3 sets with deeper searches for the hard ones.
- Full-knowledge benchmark: 31 of 32. 30 random deals plus 2 hard real ones; real deal "C", unsolvable for the first version, now solves from the opening in ~60 s, verified by replay.
- Keep the "most promising line" short. Without a win the solver returns its best-scoring line. Charging 40 points per move stopped it from proposing 70 moves of tidying for a tiny gain: 245 → 208 moves per won deal in simulation, same win rate.
7. Guess mode: when walls become the problem
The biggest improvement came from reading the logs.
One deal went 14 attempts, about 89 minutes, without a win, and attempts 9–13 learned no new card at all — each replayed an almost identical ~65-move line. By then 88 of 94 hidden cards were known. From the cards not yet seen, the 6 unknown ones could only be {2♠, 6♠ ×3, 9♠, 9♥}: 120 arrangements.
An offline test settled it:
| Method | Result |
|---|---|
| Unknown cards as walls | no solution in 60 s |
| Fill the 6 cards at random, then solve (6 fillings) | 2 solved to a full win in ~12 s |
It wasn't bad luck. It was the planning. With walls, every line that has to pass through those six cards is forbidden — and the lines that remain are all dead ends. So the bot walked into the same dead end every attempt.
The fix is determinization. When ≤ 16 cards are unknown and the walls search finds nothing:
- Fill every unknown slot with a random draw from the cards not yet seen (the suits of completed runs are inferred from the remaining counts) and solve the whole deal.
- Follow the first guess that wins. Each time a real unknown card turns up, check it: right → keep following; wrong → guess again with one more card known.
- After a round of 6 failed guesses, don't try again until a new card is learned. Otherwise every move could cost a minute.
I also cached failed exact searches per position and knowledge. Before, every Undo All redid the 90-second opening search with nothing new to go on.
Results:
- Replaying the stuck deal offline, with the 6 unknown cards filled in as several possible "truths": the old logic lost all 5 over 5 attempts each, stuck at 88 known cards. The new logic won both fillings it finished, on attempt 1 and attempt 2.
- On the phone, another deal sat at 80 known cards for attempts 6–8 with the old logic. With guess mode, attempt 9 learned 3 new cards and attempt 10 won.
8. Undo, and rewinding to the cheapest good position
The game also has a plain Undo. An earlier experiment had used it the naive way — at a dead end, undo a fixed 8 moves and try another line — and it lost to simply restarting: 28/42 against 41/42. A local step back never re-plans the whole deal.
With guess mode in place I tried a different use. At a dead end, look for a winning line — with everything known, guesses included — from a few earlier positions of this attempt (5, 10, 20, 40 moves back) and from the opening, then go to the cheapest one:
cost = undos × 1.1 s + remaining moves × 2.4 s (Undo All = 6 s)
Undo All is just the "back to the opening" case, so this can't do worse than restarting.
First I checked that Undo is exact. 20 undos in a row on the phone, each compared with the position the database had recorded before that move: plain moves, two stock deals, one card flip (the card went face down again). 20 of 20 matched.
Then I made it fast. One undo at a time — raise the toolbar, tap, confirm with a screenshot — costs ~3.6 s. Tapping in batches, then reading the board once and locating it in this attempt's list of positions (which also corrects a swallowed or doubled tap), did 7 undos in 7.4 s: ~1.06 s per undo, cheaper than a move. The per-attempt list mirrors the game's undo stack; anything unplanned invalidates it and the bot falls back to Undo All.
Simulation, 42 random deals, guess mode on in both:
| Approach | Won | Actions per won deal (median / mean) |
|---|---|---|
| Undo All only | 42/42 | 255 / 319 |
| Smart rewind | 42/42 | 210 / 251 |
Were 5 / 10 / 20 / 40 the right numbers? Honestly, I picked them by feel. So I compared three candidate sets:
| Candidates | Won | Actions per win (median / mean) | Est. phone time | Probes per deal |
|---|---|---|---|---|
| fixed: 5/10/20/40 + opening | 42/42 | 202 / 260 | 10.3 min | 6 |
| events: before the last 6 deals/flips + opening | 42/42 | 230 / 272 | 10.8 min | 11 |
| dense: every 5 moves, up to 60 back | 42/42 | 229 / 261 | 10.2 min | 19 |
Rewinding at all is where the gain is. Which positions you try barely matters — and every probe costs real seconds on the phone, so the cheapest set stays. The next real lever is line length: the solver finds a winning line, not the shortest. Lines found for the same situation ranged from 120 to 177 moves; those 57 moves are over two minutes.
The first live rewind: 24 s to probe 5 candidates, back 10 moves in 11 s, then a win in 126 moves — the line from the opening was 177.
9. Running unattended: the long tail
The algorithms took less time than getting the bot to run for hours on a real phone without me. Almost everything here first happened during live play:
| What happened | Why | Fix |
|---|---|---|
| The bot crashed | On a USB drop only screenshots waited for the device; a tap or drag raised straight through | Every adb command waits for the device (up to 1 h) and resends — it never reached the phone, so a resend can't play a move twice |
| It nearly ended a deal | During rotation back to landscape the screen was unrecognisable; the old rule "unknown screen → new game" tapped New and the game asked "end your game?" | At start-up only a win / new-game / level-up / rewards screen may lead to a new game; anything else is waited out |
| Stuck after a win | "New Game" jumped straight into a Bonus Game (orange header, New becomes Retry); the bot thought it was on the old deal and tapped Retry | Tapping New Game on the win screen counts as starting a new game |
| Halted on the winning move | The last run leaves in a win animation; the board never matches, so the bot tried to undo it | A move that wins by the rules is accepted without a board check |
| It clicked an ad | A video ad plays after Play; a pixel on it looked like the game's gold confirm button, the bot tapped it and opened the advertiser's page, then took the white page for the board | After Play, tap nothing until a real opening board passes the full check; if none comes, press Android Back |
| The game froze | The XP animation after completing a run hung: clock stopped, input ignored, process alive, no ANR | After 6 moves in a row fail, send the game to the background and relaunch it (it restores its own save); halt only if that doesn't help |
| The OS restarted the app | Killed in the background; it came back on a loading screen, then the app's game list | Wait; on the game list tap Spider, and the game resumes the deal from its save |
| New pop-ups | Weekly Rewards, Level-Up, … | New dialog kinds, each probe checked against every old screenshot first |
Two rules came out of this:
- On a screen you don't recognise, wait — don't tap. The two most dangerous incidents — nearly ending a deal and clicking an ad — both came from acting on an unknown screen.
- Run every new pixel probe over all historical screenshots before it goes live. Hundreds of old screenshots are a free regression suite. The Weekly Rewards probe had 0 false hits on 493 of them, the game-list probe 0 on 875.
10. What didn't work
- An LLM choosing the move. Code enumerated the legal moves and their plain facts; an external LLM only picked "the best one". Refereed by the full-knowledge solver on 30 real positions, the share of choices that kept the position winnable was LLM 70%, Rust solver 90%, Python beam search 100%, random 77%. It also got stuck in full playthroughs. It's not used.
- Fixed-length backtracking: 28/42.
- Voting over several guesses (play the opening moves most solved guesses agree on): no measurable gain over one guess in simulation, so the simpler version stayed.
- Capping the line and re-searching every few moves: worse and slower.
- Denser rewind candidates: three times the probes, same time.
11. Data
Every deal, attempt and move went into SQLite: the position before the move, the move, who chose it, what it turned up, and how the attempt and the deal ended — meant as training data later.
As of September 26, 2026:
| Metric | Value |
|---|---|
| Deals recorded | 27 (19 with every card known) |
| Deals finished | 20, all won |
| Attempts | 153 |
| Moves | 11,052 (2,896 turned up a card, 228 completed a run) |
| Chosen by | promising line 4,155 · guess 2,716 · heuristic 2,352 · imported early logs 1,800 · exact 29 |
The 7 unfinished deals were interrupted or replaced (for example, the game switched deals during a disconnect), not given up.
12. What I'd tell myself at the start
- Look for the game mechanic you can exploit. "Undo All restores the same deal" mattered more than any search trick.
- Stagnation in the logs means the planning is wrong. Five attempts that learned nothing were the whole clue for guess mode.
- Reproduce offline before changing code, and measure before and after with the same ruler.
- Be honest about parameters. 5/10/20/40 was a guess. The experiment said it doesn't matter — but only the experiment could say that.
- Reliability is half the work. A 100% win rate that needs a human every few hours is not unattended.
Next: search a few more seconds for a shorter line once a win is found, compare rewind candidates on real deals rather than random ones, and train a move evaluator on the recorded games.
If you're automating something that only exposes a screen — a game, a legacy app, a device — and want to compare notes, get in touch.
让机器人赢下每一局 Grandmaster 蜘蛛纸牌:视觉、搜索、猜牌与撤销
📖 English version above — scroll up for the English translation.
写于 2026 年 9 月。文中数据取自机器人的对局数据库,截至 2026 年 9 月 26 日。
《Microsoft Solitaire Collection》里 Grandmaster 难度的蜘蛛纸牌是双花色、104 张牌,很多局只有极少数走法能赢。我想要一个机器人在我的 Android 手机上玩它:开了的局一局都不放弃,而且能不用我管、连续打上几个小时。
机器人不能用任何人没有的东西:截图、看牌、规划、拖一张牌、检查结果。不读内存,没有游戏接口。
截至 9 月 26 日,它的对局数据库记录了 20 局打完、20 局全胜,共 153 次尝试、11,052 步。
这篇文章讲真正重要的问题、我试过什么、什么真正管用——也包括没管用的。
1. 问题长什么样
双花色蜘蛛纸牌:54 张牌发成 10 列(其中 44 张面朝下),另外 50 张在发牌区,每次发 10 张。同花色从 K 到 A 连成一串就收走,收满 8 串获胜。开局时你只看得见 10 张牌。
限制:
- 只有 ADB:
screencap、input tap、input swipe,其余一切都从像素里推断。 - 手机横屏 2376×1080。牌列变长会被压缩,底部有一条会升起来盖住牌的工具栏,系统导航栏也挡着一部分屏幕。
- 一局都不放弃,并且无人值守,一局接一局。
2. 架构
图里都是模块名和命令名,各模块的职责见下表。
adb screencap ──► vision.py ──► brain.py ─────────────► executor.py ──► adb tap / swipe
▲ screen → exact solve drag, verify, │
│ Observation → guessed solve undo, dialogs │
│ → promising line │ │
│ → beam search ▼ │
│ │ ▲ knowledge.py │
│ ▼ │ (card memory) │
│ fastsolver (Rust) record.py │
│ full-game search (games.db) │
└────────────────────────────────────────────────────────────────────────────┘
| 模块 | 职责 |
|---|---|
vision.py | 截图 → 牌面观测:每列暗牌数和明牌、发牌区剩余次数、工具栏、对话框 |
rules.py / model.py | 纯规则引擎:合法走法、执行走法(自动翻牌、自动收牌)、Zobrist 哈希 |
knowledge.py | 牌的记忆:("H", 列, 深度) 与 ("D", 剩余发牌次数, 列) → 牌 |
fastsolver/ | Rust 全局求解器 |
brain.py | 精确解 → 猜牌解 → 最有希望的路线 → Python 束搜索;回退规划 |
executor.py | 拖动、校验、重试 / 撤销 / 重新同步、对话框、Undo All、开新局、智能回退 |
record.py | 每局、每次尝试、每一步写进 SQLite |
3. 视觉:找边缘,别信固定间距
牌列越长,游戏就把牌压得越紧,固定的纵向间距根本活不下来。活得下来的是:每张牌的上边缘都是一条横跨牌面、2–4 像素的浅灰线,面朝下的牌线下是蓝色,面朝上是白色。每一列都做一次自上而下的边缘扫描,数出暗牌和明牌。
被压住的牌只露出一条窄边,所以点数用左上角的字形做模板匹配,花色看右上角的小花色符号。模板按位置对齐,由标注过的截图生成。
单个字形可能认错,所以整盘牌要一起校验:总数必须刚好 104 张(包括收走的),任何一张牌出现的次数都不能超过它实际的张数。有一次,一张牌以 0.748 的分数(阈值 0.75)让机器人停了下来,而全盘校验其实全部通过。现在"全盘自洽、只有几张牌略低于阈值"会被接受并记录下来。
坑都出在"有东西挡住了牌面":
- 导航栏挡住了发牌区的底部。 改为数叠放的牌顶,而不是看牌堆底。
- 长牌列让检测器看不见工具栏。 原来判断"工具栏是否升起"只看 "New" 按钮下方一小块区域。一列长牌刚好垂在那块区域上,升起的工具栏就被判成没升起——机器人升不起工具栏,而且扫描牌列时一路扫进了工具栏,把它当成多出来的一张牌(21 张而不是 20 张),接着去撤销一步完全正确的走法。修复:沿整个工具栏宽度取多条色带的中位数。在 158 张历史截图上,有工具栏时 ≤ 64,没有时 ≥ 123。
- 特别长的牌列冷启动读不准。 每张牌只露出约 30 像素,字形被切掉一半。现在启动时会取上次运行保存的最后局面(以及它将要走那一步之后的局面),只要屏幕与之吻合就用它续打。
4. 执行:可靠地挪动一张牌
- 太快就成了"甩"。 快速的长距离滑动会被当成 fling,游戏会把牌放到别处。手指速度控制在约 2.5 像素/毫秒以下,最短 200 毫秒。
- 游戏会吞点击。 只清除提示光效的那一下点击什么也不做,应用空闲后的第一下点击也一样。每个按钮都走
tap_until:点、看、没反应再点。 - 工具栏的箭头在导航栏下面,永远点不到。改为点空白桌面来升起工具栏。
- 每一步都校验。 规则引擎预测下一个局面,截图与之比对。不一致时:重读 → 慢速重拖 → 撤销并以屏幕为准重新同步。
- 结果确定的走法成批发送。 源列几何位置仍然有效时连续发送,最后统一校验一次。
最后每一步(含校验)约 2.4 秒,其中光截图就要约 0.8 秒。
5. 一切的基础:牌的记忆
Undo All 会把这一局还原到开局,每一张牌都在原来的位置。
所以输掉一次尝试并不白输。它翻开的每一张暗牌、每一次发牌发出的每一张牌都会被记下来:
("H", 列, 深度)——某列暗牌堆某个深度上的牌;("D", 剩余发牌次数, 列)——某次发牌发到某列的牌。
下一次尝试一开始就知道得更多。尝试次数也会保存,机器人重启后接着数。开局变了就清空记忆;记忆和屏幕一旦矛盾,以屏幕为准。
这一个观察,把一个隐藏信息游戏变成了可以一局一局学习的问题。它比本文里任何搜索优化都值钱。
6. 求解器
fastsolver 是用 Rust 写的并行最优优先搜索,搜索整局游戏。已知的牌照常参与;未知的牌是"墙"——翻出来就停在那里,任何路线都不能依赖它是什么牌。
让它够快的几件事:
- 紧凑的搜索树。 节点只存"父节点 + 走法"(16 字节);完整局面每 8 层才存一次,其余靠重放恢复;局面是定长数组,每一步都不分配堆内存。每个节点约 34 字节(原来约 130),速度约快 3.5 倍。
- 分阶段组合。 先用 12 套权重各跑一个小预算(多数局面 1 秒内解出),再用 6 套、3 套权重跑更深的搜索来啃难局。
- 完整信息基准:32 局解出 31 局。 30 个随机局加 2 个真实难局;真实难局 C 第一版解不出,现在约 60 秒从开局解出,并经重放验证。
- "最有希望的路线"要短。 找不到赢法时,求解器返回评分最高的路线。给每一步加 40 分的代价后,它不再为了一点点分数提出 70 步的整理动作:模拟中每赢一局所需步数 245 → 208,胜率不变。
7. 猜牌模式:当"墙"本身成了问题
最大的改进来自读日志。
有一局打了 14 次尝试、约 89 分钟都没赢,第 9–13 次尝试一张新牌都没学到——每次都在重复一条几乎一样、约 65 步的路线。那时 94 张暗牌已知 88 张。根据还没见过的牌,剩下 6 张只可能是 {2♠, 6♠ ×3, 9♠, 9♥}:一共 120 种排列。
一次离线测试就说明了问题:
| 方法 | 结果 |
|---|---|
| 未知牌当"墙" | 60 秒无解 |
| 随机填满这 6 张再求解(6 种填法) | 2 种约 12 秒解出完整赢法 |
不是运气不好,是规划方式有问题。当成墙时,所有需要穿过这 6 张牌的路线都被禁止——而剩下的路线全是死路。所以机器人每次尝试都走进同一条死胡同。
修复办法是确定化(determinization)。当未知牌 ≤ 16 张、且"墙"模式找不到赢法时:
- 用还没见过的牌随机填满所有未知位置(已收走牌串的花色由剩余计数推断),求解整局。
- 跟着第一个能赢的猜测走。每翻出一张真正的未知牌就核对:猜对了继续走;猜错了就带着多知道的这一张重新猜。
- 一轮 6 次猜测都失败后,在学到新牌之前不再重试,否则每一步都可能花一分钟。
我还把失败的精确搜索按"局面 + 已知信息"缓存起来。以前每次 Undo All,开局那 90 秒的搜索都在没有任何新信息的情况下白做一遍。
效果:
- 离线重放那局卡住的牌,把 6 张未知牌填成几种可能的"真相":旧逻辑 5 种填法全部 5 次尝试不赢、始终停在 88 张;新逻辑跑完的 2 种填法分别在第 1 次、第 2 次尝试获胜。
- 在手机上,另一局在旧逻辑下第 6–8 次尝试一直停在 80 张已知牌。换上猜牌模式后,第 9 次尝试学到 3 张新牌,第 10 次尝试获胜。
8. 撤销:退回到代价最低的好局面
游戏还有普通的 Undo(撤销)。早先有个实验用的是最朴素的方式——死局时撤销固定 8 步、换一条路线——结果输给了直接重开:28/42 对 41/42。局部后退永远不会重新规划整局。
有了猜牌模式之后,我换了一种用法。死局时,用当前所有已知信息(包括猜牌),在本次尝试的几个早先局面(往回 5、10、20、40 步)以及开局里找赢法,然后去代价最低的那一个:
代价 = 撤销次数 × 1.1 秒 + 剩余步数 × 2.4 秒 (Undo All 记 6 秒)
Undo All 只是"退回开局"这一特例,所以这不可能比只用重开更差。
先确认 Undo 是精确的。 在手机上连续撤销 20 次,每一步都与数据库记录的"走这一步之前的局面"比对:普通走法、两次发牌、一次翻牌(翻开的牌重新盖了回去)。20 次全部一致。
再让它变快。 一次撤销一步——升工具栏、点撤销、截图确认——约 3.6 秒。改成成批连点,然后读一次盘面、在本次尝试的局面列表里定位它(顺便纠正被吞掉或多出来的点击)之后,7 次撤销用了 7.4 秒:约 1.06 秒一次,比走一步还便宜。这份局面列表与游戏的撤销栈一一对应;一旦发生计划外的情况就作废,退回 Undo All。
模拟,42 个随机局,两边都开猜牌模式:
| 方法 | 胜局 | 每局操作数(中位 / 平均) |
|---|---|---|
| 只用 Undo All | 42/42 | 255 / 319 |
| 智能回退 | 42/42 | 210 / 251 |
5 / 10 / 20 / 40 选得对吗? 老实说,是凭感觉定的。于是比较了三种候选集合:
| 候选局面 | 胜局 | 每局操作数(中位 / 平均) | 估计手机耗时 | 每局探测次数 |
|---|---|---|---|---|
| fixed:5/10/20/40 + 开局 | 42/42 | 202 / 260 | 10.3 分钟 | 6 |
| events:最近 6 次发牌/翻牌之前 + 开局 | 42/42 | 230 / 272 | 10.8 分钟 | 11 |
| dense:每 5 步一个,最多退 60 步 | 42/42 | 229 / 261 | 10.2 分钟 | 19 |
收益来自"会回退"本身,试哪些局面几乎无所谓——而每次探测在手机上都是真金白银的秒数,所以保留最省的那组。下一个真正的杠杆是路线长度:求解器找到的是"一条"赢法,不是最短的。同一个局面找到的赢法从 120 步到 177 步不等,这 57 步就是两分多钟。
第一次实机回退:24 秒探测 5 个候选,11 秒退回 10 步,然后 126 步获胜——从开局重来的路线是 177 步。
9. 无人值守:长尾问题
算法花的时间,比不上让机器人在真手机上不用我管、连跑几个小时。下面几乎每一条,都是实机运行时第一次冒出来的:
| 发生了什么 | 原因 | 修复 |
|---|---|---|
| 机器人崩溃 | USB 掉线时只有截图会等设备,点按和拖动直接抛异常 | 所有 adb 命令都等设备回来(最多 1 小时)再重发——命令根本没到手机,重发不会重复走牌 |
| 差点结束一局 | 转回横屏的过程中画面认不出来,旧规则"未知画面 → 开新局"点了 New,游戏问*"是否结束本局?"* | 启动时只有胜利 / 新局 / 升级 / 奖励画面才能开新局,其他一律等待 |
| 赢了之后卡住 | "New Game" 直接进入了 Bonus Game(橙色顶栏,New 变成 Retry);机器人以为还在旧局,点了 Retry | 在胜利画面点了 New Game 就算已经开了新局 |
| 在制胜一步上停机 | 最后一串收走后是胜利动画,盘面永远对不上,机器人想撤销它 | 按规则会获胜的那一步直接判赢,不校验盘面 |
| 点开了广告 | 点 Play 后插播视频广告;广告上某个像素像游戏的金色确认按钮,机器人点了,打开了广告主网页,又把白色网页当成了牌桌 | 点 Play 之后什么都不点,直到真正的开局盘面通过全盘校验;等不到就按 Android 返回键 |
| 游戏卡死 | 收串后的 XP 动画卡住:计时停止、输入无效、进程还在、系统没报无响应 | 连续 6 步不生效时,把游戏切到后台再拉起(它会从自己的存档恢复),还不行才停机 |
| 应用被系统重启 | 在后台时被杀掉,回来是加载画面,然后是应用首页 | 等待;在首页点 Spider,游戏会从存档恢复这一局 |
| 新弹窗 | 每周奖励、升级…… | 新增对话框类型,每个探针上线前先在全部历史截图上验证 |
从中得出两条规则:
- 遇到不认识的画面,只等待,不点击。 最危险的两次事故——差点结束一局、点开广告——都来自"对未知画面采取了行动"。
- 每个新的像素探针,上线前先在全部历史截图上跑一遍。 几百张历史截图就是免费的回归测试集。每周奖励弹窗的探针在其中 493 张上 0 误报,应用首页的探针在 875 张上 0 误报。
10. 没成功的尝试
- 让大语言模型选走法。 代码枚举合法走法和它们的客观事实,外部 LLM 只负责挑"最好的一步"。在 30 个真实局面上由完整信息求解器裁判,"选完之后仍然能赢"的比例:LLM 70%、Rust 求解器 90%、Python 束搜索 100%、随机 77%。完整对局里它也会卡住。没有采用。
- 固定步数的局部回退:28/42。
- 多次猜测投票(按多数猜测解一致的开头几步走):模拟中与单次猜测相比没有可测的收益,保留了更简单的版本。
- 截断路线、每几步重新搜索:更差也更慢。
- 更密的回退候选:探测次数多三倍,时间一样。
11. 数据
每一局、每次尝试、每一步都写进了 SQLite:走之前的局面、走法、由谁决定、翻出了什么、这次尝试和这一局的结局——打算以后当训练数据用。
截至 2026 年 9 月 26 日:
| 指标 | 数值 |
|---|---|
| 记录的牌局 | 27(其中 19 局所有牌都已知) |
| 打完的牌局 | 20,全部获胜 |
| 尝试次数 | 153 |
| 走法 | 11,052(翻出暗牌 2,896 次,收串 228 次) |
| 由谁决定 | 最有希望的路线 4,155 · 猜牌 2,716 · 启发式 2,352 · 早期日志导入 1,800 · 精确解 29 |
没打完的 7 局是被中断或被替换的(比如断线期间游戏换了一局),不是放弃。
12. 如果能回到开始,我会告诉自己
- 先找能利用的游戏机制。 "Undo All 还原同一局"比任何搜索技巧都重要。
- 日志里的停滞,说明规划方式错了。 五次尝试一张新牌都没学到,这就是猜牌模式的全部线索。
- 改代码之前先离线复现,改之前和改之后用同一把尺子量。
- 对参数要诚实。 5/10/20/40 是猜的。实验说它无所谓——但只有实验能说出这句话。
- 可靠性是一半的工作量。 一个 100% 胜率、却每隔几小时就要人去点一下的机器人,不叫无人值守。
接下来:找到赢法后多花几秒找一条更短的;在真实牌局而不是随机局上比较回退候选;用记录下来的对局训练一个走法评估模型。
如果你也在自动化一个只暴露屏幕的东西——游戏、老旧系统、某台设备——想交流一下,欢迎聊聊。