Polaris
One local dashboard where AI agents hand work back to me
A desktop dashboard that turns what my AI agents need from me into cards I can act on, such as questions to answer, commands to copy and run, and revision drills, served by a server only my own Mac can reach.
Problem
My agents run unattended, but some things only I can do, such as answering a question or running a command their safety rules forbid. Those requests used to sit in reports, easy to miss.
Approach
One window showing what the terminal pane shows, as a second surface on the same facts. Every action runs the pane's own code, so a click and a keypress do the same thing.
Architecture
A Python server of 1,873 lines loads the pane as a module and computes no state of its own. It binds to 127.0.0.1 only and mints a token at start. It refuses any Host header that is not local, and every write must carry that token. A build script wraps it as a Mac app with a native window, installing nothing.
What I built
Questions. One queue of everything waiting on me, oldest blocker first. Any agent run can ask by writing a question file, and an answer can queue follow-up work, so a yes launches it and a no launches nothing. Finished work pops up as a pin when Polaris next opens.
Run these. When a session hits an action it may not take, it hands me the exact command instead. A card shows each with Copy, Done and Drop. A command must be one line from an absolute directory, or it is refused. A checker marks each safe, unverified, done or stale, and destructive steps such as a forced push count as stale. It runs every 10 minutes, or sooner when code a command names changes, and a stale row has no Copy button.
Revision drills. Sessions with formatted questions, one model call per turn returning JSON for the page. Answers are saved through the command-line tool's own code path, so both leave identical records.
Terminals and the phone page. Real shells in tabs, using a bundled xterm.js and one long-poll for every tab. A static companion page adds Done buttons and answer boxes on the phone. One applier does all its writing, nothing is fuzzy matched, and a tapped row says "sent", never "done".
Results
Parity with the pane is checked by a self-test that drives every action over a synthetic tree, with every writer stubbed. The server's self-test confirms a send without the token, or from another Host, is refused.
What didn't work
The first Run these commands were multi-line blocks, and a plain shell ran the comments as commands. One had gone stale overnight and would have deleted tracked files. The Terminal tab first showed only a drawing of a terminal. Replaying stored output also replayed tmux's queries, typing stray characters into the shell, so replayed bytes are now never answered.
Stack
- Python
- JavaScript
- WebGL2
- xterm.js
- pywebview
- tmux