I Replaced My Paper Budget With a WhatsApp Bot. Here's What Broke.

TLDR: I built a family budget WhatsApp bot on Hermes Agent. It reads my credit card SMSs, tracks every category against the budget, and answers questions in Hebrew or English. The rule that made it trustworthy: the AI understands the question, but deterministic Python does all the math. The code was the easy part. The plumbing, and an agent with more access than I remembered giving it, were not.

Once a week, I sit down with a piece of paper. I go over every expense from my credit card, add them up by category, and work out how much budget I have left for the rest of the month.

It works. It has worked for years. It’s also tedious, it’s always a week out of date, and the only person who actually knows the numbers is me.

Here’s the thing though. Every single one of those charges already arrives on my phone as an SMS from the credit card company. The data was there the whole time. I just wasn’t doing anything with it.

I asked Claude Code to help me plan this:

create a Bot - that both I and my wife can interact with over Whatsapp

It should track the budget from the credit card messages, send a weekly status, alert me if I go over budget, and tell me what’s left whenever I ask. I already had most of the pieces: a Hermes Agent running on an old Mac mini at home, a spare WhatsApp number, and IFTTT on my Android phone.

One day later, the paper is (almost) retired. Here’s how it went. What worked, what didn’t, and why I made the choices I did.

The one rule: numbers never come from the AI

Before a single line of code, I started with one rule that shaped everything else.

The LLM never does the math.

LLMs are great at understanding when I ask “how much is left in the clothing budget” or “125 Delek” (when I say Delek, I mean fuel). They are not great at adding up 60 charges the same way every time. And this is my family’s money. “Mostly right” is not good enough.

An LLM is a great interpreter and a terrible accountant.

All the parsing, the math and the alerts are plain, deterministic Python on top of SQLite. The AI only does two things:

  1. It turns chat messages, in Hebrew or English, into tool calls.
  2. It guesses the category of a merchant it has never seen before.

And when it answers, it relays whatever the tool returned, word for word. It isn’t allowed to “improve” the numbers.

How does the WhatsApp budget bot work?

Credit card SMSs come in through IFTTT and a webhook, deterministic Python records them in SQLite, and WhatsApp is where my family talks to it. Here’s the whole flow:

When a charge comes in:

flowchart TD
    SMS["💳 Credit card SMS"] --> IFTTT["Android phone<br/>IFTTT applet"]
    IFTTT -->|"HTTPS POST + shared secret"| Router["Home router<br/>one forwarded port"]
    Router --> Caddy["Caddy<br/>TLS, one path only"]
    Caddy --> Webhook["Hermes webhook<br/>localhost only"]
    Webhook --> Python["Python<br/>parse, dedupe, categorize"]
    Python --> DB[("SQLite")]
    Python --> WA["WhatsApp"]
    WA --> Recorded["💳 Recorded, with what's left"]
    WA --> Ask["❓ New merchant, which category?"]
    WA --> Over["🔴 Over budget"]

When someone asks a question, or it’s time for a report:

flowchart LR
    Question["💬 WhatsApp question"] --> Agent["Hermes agent<br/>budget profile"]
    Agent --> Tools["Budget MCP tools"]
    Tools --> Python["Same Python code"]
    Python --> DB[("SQLite")]
    Python --> Reply["Reply, relayed word for word"]
    Cron["⏰ Hermes cron"] -->|"weekly report, cycle summary, nightly backup"| Python

A charge comes in, gets recorded, and WhatsApp shows what’s left in that category. A new merchant gets a ❓ with the AI’s best guess, and nothing counts against a category until I confirm it. I answer “food”, and the bot never asks about it again.

Why I made the choices I did

Every component had to earn its place. Here’s the reasoning behind each one.

A separate profile and a separate number

My main Hermes assistant is a power user. It has a terminal, it can run code, it has a browser. Do I want anyone’s WhatsApp messages going to an agent that can run shell commands on my Mac? Hell no!

The budget bot got its own Hermes profile, its own gateway and its own WhatsApp number. Its WhatsApp side has exactly nine budget tools, exposed over MCP. No shell. No files. No browser. No memory. If you’ve read my post on tool misuse in the OWASP Agentic Top 10, you know why.

The smallest blast radius is a tool that doesn’t exist.

IFTTT, but only for charge messages

IFTTT forwards the SMS to a webhook. The trigger matches a phrase that only appears in charge messages (“transaction approved”), and not the name of the card company. Why? Because the card company also sends me login and verification codes. 2FA codes never leave the phone.

IFTTT can’t sign requests with HMAC, so the webhook uses a plain shared secret in a header. Not perfect. But it only travels over TLS, and it’s one layer out of several. Good enough for me.

Port forwarding, not a tunnel

The first plan used Tailscale Funnel. I asked a simple question: “Why does this need tailscale?”

It didn’t. I can forward a port on my home router, I’m not behind carrier-grade NAT, and I already had a domain on Cloudflare. ddclient keeps the DNS record current. One less dependency, one less account.

Caddy in front of Hermes

I’d noticed that Hermes has its own webhooks page, so why put Caddy in front of it? I asked exactly that, and said “don’t make changes just explain your thinking”.

The answer was a good one. That page is the Hermes dashboard, which should never be anywhere near the internet. The webhook listener itself speaks plain HTTP and has no path filtering. Caddy adds a real TLS certificate (through a Cloudflare DNS challenge, so ports 80 and 443 stay closed), exposes a single path, caps the body at 64KB, and keeps the secret header out of the logs. Hermes only listens on localhost.

The budget follows the credit card, not the calendar

My month starts on the day the billing cycle starts, not on the 1st. Food is the only weekly category. The monthly amount is split over the Sundays in the cycle (4 or 5 weeks), and a cheap week makes the next one bigger. Everything else is monthly.

Alerts? I was very clear: “Do I get more alerts from the bot (spoiler: I do not want to)”. One alert per category, per period, when it goes over. That’s it.

What didn’t work

This is the part I find most interesting. The code was the easy part. Everything around it was not.

“I DM’ed status, nothing happened”

The first real test. I sent the bot “status”. Silence.

The WhatsApp bridge reads its allowlist only from an environment variable, not from the config file where I’d put it. Everyone got dropped silently, including me. No error. No log line. Just… nothing.

The bot forgot it had tools

Fixed the allowlist, sent “status” again. This time I got a cheerful, generic assistant who offered to load a budget skill, and to get to know me better by building a profile of me.

Hermes has a tool search feature that hides MCP tools behind a search step. The model never searched, and answered without them. I turned tool search off, wrote a dedicated persona, and turned off memory and onboarding. The bot should not be building profiles of my family. Ever.

“ignore” → “I only handle the family budget”

A charge came in that isn’t part of my budget, and I replied “ignore”. The bot told me it only handles the family budget. Thanks.

The ❓ question was sent by the webhook, and Hermes doesn’t copy webhook deliveries into the chat session. When I replied “ignore”, the agent had no idea what I was replying to. The fix: a bare reply now answers the newest pending charge. Simple, once you know it.

The AI moved faster than I did

Two moments stood out, and neither of them was a bug in the code. Both came from Claude Code, the agent building the bot, not from the bot itself.

Early on, Claude Code SSH’d into my Mac mini to look around. Read-only, but it didn’t ask first. I had to stop and ask, “Where did that info come from?”

Later, it couldn’t test the public URL from inside my network. Its workaround? It used an IFTTT MCP server I had connected months ago (yes, I wrote a whole post about it, and I still completely forgot) to hit the URL from IFTTT’s cloud. My reaction: “where did you get access to IFTTT from? I don’t remember giving you credentials or my account?”

Nothing bad happened. But that’s exactly the point. I’ve written about identity and privilege abuse in agents, and here it was, on my own desk.

Your agent has every credential you ever connected, and it will use whatever gets the job done.

I set clearer boundaries (read-only until I say otherwise). Later, once I trusted the process, I told it “can you run this all for me?”. Trust gets earned one step at a time.

The free models are slow

Replying “ignore” took 85 seconds on the free-tier model. Right answer, painfully slow. Charges never touch the chat model, so those are instant. Chatting with the bot? Patience required. That’s the price of free.

What worked

Honestly? More than I expected, and faster.

  • The core was fast. From an approved plan to 63 passing tests took about 45 minutes, with a full local simulation before anything touched the Mac. By the end of the day it was 100 tests.
  • The first real charge just worked. It arrived with slightly different wording from the sample SMSs, and it still parsed.
  • My old spreadsheet turned into rules. I had January through August credit card charges in a spreadsheet: about 1,400 charges, which became about 420 merchant rules. I made the calls (“Medical charges, should only be pharmacies”, “BIT/Paybox can be skipped”), and for the 73 merchants that could go either way, “let me decide”. Fuel stations get an amount rule. A big charge is fuel, a small one is the shop, which makes it food.
  • “Ask me” rules. Some merchants sell everything. For 19 of them I said “ask me if there is a charge”. Those never get a guess and never get learned, unless I say “always”.
  • Safety nets. A backup before every change on the Mac, rollback steps written down, and a redaction pass before the first commit. That pass caught the real last four digits of a card, sitting in a code comment. Imagine that on GitHub.

All in all, one day. About 1,800 lines of Python, 1,000 lines of tests and around 9,000 words of documentation, so future me can rebuild this on a new Mac without reverse-engineering what present me did. Code you didn’t write is still code you own.

What’s still open

Refunds and installments haven’t been tested against real SMSs. And I am running it side by side with the paper for two weeks before the pen retires for good.

What I learned

  1. Keep the AI away from the math. Let it understand people. Let code count the money.
  2. The AI was never the hard part. Allowlists that only live in environment variables, sessions that can’t see webhooks, tools hidden behind a search step… the plumbing ate most of the day.
  3. Ask “why” about every component. Two questions (“why Tailscale?”, “why Caddy?”) removed one dependency and justified another.
  4. Your agent has all your keys. It will use whatever works, including the MCP server you forgot you connected. Set the boundaries before it starts, not after.

The pen isn’t retired yet. But it’s already packing its bags.

Have you built something like this for your household, or are you still on paper? I would be very interested to hear your thoughts or comments, so please feel free to ping me on Twitter or LinkedIn.