Implementing AI / After the hype

Yes, we are also sick of hearing about fucking AI.

The buzzword has been beaten to death by people who automated one email and called themselves futurists. Underneath the noise is genuinely useful infrastructure. Let us show you the sauce.

Skip the keynote

Submitted securely through Formspree.

First principles

AI is not a strategy.
It is a component.

A model can read, reason, classify, draft, compare, extract and sometimes use tools. That does not mean it understands your business, should touch every process or gets permission to freestyle inside the bank account.

Useful implementation starts with a real job, gives the model the right context, picks the right level of intelligence, limits what it can do and measures whether the workflow is better than before.

01 / The actual sauce

It is not one
magic prompt.

The chat box is the visible bit. The quality comes from the system around it.

01

Pick a real business job.

“Use AI” is not a brief. “Turn every sales call into a clean summary, follow-up draft and CRM update within two minutes” is.

02

Engineer the context.

The model needs the right slice of your offers, customers, tone, SOPs, current records, examples and constraints at the moment it works. More context is not always better. Relevant context is.

03

Write prompts like operating instructions.

State the outcome, inputs, constraints, process, examples and definition of done. Prompting is not collecting magic phrases. It is learning to specify work clearly and test the output.

04

Route work to the right model.

Do not pay frontier-model prices to label a support ticket. Do not ask the cheapest tiny model to make a high-stakes commercial judgement. Different work deserves different brains.

05

Give agents narrow tools.

Reading a task list is different from changing it. Drafting an action is different from executing it. Expose only the tools required, keep an audit trail and put approval in front of risky writes.

06

Evaluate reality.

Track accuracy, time saved, correction rate, latency, token cost and commercial outcome. If humans spend longer fixing the output than doing the work, the demo failed.

02 / Things we have actually built

Not a concept deck.
Working systems.

01

Neo HQ + Hermes

An agent connected to the business, not just another chat tab.

We built a private operations HQ, then connected a Hermes agent running on a VPS through a controlled tool bridge. It could discover authorised tools for tasks, clients, pipeline, revenue context and ad-account status without opening the entire system to the internet.

  1. Business data stays in the private HQ
  2. Hermes receives only the context and tools needed
  3. The agent reads the current state and prepares useful work
  4. Tool activity is visible instead of hidden behind “trust me”
02

Operator HQ approvals

AI proposes. A human decides. The system records both.

For an operating dashboard, we created an agent bridge that returned normal advice plus structured proposed actions. The first executable write surface was deliberately narrow: create a task only after explicit approval.

  1. Operator asks for help in plain language
  2. Agent returns an answer and proposed actions
  3. Operator reviews, approves or rejects each action
  4. Approved work is executed and kept in durable state

03 / What this could look like for you

A real workflow.
Step by step.

Say you run a service business and leads arrive through forms, phone calls and email. Here is a useful implementation. Not a chatbot waving from the bottom-right corner.

  1. Trigger

    A lead arrives.

    The system validates the contact details and pulls the source, page, service and campaign context.

  2. Understand

    A fast model classifies it.

    Job type, location, urgency, likely value and missing information are returned in a strict structure.

  3. Act

    The workflow routes it.

    High-intent work goes to the right person. A useful reply is drafted in your tone. The CRM record and follow-up task are prepared.

  4. Control

    Risk decides the approval.

    A simple acknowledgement can send automatically. A quote, promise or unusual case waits for a human.

  5. Learn

    The outcome closes the loop.

    Booked, lost, junk or no-answer data returns to the system so prompts, routing and acquisition improve.

04 / Different models, different work

Stop using a Ferrari
to file receipts.

Frontier reasoning

For hard thinking.

Strategy, complex planning, messy document analysis, coding and decisions where a better answer is worth the extra latency and cost.

Fast, lower-cost models

For volume.

Classification, extraction, formatting, routing, first drafts and repetitive work with clear rules and cheap verification.

Multimodal & specialist models

For the right input.

Calls, screenshots, plans, invoices, photos, video or realtime voice. The best text model is not automatically the best model for every medium.

Open & Chinese models

For options and cost pressure.

Qwen, DeepSeek, Kimi and open-weight models can be strong, cost-effective parts of a router. They are not automatically the right choice. Data handling, hosting, reliability, provider risk and output quality still need testing.

Model loyalty is not a business strategy. Route by quality, speed, privacy and cost—then keep an exit door.

Route

Use cheaper models for clear, high-volume work.

Retrieve

Send the relevant records, not the entire company drive.

Cache

Reuse stable instructions and repeated context where providers support it.

Batch

Run non-urgent work efficiently instead of demanding instant answers.

05 / Context engineering

The model is smart.
It still does not know you.

A good prompt with bad context produces polished guessing. The goal is to assemble the smallest trustworthy packet of information needed for this job, right now.

Business truth

Offers, margins, customers, territory, policies, tone and what “good” actually means.

+
Live state

The current lead, job, client, account, task, document or conversation.

+
Instructions

Outcome, constraints, examples, tools, edge cases and a definition of done.

+
Control

Permissions, approvals, structured output, logging, evaluation and a fallback path.

06 / Prompting without wizard cosplay

Clear thinking in.
Useful work out.

01

Give it a job.

One outcome. One owner. One definition of done.

02

Show the standard.

Real examples beat ten paragraphs of adjectives.

03

Name the constraints.

What it may use, must avoid and should do when uncertain.

04

Demand structure.

Use fields, schemas or checklists when another system consumes the result.

05

Test ugly cases.

Missing data, contradictions, nonsense input and customers doing weird customer shit.

06

Version it.

Prompts are operating logic. Change them deliberately and compare outcomes.

07 / Find a useful first move

What work would you
never miss?

Tell us what repeats, what requires judgement, where the data lives and what going wrong would cost. We will map a sensible first implementation—even if the answer is “do not use AI here.”

hello@goodproblems.com.au
Start the free diagnosis