Capabilities / Software, Apps & AI / Custom AI

AI with human hands on the wheel.

Assess. Build. Govern.

Custom assistants, automation, and model integration built for a specific job, with the guardrails written before the feature. We will also tell you when you do not need AI at all.

HAUS XXIV builds custom AI assistants, retrieval systems, automation, and governance. Every engagement begins with a paid assessment producing a written go or no-go recommendation and a cost model per user and per query. A meaningful share of assessments end with a recommendation not to build.

Most AI projects are a demo that never became a product.

It looks incredible in the meeting. It works on the four examples someone picked. Then real people use it, it confidently invents an answer, someone senior loses trust, and the whole thing gets quietly shelved eight months and a lot of money later.

01 / WHAT WE BUILD

Four capabilities. One honest answer.

Every engagement starts with whether you need this. We map the workflow you want to improve, estimate what AI would actually change, and model the running cost per user and per query before anybody commits.

WHAT SHIPS

  • Workflow mapping of the process in question
  • Go or no-go recommendation with the reasoning written out
  • Cost model per user and per query at realistic volume
  • Risk assessment covering accuracy, privacy, and compliance
  • Data readiness review of what you actually have
  • Phased build plan, if the answer is yes

Systems that answer questions over your own documents, data, and history. Internal knowledge assistants, customer-facing support, and search that understands a question rather than matching keywords.

WHAT SHIPS

  • Retrieval architecture over your documents and data
  • Source citation on every generated answer
  • Evaluation harness with a scored test set
  • Refusal behavior for questions outside its knowledge
  • Conversation logging with full audit trail
  • Admin interface for reviewing and correcting answers

Document processing, classification, drafting, extraction, and the repetitive judgment work that eats a team's week. The tasks where a human is currently doing something a machine can do at ninety-five percent accuracy, and where the remaining five percent is exactly why the human stays in the loop.

WHAT SHIPS

  • Automated pipeline integrated into your existing systems
  • Human review checkpoints at the decisions that carry weight
  • Confidence scoring with automatic escalation thresholds
  • Accuracy measurement against a labeled test set
  • Full processing logs for audit and dispute
  • Fallback behavior for when the model is unavailable

The part nobody demos and everybody needs. What the system is allowed to do, what data it can reach, who can see the logs, and how you prove any of it to a regulator, a board, or a client who asks a hard question.

WHAT SHIPS

  • Written AI use policy for the system
  • Data handling, retention, and residency documentation
  • Access control and permission model
  • Prompt injection and misuse testing
  • Audit log design that satisfies a real review
  • Disclosure language for anyone interacting with the system

02 / NO AMBIGUITY

What actually lands in your hands.

Assessment

  1. Workflow map of the process under review
  2. Written go or no-go recommendation
  3. Cost model per user and per query at real volume
  4. Risk register covering accuracy, privacy, and compliance
  5. Phased build plan with a cost range per phase

Build

  1. Source code in your repository from the first commit
  2. Evaluation harness with a scored, labeled test set
  3. Model integration built to be swapped, not locked in
  4. Human-in-the-loop checkpoints where decisions carry weight
  5. Audit logging on every generated output
  6. Admin tooling for review and correction

Governance & handover

  1. Written AI use policy and disclosure language
  2. Data handling and retention documentation
  3. Prompt injection and misuse test results
  4. Running cost dashboard
  5. Recorded admin training and a written runbook

03 / HOW AI GETS BUILT

Five phases. Assessment first, and it can end there.

A meaningful share of AI assessments end with a recommendation not to build. That is a successful engagement. You keep the document and the money you did not spend.

Two to three weeks, priced separately. We map the workflow, review your data, model the running cost, and give you an honest recommendation. If the answer is no, this is where the project ends and you are better off for it.

Before building the feature we build the way to score it. A labeled test set drawn from your real cases, with a target accuracy agreed in writing. Without this, "it seems pretty good" becomes the entire quality process, which is how AI projects die.

Retrieval, integration, interface, and the human checkpoints. Two-week cycles with a working demo each time, scored against the harness so progress is a number rather than a feeling.

Guardrails, access control, audit logging, and adversarial testing. We try to make it behave badly on purpose, including prompt injection and the questions it should refuse. Then we write the policy and the disclosure language.

Staged rollout to a small group first, because real users ask questions nobody on the project team imagined. We monitor accuracy and cost in production, then train your team to run it and read the logs.

01 / 05

  1. 01

    Assess

    Two to three weeks, priced separately. We map the workflow, review your data, model the running cost, and give you an honest recommendation. If the answer is no, this is where the project ends and you are better off for it.

    Ends with: Written recommendation, cost model, and risk register.

  2. 02

    Measure first

    Before building the feature we build the way to score it. A labeled test set drawn from your real cases, with a target accuracy agreed in writing. Without this, "it seems pretty good" becomes the entire quality process, which is how AI projects die.

    Ends with: Evaluation harness and an agreed accuracy threshold.

  3. 03

    Build

    Retrieval, integration, interface, and the human checkpoints. Two-week cycles with a working demo each time, scored against the harness so progress is a number rather than a feeling.

    Ends with: Working system in staging with published accuracy scores.

  4. 04

    Govern

    Guardrails, access control, audit logging, and adversarial testing. We try to make it behave badly on purpose, including prompt injection and the questions it should refuse. Then we write the policy and the disclosure language.

    Ends with: Governance documentation and misuse test results.

  5. 05

    Launch and watch

    Staged rollout to a small group first, because real users ask questions nobody on the project team imagined. We monitor accuracy and cost in production, then train your team to run it and read the logs.

    Ends with: Production deployment, cost dashboard, recorded training, and a runbook.

04 / WHY THIS HAUS

Three reasons, and the code behind them.

PRINCIPLE 12

AI is a tool. Not a shortcut.

We build AI that accelerates people, not AI that replaces them quietly.

Every system keeps a human at the decisions that carry weight, and logs enough that you can audit what it did and why. The more powerful the tool, the higher the bar for the human holding it. We don't replace evolving. We replace laziness.

PRINCIPLE 10

No ego. All output.

We will tell you not to build it.

A real share of the AI work we get asked for should not exist. Sometimes the answer is a report, a rule, or a tool you already pay for. We say so in the assessment, on the record, even though the build is worth far more to us than the honest answer.

PRINCIPLE 07

Conviction over convenience

The guardrails get written before the feature.

Governance is not a phase we add if the client asks. Access control, audit logging, refusal behavior, and disclosure language are scoped into every build. For regulated, education, and nonprofit work, that conversation happens first.

05 / THE STACK

  • Claude
  • OpenAI
  • Open-weight models
  • LangChain
  • pgvector
  • Pinecone
  • Python
  • TypeScript
  • PostgreSQL
  • AWS Bedrock
  • Vercel AI SDK
  • Langfuse

The model layer moves every few months, so we build the integration to be swapped rather than welded on. Your system should survive a vendor changing their pricing, their terms, or their roadmap without a rewrite.

See the full Software, Apps & AI discipline

06 / STRAIGHT ANSWERS

AI questions, answered plainly.

An assessment runs low five figures and is contracted separately. Build cost depends entirely on what came out of the assessment, and most first phases land in the mid five to low six figure range. Running costs are modeled per user and per query during the assessment, because inference cost is the line that surprises people six months in, not the build.

Not on any architecture we build. We use enterprise API tiers with training explicitly disabled, or open-weight models running in your own infrastructure when the data cannot leave. Data handling and residency are documented in writing during the assessment, before you commit to anything.

It will, occasionally. Any vendor telling you otherwise is selling. The question is what happens next, and that is what governance is for. Citations so answers can be checked, confidence thresholds that escalate to a human, refusal behavior for questions outside its knowledge, and logs that let you reconstruct exactly what happened.

Almost certainly not. Training a custom model is expensive, slow, and rarely better than a good retrieval system over your own data using an existing model. We will tell you if you are the exception. Most companies asking this question need better retrieval, not a bigger model.

Increasingly yes, and this is where AI is genuinely useful rather than fashionable. A small team automating document processing or answering repeat questions gets back real hours every week. We scope those engagements at a fraction of enterprise builds, because community is non-negotiable in the Code of 24 and that has to survive contact with a price list.

Not sure whether you need AI or just better software?

Tell us the workflow that is eating your team's week. We will tell you honestly which one it is.