Why we build on the same AI we deploy for clients
There is a well-known phrase in the construction industry: 'build what you sell'. A roofing contractor who leaks is not in business for long. The same principle applies to AI consultancy. If we recommend self-hosted, privacy-first AI infrastructure to our clients, we ought to run our own business on exactly the same stack.
At Hynt Digital, we do. Every internal workflow — client research, email triage, document drafting, meeting transcription, and project management — runs on the same locally-hosted AI models we deploy for our clients.
This is not a marketing gimmick. It is a practical discipline that makes us better at what we do.
When you run your business on the same infrastructure you recommend, you discover the rough edges first. You learn which models work well for document summarisation and which hallucinate under pressure. You know exactly how much hardware is needed to run a 7B model versus a 13B model in real-world use. You experience the same constraints your clients will face — and you solve them before they become problems.
Our internal stack uses Ollama to serve open-source models — currently Mistral and Llama 3.1 variants — on dedicated hardware in our own office. A local vector database indexes our internal knowledge base, client precedents, and project history. Our AI assistant drafts emails, suggests research directions, and helps us maintain consistency across client work.
The open-source models we deploy are not experimental. Mistral 7B scores 62.5% on the MMLU benchmark — competitive with GPT-3.5 (70%) and approaching GPT-4 class performance on many business tasks. Llama 3.1 8B scores 73% on MMLU, demonstrating that a model running locally on a mid-range server can match the capability needed for most SME workflows. These are not toy models — they are production-grade tools that happen to run on your own hardware.
Crucially, because everything runs locally, there is no data leakage. Client information discussed in internal emails never touches a third-party API. Meeting transcriptions stay on our own hardware. This is the same privacy guarantee we offer every client.
When a client asks whether local infrastructure really works in practice, we do not hand them a brochure. We show them our own setup. We let them talk to our internal assistant. We explain that the system we are building for them is the same system we use every day.
That level of transparency builds trust in a way no sales deck ever could.
Based in West Wales, we serve clients across Carmarthenshire, Ceredigion, Pembrokeshire, and Swansea — and every single one of them benefits from the fact that we have already tested our own recommendations on ourselves. Book a Discovery Audit to see how self-hosted AI could work for your business.
Sources:
- Ollama: Open-source model server — ollama.com
- Mistral 7B: MMLU benchmark score 62.5% — mistral.ai
- Llama 3.1 8B: MMLU benchmark score 73% — meta.ai/llama
- UK ICO: "AI and data protection" — data processing accountability
- EU AI Act: Regulation 2024/1689 — transparency requirements for AI systems
Got a project in mind? Let's talk.
Whether you are just starting to explore AI or ready to build something specific, we can help. Start with a free 15-minute call to scope your idea.