Enterprise RAG in Hong Kong: build vs buy
Buy a RAG platform when your documents are clean, single-language and not confidential. Build when the data cannot leave your infrastructure, the corpus mixes English and Traditional Chinese, or every answer has to cite its source. Here is how to tell which side you are on.
When should you buy a RAG platform, and when should you build?
Buy when your corpus is clean, single-language and non-confidential, and the answers do not have to survive an audit. A platform is faster and cheaper, and an off-the-shelf local RAG product can be enough.
Build when the data cannot leave your infrastructure, when retrieval has to work across English and Traditional Chinese, or when every answer needs a citation a reviewer can check. Those three conditions are common in Hong Kong banks, insurers, family offices and professional-services firms, and they are where packaged retrieval usually breaks.
| Your corpus | Buy a platform | Build your own |
|---|---|---|
| Language | One language | English and Traditional Chinese together |
| Format | Clean prose | Rate tables, scanned annexes, images |
| Confidentiality | Can go to a hosted service | Must stay on your infrastructure |
| Answers | Helpful is enough | Must cite the source passage |
Why does enterprise RAG fail in production?
In the systems we have rebuilt, retrieval failed rather than generation. Each of these four failures is silent: nothing errors, nothing alerts, and the answer comes back wrong and fluent.
1. The table gets split down the middle
Hong Kong enterprise documents are rate tables, disclosure matrices and scanned annexes, not the clean prose tutorials assume. A naive chunker cuts a table in half and the retriever confidently returns the wrong half.
2. The query is in English, the document is in Chinese
A board pack in English cites a filing in Traditional Chinese. Single-language embeddings never retrieve the other half of the estate, so cross-lingual retrieval has to be an architecture decision, not a translation step bolted on at query time.
3. The answer has no source attached
In audit, compliance and investment work, an answer you cannot cite is a liability. Retrieval tuned for summary produces fluent paragraphs nobody can defend; retrieval tuned for citation returns the supporting passage alongside the answer.
4. Nobody built the evaluation set
Teams demo, ship, then discover edge cases in production. A labelled evaluation set built before the first deployment and re-run on every release is what turns a demo into a system.
How do you run RAG on confidential Hong Kong data?
Keep retrieval inside your own perimeter and design for four constraints before anyone picks a model:
- Data residency. Client data under the PDPO, the HKMA’s outsourcing module (SPM SA-2) and its 2022 cloud-computing guidance. Self-hosted retrieval and a self-hostable embedding model keep the corpus inside your perimeter.
- Explainability. The HKMA’s GenAI guidance expects an institution to explain an AI-assisted output. Citation-first retrieval is the practical answer: the source paragraph ships with the answer.
- Human oversight. High-risk outputs need a review gate, and the system should refuse rather than guess when the library does not cover the question.
- Audit logging. Log the query, the retrieved evidence and the generated answer, so a disputed output six months later can be reconstructed.
How do you make retrieval work in English and Traditional Chinese?
Design for it from the start: a self-hostable multilingual embedding model such as BGE-M3, language-aware chunking, and hybrid vector and keyword search so exact identifiers, such as fund codes and clause numbers, still match. Then test it with questions in one language against documents in the other, in the evaluation set, before launch.
Who does RAG in Hong Kong?
The market splits into a few kinds of provider, each right for a different buyer:
| Provider | What they offer | Pick them when |
|---|---|---|
| HKPC | Institutional programmes and showcases | You want funding, training and a first look |
| HKU RAG-Anything | Open-source multimodal RAG framework | Your team can run and harden open source itself |
| ThinkCol | Local GenAI consultancy with a RAG chatbot platform | You want a platform plus a project team |
| GPTBots / Aurora Mobile, Wavenex, MasterConcept | Upload-your-documents chatbot products | The corpus is clean and answers need not survive an audit |
| UD (InfiniAI), DMS Solutions (elDoc) | Private-AI and document-intelligence products | Your documents fit a packaged product |
| Super Cat Technology | Embedded engineers who build on your infrastructure | The corpus is bilingual, semi-structured and confidential |
Super Cat is vendor-neutral on build versus buy: it earns on engineering time, not licences, and will point you to a platform when a platform is the better fit.
What does a custom build involve?
It starts with a review of the document estate: formats, languages, authoritative versions and access rights. Then comes an ingestion and retrieval build tuned to that estate, an evaluation set and instrumentation. Timelines depend on the corpus, not the model; scanned bilingual filings take longest to ingest.
Super Cat has shipped three such systems, all on client infrastructure: retrieval over a Big 4 firm’s compliance library, private retrieval for a family office, and a multimodal RAG application over text, tables and images.
Frequently asked questions
What is the best Hong Kong RAG solution for a regulated enterprise?
It depends on the corpus, not the model. For clean, single-language, non-confidential documents, an off-the-shelf local RAG platform can be enough. For bilingual, scanned or confidential documents where every answer must cite its source, build on your own servers.
Why do RAG systems return wrong answers when the document is in the index?
In the systems we have rebuilt, retrieval rather than generation was at fault: tables split by chunking, queries in one language against documents in another, retrieval tuned for summary rather than citation, and no evaluation set.
Can a RAG system stay entirely on our own servers?
Yes. Self-hostable embedding models and vector stores make a fully on-premise stack practical. Both Super Cat’s Big 4 compliance system and its family-office platform run entirely on the client’s own infrastructure.
Does Super Cat sell a RAG platform?
No. Super Cat embeds engineers who build retrieval on your infrastructure, on time and materials, and earns on engineering time rather than licences.
The service, the three systems shipped, and the market map.
Embedded AI teams for Hong Kong enterprises.
Eight production AI case studies.
The founder who leads the RAG work.

About the author
Cat Yung is the founder of Super Cat Technology, an AI agent engineering team in Hong Kong and London, and works as an AI expert, data scientist and fractional CTO. Over a decade in production AI, NLP and machine learning; Expert-Vetted top 1% on Upwork with a 100% Job Success Score (source: Upwork profile, October 2026).