Its recommendations are based on its own measurements. A proprietary testing process evaluates a range of models—from coding assistance to speech-to-text—over the course of a weekend on specified hardware. At the same time, it answers the three questions that repeatedly arise in projects: How well does the hardware fit the use case? How quickly does the model respond? And how many people can use it simultaneously?
Good Ideas Shouldn’t Fail Because of Sensitive Data
b.telligent often hears the same story in client conversations: AI may only be used with non-sensitive data, while the applications that would deliver the greatest benefits are often left on the table. Why? Because the underlying data is too sensitive. There is also legal uncertainty. Whether data may be transferred to the United States must be reassessed on an ongoing basis; following a decision by the US Supreme Court this summer, the European Data Protection Board asked the European Commission in July 2026 to review the basis for these transfers. Companies that operate models in-house are on the safe side here, because their data simply stays where it is.
“Anyone deploying an AI application on a US service today is building on a foundation that could shift within a few months. After Safe Harbor and Privacy Shield, this would be the third disruption in ten years,” says Sebastian Amtage, Founder and Managing Director, Germany, at b.telligent. “And that matters greatly because language models are becoming just as important as ERP systems and databases. That is why it is valuable to be able to decide for yourself where they run, which version you use, and when changes are made.”
Measured Precisely, Not Estimated Broadly
There are many figures circulating on how much computing power a language model requires and how well it performs—but these figures often cannot be replicated in practice. The standard leaderboards provide little help: They assess how intelligently a model responds, not how many colleagues can work with it at the same time or how smoothly it delivers responses. That is why b.telligent conducts its own measurements, updates the results whenever new models or hardware become available, and, where needed, tests a client’s preferred models on the client’s own servers.
“The question is rarely whether an openly available model is good enough,” says Kai Kalchthaler, Managing Director, Germany, at b.telligent. “The question is what it delivers on our hardware, how many people can use it simultaneously, and what it costs per month. These are precisely the figures missing from most business cases—and without them, every purchase becomes a bet. That is why we run our own measurements and disclose how we measure. This allows anyone to verify whether our recommendation holds up.”
Your Own Model in Three Steps
The process starts with an assessment: Which use cases offer the greatest potential, how sensitive is the underlying data, and what technology is already in place? The result is a prioritized roadmap that can be implemented step by step. The second phase focuses on selecting the model and suitable hardware, as well as determining how the model will access the company’s own documents. Finally, the environment is set up, the first application is connected, and operations are handed over to the client’s team.
“In-house” does not necessarily mean running your own servers in the basement: A self-controlled environment with a European provider can serve the same purpose. The first step is a no-obligation conversation about which use cases are worth pursuing in the first place. The contact person is Dr. Michael Allgöwer, Management Consultant Data Science & AI at b.telligent.
Learn more: Local LLMs: Use AI Without Giving Up Control of Your Data