Under the surface of the explosion of AI companies and products there is an infrastructure battle taking place – it is a battle to control access to your content and make it a part of an AI company’s training data. It is a battle to control the path of your information.
There is a spectrum of AI usage types today:
- Individual users on a chat product using a web browser for chatgpt.com or claude.com or similar.
- Individual users moving to the desktop client of one of these vendors now packaged as “unified” desktops incorporating chat, coding and work features
- The enterprise account equivalents of either of the first two, with account controls, budgeting, etc.
- Token/consolidation products like software development tools which let you pick AI endpoints for coding. Some of these are configured with your individual API keys to your different accounts and others provide an API token to their service and charge an “uplift” as they route your traffic, basically token re-selling.
- More comprehensive model routers like OpenRouter (just acquired by Stripe) which let you multiplex and set workflows for where to send different types of queries, and let users pick AI targets dynamically as well.
All of these are trying to ensure that they are in path to your interactions with the LLMs, including potentially your own private LLMs. Why? Because they want to capture your users’ interactive sessions with the LLMs.
Hot in the recent news is it appears that mathematicians who spent countless hours independent of LLMs pursuing a significant mathematical challenge, applied their domain and research knowledge to guide and direct their “interrogatory sessions” with an LLM – only to have the LLM company take that synthesized knowledge as their own and claim to have solved the mathematical problem.
The business enterprise stakes are perhaps smaller but omnipresent.
The user might be an expert lawyer, with a deep capable memory, depth of experience in industrial operations litigation, who is using LLMs in their work. The captured, iterative conversation stream from that person, as they ‘leave’ an LLM session satisfied with the result of the interactions, is likely to be valuable. The human user directed a synthesis of the knowledge available to the LLM, and the person’s knowledge and lived actions. This synthesis is likely to have never been captured in the training set. UNTIL NOW.
The user might be a software designer who through the LLM interactions has that ‘aha’ moment about how to tackle a problem, creating new knowledge. Even if the LLM tried to “one shot” a response to the designer’s initial prompt, let’s assume the person responded with critique and re-direction. This interrogative stream likely contains, if not entirely new knowledge to the LLM, then curated knowledge, which adds to the value of the LLM provider. IF CAPTURED.
Everyone wants you in-path to use your knowledge synthesis as the new training data. Every day the AI services are being fed “AHA” moments that experts have had while making use of the LLMs as a super knowledge base – but in fact shaping, guiding, let’s say providing, much of the intelligence.
Which leads me to repeat “Kerpan’s Law of AI”:
There are two types of AI companies:
- those that tell you the are using your content
- those that declare they are not, but are lying
If you do not control your path to LLMs this will be your forever situation. At Cohesive Networks we think organizations are on a journey from early adoption of LLM’s as a Service (ChatGPT, Claude, Kimi, etc.) and over time will gradually migrate many of these interactions to self-hosted AI (regardless of in-cloud, hosting provider or on-premise). The key point will be explicit decisions of how much to share with the “as a service” players.
In the past most of us have been too busy, too introverted, too scared, too worried about HR and Legal, to post detailed interactive transcriptions of our thoughts, of team brainstorms, of the fragment of solutions which are not the heart of a product or service, but provide insight and capabilities to products and services. BUT, in the confessional of the chat interface, given access to something in many ways worse than search, but in other ways better, a more comprehensive knowledge base, we type, guide, shape, drive to new knowledge and give it away. No, in fact pay for the privilege.
I am not privy to any such fact, but, the LLM as a Service companies have to have a scoring mechanism in place for queries or users, perhaps both. I would surmise that users whose primary usage is ‘What’s the best cat food for Calicos in February?’ or ‘How do I clean oil stains off my garage floor?’, while perfectly fine customers, might have a low score for bringing “additive knowledge” to the training data. But there are some users who you might argue should be paid for interacting with the LLMs.
This is why they want your interactions to flow through their path. Some subset of those interactions are worth their weight in gold as higher quality training material than much of what has been trained on from Reddit, Stack Overflow, etc. Maybe “better” is even the wrong word. It is at the very least tapping into to the knowledge creation capabilities of a vastly larger number of people.
WHAT DO YOU DO?
Our belief is that there are the technical components and policy elements to this situation.
In summary, you will need to have corporate policies that define acceptable AI usage. This will include approved vendors, tools, and modes of usage. I am not sure what the “carrot” will be, but the “stick” could be something like “violating this policy is grounds for termination, and possibly further ramifications”. Frankly there will be a number of cases where you will be “catching the cows just out of the barn”, but at least quickly and not long after. Your policy might allow OpenAI and Anthropic but only via a manged corporate account, maybe browser-based usage only, whilst their desktop applications are not allowed (unless perhaps on an sterile desktop with limited connectivity). Of course there needs be strictures against feeding PII or PCI data, details of secret projects, corporate actions (possible acquisitions/divestitures). These details will have to be stated in policy for the protection of your company, the employees (clear expectations, no surprises), and your customers.
You will need technical guardrails and enforcement. For example, at Cohesive we have had controls like virus scanners and device management. Moving forward we will be implementing even stricter controls so that many of the AI wrapper applications can be prevented from running on our laptops/desktops. (My personal view is ‘Only in a container in a VM on a Linux host, in the cloud’).
So far we have proposed policy of “don’t do that or you will get in trouble” and “we won’t let you do that on your work machine”, but what about everything else? What are the technical components?
At Cohesive we have created an in-path AI Firewall in a virtual appliance that we call WaiF™.
It works in conjunction with our VNS3 Network Platform. Overall, we think a network virtualization platform is a prerequisite to meet the challenges in a world of LLMs. While an entire class of vendors have declared the VPN as dead, we are an outlier, put everyone on a VPN, essentially an AI-VPN, at least for accessing LLMs. Our WaiF solution integrates to your DNS, collaborates with your WAF if you have one, and creates both alert-in-the-wire and block-in-the-wire capabilities while capturing all of the traffic to your approved LLMs, sending the data to object storage, multi-modal databases, data lake infrastructure, etc. that you already have in place. A simple place to start is getting it all into object storage with a lifecycle retention policy and auditable/browsable by a self-hosted ElasticSearch dashboard.
Aren’t we at Cohesive just trying to get in your path too? No, we are letting you control your path, we aren’t in control of this infrastructure you are.
This path gateway approach is the beginning of the inference/intelligence/AI journey every organization is embarking on. Where any given organization ultimately arrives is “to-be-determined”. There are businesses where using LLMaaS maybe fine, without ever moving to private or local AI. BUT, even these organizations need an audit trail, need the best practice of having observability, and in large part visible controls over the usage of AI for their business.
For other organizations the AI Firewall with a network platform is the junction point, the place to hook in additional elements, semantic scoring for example. You can use your organization’s knowledge of itself to score interactive sessions, determining content which just went off to an LLMaaS, but should now be captured separately as knowledge for your private AI. As the technology evolves you will be able to “front run” specific users and query types to go immediately to local AI as the frontier model capabilities of today become the local AI of tomorrow. The accomplishment will be keeping more of your specific intellectual property and accumulated intellectual capital in-house.
As the technology is moving so quickly in multiple dimensions, our primary advice is “start”. You can’t wait for the final “all singing, all dancing” product as it is unlikely the industry is going to be stable anytime soon. Have joints, junctures, hooks to be in control, and among these, control your path.
