AI on-premise: When to Collect Data In-House

24/09/2026

AI on-premise: When to Collect Data In-House

A Lesson from a Factory in Long An

I still clearly remember a late afternoon last year, sitting in the control room of a food processing plant in Long An. The production director pointed to a large screen where the status "Model Confidence: 12%" was blinking continuously. He sighed, his voice heavy with frustration: "We bought expensive AI on-premise software, integrated all our cameras and production data into the system, yet it still makes more mistakes than a human." This is a typical scenario we often encounter when Vietnamese businesses begin deploying internal AI. They want to keep data within the factory walls, which is entirely reasonable from a security standpoint. But they forget that raw data does not automatically turn into knowledge. It needs to be "cooked" properly. If the input ingredients are flawed, the output will be nothing but waste. This article is not a dry checklist. It is a collection of real-world stories about the crossroads you will face when deciding to take control of your training data.

AIVISION solution demo
A short video covering the main points of this article.

The Scale Dilemma: When Is Enough Enough?

Many technical leaders share the same mindset: "The more data, the better." This is the most common trap. I once saw a client in the lubricant industry collecting up to 50,000 images daily from their bottling line. The number sounds impressive. But in reality, 90% of those were duplicate frames or had poor lighting quality. The result? The model overfitted to the specific lighting conditions of the 2 AM shift and was completely blind when the day shift started. Collecting training data is not a marathon; it is a selective exercise. You need enough diversity for the model to understand the essence of the problem, but not so much that it gets buried in noise. A practical rule we often apply at AIVISION, when working with projects at Masan or Meat Deli, is to focus on annotation quality rather than absolute quantity. Around 3,000-5,000 accurately and diversely labeled samples are usually more effective than 50,000 chaotic ones. Ask yourself: Does our data accurately reflect the reality of the production line, or is it just random moments?

The Trap of Balanced Data

Imagine you are a factory security guard. Over 100 days, no thieves appear. You only see 99 peaceful days and 1 incident. If you train an AI to identify thieves based solely on that data, it will consider "no thieves" as the normal state and ignore all signs of abnormality. This is the data imbalance problem. In manufacturing, defects or faulty products often account for only about 1-2% of the total volume. If you collect training data passively, the system will be "blind" to the very things it needs to detect. The solution is not to wait. You must actively create data or use synthetic data techniques to augment rare cases. I once worked with an instant noodle manufacturer where packaging errors occurred only a few times a week. We could not wait for enough data. Instead, we used data augmentation techniques and simulated various error types. As a result, the model began to "see" small details that the human eye easily missed. What is the condition for this to work? You need a technical team capable of evaluating whether synthetic data truly resembles real data. If not, you are training AI based on illusions.

Where to Keep Data: The Trade-off Between Speed and Safety

The term "AI on-premise" sounds simple: place the server in the factory. But reality is much more complex. Do you want training data to remain on edge devices for fast processing, or push everything to a central data center to train large models? Each choice has its price. If you keep data at the edge, response speed is very fast, but updating the model becomes a logistical nightmare. You must physically move devices or use data transfer methods to update algorithms. Conversely, if you push everything to the center, you have massive computing power, but latency can increase, affecting real-time quality control processes. For high-speed packaging lines, even a 200ms latency can lead to thousands of defective products passing through. I recommend considering a hybrid architecture. Raw data stays at the edge, but typical samples and error data are sent to the center for periodic retraining. This is how we deployed for a partner in the beer industry in Thailand. They needed high speed but also needed the model to adapt to seasonal changes in product labels. This balance requires continuous operations, not a one-time setup.

When NOT to Do It Yourself

This is the most important part. There are times when you should stop and realize that collecting and processing data in-house is a bottomless pit. Consider the following signs. First, your production line changes continuously. If you change packaging every quarter and machinery every year, old data will quickly lose value. Data maintenance costs will erode the profits from software cost savings. Second, you do not have a team of computer vision experts. Data labeling is not just drawing rectangles. It is understanding the production context. If the labeler does not understand why a scratch is a defect, they will label it incorrectly, and the model will learn incorrectly. Third, your production process is too complex and unstable. If each batch has unique characteristics, standardizing data will take more time than the benefits it brings. In these cases, the most effective solution is to use foundation models that have been pre-trained on general data, and only fine-tune them with a small amount of specific data. Do not try to catch fish by hand if you can use a net. Sometimes, humility in acknowledging technical limitations yields better results for the business.

Stability Is a Process, Not a Destination

Deploying internal AI does not end when you successfully press the "Train" button for the first time. It has just begun. The model will drift when the environment changes: lighting shifts with the seasons, machinery degrades, and workers change. If you do not have a continuous process to collect, evaluate, and update training data, your model will only be stable for the first few weeks, then gradually become useless. I often compare AI data to fresh food. It has an expiration date. You need a smooth-running data pipeline where new data is automatically detected, quality-assessed, and fed into the retraining process. At AIVISION, when working with multinational corporations like TTN or partners in Mexico and the Philippines, we emphasize this aspect. We build model performance monitoring systems that alert immediately when accuracy drops below the acceptable threshold. You do not need to build everything from scratch. But you need to understand that maintaining model stability requires operational discipline that is just as important as initial development. Ask your organization: Are we ready for this "invisible" but costly work, or are we just viewing AI as a one-time purchase, one-time use software?

There is more here than one article can hold. Keep reading on the AIVISION blog, look at our display scoring solution, or get in touch.

Related insights

See all insights