A few days ago, Splash247 published a quote from Daniel Weiss, general manager of shipping strategy at Vale, on certain maritime AI products: "It's a solution looking for a problem, not the other way around." I read it twice. That's exactly it.
Lloyd's Register counts 420 organizations active in maritime AI development, up from 276 last year, a 52% increase in twelve months. A Thetius/Marcura study found that 81% of maritime companies are running AI pilots, meaning controlled test deployments on a single vessel or route before any fleet-wide rollout, while only 11% have formal policies governing how to scale them. The market nearly doubled, 81% is testing, and one in nine knows what it will do if it actually works. Call it enthusiasm without a plan.
The reason isn't a mystery. The industry just doesn't say it out loud: shipping is running AI tools on data that isn't ready for AI. The Lloyd's Register Digital Maturity Index puts the sector at 2.1 out of 4. Data standardization, without which no AI model performs reliably, sits at 2.45. In practice: operational data entered manually by crew who round numbers at the end of a watch. One vessel logging fuel consumption in metric tonnes per day, another reporting in kilograms per nautical mile.
Sensor outages that go unlogged. Fields with inconsistent definitions across vessels in the same fleet. Data that only reaches shore when a ship enters port. Garbage in, garbage out. Then everyone wonders why the AI failed to deliver. Clean data, for comparison, means sensor readings captured automatically at regular intervals across the entire fleet, transmitted continuously rather than at port, stored in a single standardized format regardless of vessel age or flag state, with field definitions that mean the same thing on a bulk carrier in Rotterdam as on a tanker in Singapore. Very few shipping companies have this. Most are somewhere in between: partial automation, partial manual entry, multiple formats coexisting, and someone onshore trying to reconcile them before the AI ever sees a number.
Maritime has a structural problem here that few industries share. A manufacturing plant or logistics warehouse can have data silos too, but they're fixed and they can be integrated. A fleet of vessels moving across oceans for months at a time is a different problem entirely. Data is collected across different systems, from different vendors, in different formats, and the fleet manager in Piraeus does not see the same picture as the superintendent onboard. Connectivity is intermittent, sensor maintenance depends on who's on watch, and the incentive structure at sea rarely rewards meticulous data entry. In that environment, an AI voyage optimization tool can generate useful recommendations, provided it receives clean, standardized data from across the fleet.
Without it, the tool simply produces wrong answers faster. Behind that 11% figure, there is no single explanation. Some companies haven't seen results reliable enough to justify formal policies. Others don't know what AI governance means and assume it falls to the vendor.
A third group hasn't been told that what they're already running should have governance. All three have something in common: they don't know exactly what the AI they purchased is actually doing. Governance requires knowing what you're governing. If a pilot ran on one vessel for six weeks and the AI made recommendations the crew either followed or didn't, it's genuinely difficult to know what worked, why, and whether it can be replicated fleet-wide. That is not a policy failure. It is a measurement failure. A measurement failure means the pilot was designed to run the tool, not to prove anything. No success metric agreed before go-live, no independent baseline documented, no plan for how to evaluate the results afterward. The six weeks ended, someone asked whether it worked, and the honest answer was: we don't know, because we never set up the conditions to find out. That is fixable. But only if you decide before the pilot starts that measurement matters.
The successful cases share something the sales decks tend to skip. Cargill deployed virtually its entire time-chartered fleet through ZeroNorth. VTS Shipping recorded $4.6 million in verified bunker savings in 2025.
Both had invested in their own data infrastructure before buying AI tools. That didn't happen in a six-week pilot.
If you want to be in that category, start with an honest inventory: how many systems store operational data across your fleet, how many of them connect to each other, and how many rely on manual entry. If you don't know, that is your starting point. Not a new tool purchase. When evaluating vendors, four questions separate those with genuine evidence from those with good presentations. What was the baseline before the tool? Every "we saved 8%" is a comparison against something; vendors have a documented habit of selecting favorable comparison windows, and if you don't know the methodology, you don't know the number. Who verified the results besides the vendor? A vendor measuring its own performance is a student grading their own exam. Do the results hold across different vessel types, or only on the newer ships with reliable sensor coverage that tend to appear in pilots? And is there a live customer I can actually speak with, not a written testimonial, because testimonials don't say what went wrong or whether the savings held into year two.
What's missing in most cases isn't better AI. It's data workflow, clear processes, and the decision to understand what's running in your systems before purchasing the next tool. And "what's running in your systems" is not only a performance question. It's a security question too. That's a whole other chapter.
By Maria Bartzoka
Maria Bartzoka is an AI specialist with a focus on AI security and CEO of ThetaAI (thetaai.gr).