-
Featured services
2026 Global AI Report: A Playbook for AI Leaders
Why AI strategy is your business strategy: The acceleration toward an AI-native state. Explore executive insights from AI leaders.
Access the playbook -
Services
View all services and productsLeverage our capabilities to accelerate your business transformation.
-
Services
AI
-
Services
Application Services
-
Services
Business Process Services
-
Services
Cloud
-
Services
Connectivity
-
Services
Consulting
-
Services
CX and Digital Products
-
Services
Cybersecurity
-
Services
Data and Analytics
-
Services
Digital Workplace
-
-
Services
Enterprise Networking
-
Services
Enterprise Application Platforms
-
Services
Global Data Centers
-
Services
Infrastructure Solutions
-
Services
Sustainability Services
Accelerate outcomes with agentic AI
Optimize workflows and get results with NTT DATA's Smart AI AgentTM Ecosystem
Create your roadmap -
-
-
Insights
Insights
Recent Insights
-
The Future of Networking in 2025 and Beyond
-
Using the cloud to cut costs needs the right approach
When organizations focus on transformation, a move to the cloud can deliver cost savings – but they often need expert advice to help them along their journey
-
Make zero trust security work for your organization
Make zero trust security work for your organization across hybrid work environments.
-
-
2026 Global AI Report: A Playbook for AI Leaders
Why AI strategy is your business strategy: The acceleration toward an AI-native state. Explore executive insights from AI leaders.
Access the playbook -
-
Discover how we accelerate your business transformation
-
About us
CLIENT STORIES
-
Liantis
Over time, Liantis – an established HR company in Belgium – had built up data islands and isolated solutions as part of their legacy system.
-
Randstad
We ensured that Randstad’s migration to Genesys Cloud CX had no impact on availability, ensuring an exceptional user experience for clients and talent.
-
-
CLIENT STORIES
-
Liantis
Over time, Liantis – an established HR company in Belgium – had built up data islands and isolated solutions as part of their legacy system.
-
Randstad
We ensured that Randstad’s migration to Genesys Cloud CX had no impact on availability, ensuring an exceptional user experience for clients and talent.
-
2026 Global AI Report: A Playbook for AI Leaders
Why AI strategy is your business strategy: The acceleration toward an AI-native state. Explore executive insights from AI leaders.
Access the playbook -
- Careers
Topics in this article
Your large language model knows an extraordinary amount. What it doesn’t know is your business. It’s brilliant but amnesiac.
Data lakes solve the storage problem, but storing data does not make your organization AI-ready. A data lake without context is like a library with every book piled on the floor. The information is all there; finding the right piece when you need it is the problem.
That is why organizations need a context lake. This layer connects data with business meaning, policies, metadata, lineage and relationships to provide the context that helps AI systems deliver relevant and trustworthy outcomes. That context gives AI systems a better chance of producing current, relevant answers grounded in how your business works.
In an AI factory, this context layer is the core of the architecture. That raises three important questions: Where should this organizational intelligence reside, who should control it and what platform should power it?
Locating your AI: A strategic and legal decision
If AI training data at an organization in the European Union leaves Europe, the compliance risk is high.
The General Data Protection Regulation (GDPR) constrained data movement. Now, the EU AI Act and upcoming data residency mandates will make it a legal requirement. Yet many AI platforms still assume that training workloads can run wherever computing is cheapest, rather than where regulated data resides.
AI models trained on EU data often need computing, storage and governance controls that remain within EU jurisdictions. This is why initiatives such as the European High Performance Computing Joint Undertaking have invested heavily in sovereign computing capabilities while organizations reevaluate their dependency on globally distributed AI infrastructure.
Locking yourself into non-EU cloud AI training today could leave you spending the coming years untangling contractual obligations and retraining models. Building governance in from the start is far easier than retrofitting it once AI is in production.
But the challenge extends beyond data residency. Compliance requirements — including training data provenance, model explainability, bias monitoring, auditability, security controls and the ability to respond to regulatory obligations throughout the AI lifecycle — influence how platforms are designed from the outset.
By embedding governance early, you gain a significant advantage. Within an AI factory, identity, observability, lineage, policy enforcement and security controls are integrated into the operating model. This creates AI systems that are compliant, trustworthy and ready to scale as regulations evolve.
The economics change at scale
Public cloud AI platforms offer an appealing proposition: no infrastructure to own and no large upfront capital investment. Just spin up what you need.
However, that’s not the full story. First, there are egress costs. When you move terabytes of data in and out every day, per-gigabyte charges add up quickly. Then there’s latency. Real-time inference has different economics than batch processing, and shared cloud infrastructure has performance limits. Add in fine-tuned models hosted on public cloud platforms, and vendor lock-in grows with every customization.
Public cloud AI services are excellent for experimentation and certain production workloads. At high, sustained levels of GPU use, however, the economics can start to favor private infrastructure — particularly at very large scale. Many large organizations repatriate AI workloads to private infrastructure after seeing these costs explode.
For most organizations, the answer is neither cloud nor private infrastructure alone — it’s hybrid. The real question is, what should run where, and why?
The same applies to the build-versus-buy debate. Successful enterprises build some capabilities, buy others and partner for the rest. The key is distinguishing between core intellectual property, commodities and accelerators. Build what differentiates your business, buy what doesn’t and partner where speed matters most.
An AI factory provides a framework for making these choices systematically. Doing so helps you balance cost, control and agility while achieving sustainable AI economics.
The physics still matter: Networking and power
Everyone counts GPUs, but almost nobody counts the cables. A 1,000-GPU cluster with mediocre networking can perform worse than a properly connected 500-GPU cluster. Training large models requires terabytes of data moving between GPUs every second. If the network cannot keep up, expensive computing resources spend their time waiting. At scale, high-bandwidth, low-latency networking can matter as much as the GPUs themselves.
Many organizations focus on procuring GPUs, then wonder why training runs are slower than expected. They blame hardware or algorithms, when the real bottleneck may be the network fabric connecting the processors. Network engineers understand this and need a seat at the AI infrastructure planning table.
Beneath networking sits a bigger constraint: power. Modern GPU racks can consume 40kW–100kW, while many traditional data centers were designed for only 5kW–15kW per rack. As AI deployments scale, this gap becomes a critical limitation.
Many organizations are discovering they cannot deploy AI infrastructure where they want it because of insufficient power capacity. Expanding power infrastructure can take years. Although liquid cooling improves efficiency, it cannot solve a power shortage.
The problem is timing: You can buy GPUs far faster than you can add megawatts of data center capacity.
This is why an AI factory goes beyond computing. If you’re planning AI infrastructure, power, cooling and facilities capacity have to be part of the conversation from day one.
The talent challenge is growing
Organizations spend millions on GPUs, networking and data center capacity, only to discover they lack the talent needed to run it. AI infrastructure demands expertise in GPU orchestration, large-scale networking, storage, cooling and MLOps — a rare combination that spans traditional infrastructure and AI engineering.
AI researchers often lack infrastructure expertise. Infrastructure teams, meanwhile, may have limited experience with GPU-based AI systems. Only a small pool of professionals understand both.
A practical place to start is with the employees you already have. Experienced infrastructure engineers already understand reliability, scalability and operations. They need AI-specific skills layered on top. Likewise, machine-learning teams must take ownership of infrastructure decisions. Invest in developing these capabilities internally.
Nobody builds private AI infrastructure alone
AI infrastructure partnerships are evolving, and for good reason. The old model, where vendors sold hardware, integrators deployed it and customers operated it, no longer delivers outcomes. AI stacks are too complex, interconnected and business-critical for a siloed approach.
That separation no longer works. GPU, networking, storage, software and security decisions affect one another, so the teams and vendors behind them need to work together from the design stage.
The integrator’s role has changed too: It now involves making sure those different parts are designed to work as one system. Building private AI infrastructure at scale requires partners who can design, integrate and optimize together, not simply sell products.
The hidden cost of waiting
While it may feel prudent, waiting can be the costliest AI infrastructure decision. Every quarter you wait makes three things harder:
- Talent: The best AI infrastructure engineers are joining companies that are already building and operating AI environments.
- Energy: Power availability and increasing costs make future capacity harder to secure.
- Competitors: Organizations that started six months ago have already gained operational experience, customers and lessons that cannot be learned from planning alone.
Many teams fall into the trap of waiting for the next GPU generation. But the bigger constraint is increasingly operational maturity, not just access to the latest hardware. Organizations that began their AI infrastructure journey 18 months ago are already on their second or third iteration. They have also gained an understanding of performance bottlenecks and scaling challenges.
From pipe to pipelines
GPUs are the most visible part of AI infrastructure, but they are rarely the differentiator. Real value comes from everything around them: data context, security, compliance, cost optimization, scalable networking and power, skilled operators and implementation partners. Together, these elements create pipelines that turns data and ideas into measurable outcomes.
The shift is from buying the individual pipes — GPUs, networking and storage — to building those pipelines.
When you’re on that journey, think of AI infrastructure as an integrated platform, not a shopping list that depends on someone else's roadmap and economics.
So, are you still counting GPUs, or are you building an AI factory?