As Indian enterprises race to operationalise AI, the conversation is shifting from which models to deploy to whether the underlying infrastructure can actually support them at scale. Training may capture headlines, but it is inference the continuous, real-time decision-making behind every fraud alert, chatbot response, and personalised recommendation that is quietly becoming the industry’s toughest infrastructure challenge. With over 80% of AI’s computational demand now tied to inference, latency is no longer just a user-experience metric; it is a direct cost and compliance variable, especially in a price-sensitive, regulation-heavy market shaped by the DPDP Act.
In this conversation, Mitesh Jain, Regional VP, Akamai India, unpacks why distributed, edge-first architectures are becoming non-negotiable for Indian enterprises, how sectors like BFSI and fintech are pulling ahead, and what it will take beyond talent and models for India’s AI ambitions to translate into real, scalable ROI over the next three years.

Regional Vice President
Akamai India
CISO Forum: Most of India’s AI conversation has centred on models and use cases so far. Why do you believe inference, not training, is now the real infrastructure bottleneck for Indian enterprises?
Mitesh Jain: The biggest challenge for Indian enterprises is no longer building or accessing AI models, but running AI applications that deliver fast, secure, and cost-effective experiences at scale. Every AI-powered customer interaction, security decision, recommendation, or automation workflow depends on inference, which runs continuously and places far greater demands on infrastructure than model training does. As more organisations put AI into production, they need infrastructure that can handle high volumes of real-time inference without compromising performance or user experience.
The focus now should be on bringing cloud computing closer to where data is generated and consumed. A distributed cloud and edge approach helps reduce latency, improve application performance, optimise costs, and support data residency requirements while delivering consistent experiences across diverse geographies. Enterprises that invest in inference-ready infrastructure will be better prepared to run AI reliably and deliver meaningful outcomes for their businesses and customers.
CISO Forum: When you talk to Indian enterprises today, how many are actually running AI in production versus still stuck in pilot mode? What’s the single biggest thing holding that transition back?
Mitesh Jain: Based on our conversations with customers, we are seeing a clear shift from experimentation to production. The conversation today is less about whether to adopt AI and more about how to run AI reliably, securely, and efficiently at scale.
The biggest challenge is no longer access to AI models but having the infrastructure and operational readiness to support them in production. Organisations need low latency, predictable performance, strong security, and the ability to handle growing inference workloads without driving up costs. Security readiness also remains a key barrier. In Akamai’s 2026 Partner Summit survey, only 17% of respondents said organisations are very well prepared to manage AI security risks, while 70% identified data privacy and leakage as the biggest AI-related security concern. These findings highlight that scaling AI requires not just compute, but also robust governance, security, and resilient infrastructure. A distributed cloud and edge approach helps bring compute closer to users and data, enabling organisations to run AI applications more efficiently while delivering a consistent user experience.
CISO Forum: The government has committed over ₹10,000 crore under the IndiaAI Mission to build domestic compute and GPU capacity. Does this solve the infrastructure gap, or does it address only one part of a much larger problem?
Mitesh Jain: The IndiaAI Mission is an important step in strengthening India’s AI ecosystem and expanding access to domestic compute capacity. Greater availability of GPUs will help lower barriers to AI adoption and reduce dependence on overseas infrastructure. However, compute is only one part of the equation.
To run AI effectively in production, organisations also need infrastructure that can deliver low latency, move data efficiently, support security and compliance requirements, and bring inference closer to users. AI applications are increasingly expected to respond in real time, making distributed cloud and edge infrastructure just as important as centralised compute. India’s AI ambitions will depend not only on building more compute capacity, but also on creating an infrastructure ecosystem that can support AI workloads reliably, efficiently, and at scale.
CISO Forum: Where does distributed, edge-based infrastructure have a genuine advantage over centralised cloud for AI workloads, and where is centralised cloud still the more sensible choice?
Mitesh Jain: Both centralised cloud and distributed edge infrastructure have important roles to play, and the right choice depends on the workload. Centralised cloud remains the preferred environment for model training, large-scale data processing, and workloads that require significant compute resources. It offers the scale and flexibility needed for these compute-intensive tasks.
Distributed edge infrastructure becomes critical when AI applications need to make real-time decisions. Use cases such as AI assistants, fraud detection, content delivery, security, and industrial automation require low latency and consistent performance, which are best achieved by bringing inference closer to users and data. A hybrid approach that combines centralised cloud for training with distributed infrastructure for inference gives enterprises the flexibility to optimise performance, cost, and user experience while meeting data residency and compliance requirements.
CISO Forum: Latency is often framed as a user-experience issue, but you’ve linked it directly to economics. Can you unpack the actual cost implications of running AI inference at scale in a price-sensitive market like India?
Mitesh Jain: Latency is often treated as a user experience metric, but it directly impacts the cost of running AI at scale. More than 80% of AI computational demand comes from inference, making it the largest contributor to ongoing infrastructure costs. Every millisecond of delay consumes more compute resources, and when this is multiplied across millions of inference requests, the cost of delivering AI can increase significantly. For organisations in a price-sensitive market like India, improving inference efficiency is as much a business priority as it is a technical one.
The focus, therefore, should not be on adding more compute alone, but on using infrastructure more efficiently. Reducing latency, increasing throughput, and optimising resource utilisation can significantly lower inference costs while improving application performance. Our own benchmarking has shown that the right infrastructure choices can reduce the cost per inference while delivering better performance, allowing organisations to support more AI workloads, improve service quality, and make better use of their infrastructure investments.
CISO Forum: Which Indian sectors are you seeing move fastest from AI experimentation to real, revenue-linked deployment, and what’s common to those that are pulling ahead?
Mitesh Jain: In India, the sectors that are seeing the fastest move from AI experimentation to real revenue are the fintech and BFSI domains. With one of the highest AI adoption rates in the fintech sector, India has emerged as an innovation hub, fostering solutions such as fraud detection, customised financial products,s and real-time decision-making. In e-commerce, the deployment of conversational AI drives conversion and purchase decisions.
The common factor pulling ahead is the shift in infrastructure. Enterprises are now looking at AI through the lens of a decentralised functional challenge rather than a centralised data project. They realised that if an AI application takes even a few seconds to show results, ROI drops and the end consumer’s interest vanishes. Hence, industry bodies are focusing mainly on the Inference at the Edge model, in which applications can deliver results that drive customer value rather than merely store data in a single repository.
CISO Forum: What happens when an enterprise scales its AI ambitions faster than its underlying infrastructure can support? Have you seen real examples of this backfiring in India?
Mitesh Jain: When AI ambitions surpass the infrastructure, companies hit a roadblock called “The AI Performance Wall”. At a small scale, the enterprise AI models function smoothly. But when the same is deployed to large numbers of users simultaneously, as in fraud detection and risk evaluation, centralised systems collapse. As a result, latency spikes, bottlenecks occur, and the user’s system performance drops.
In India, we’re seeing this challenge emerge across sectors as AI moves from pilot projects to mission-critical production deployments. Financial institutions using AI for real-time fraud detection, credit risk assessment and customer service need inference decisions in milliseconds. Even small increases in latency can impact transaction approvals and customer experience during peak periods.
A leading consumer internet platform in India recently adopted GPU infrastructure to power large-scale generative AI experiences, including AI-generated personalised content for millions of mobile users. While the AI models themselves are highly capable, the real challenge is ensuring GPU capacity, low-latency inference and the ability to scale seamlessly as user adoption grows. This is a great example of how infrastructure becomes just as critical as the AI model itself.
We are also seeing digital mobility and automotive platforms increasingly leverage cloud infrastructure to process vast amounts of location, vehicle and user data while supporting AI-driven services. As these businesses expand, they require cloud platforms that can scale rapidly without compromising performance or reliability, making modern IaaS a foundational enabler for their AI and digital transformation initiatives.
The same trend is visible in media and streaming, where AI is increasingly used for content personalisation, advertising optimisation, and content moderation. Major live sporting events or blockbuster releases can generate massive spikes in concurrent AI inference requests, exposing limitations in legacy infrastructure.
These examples reinforce that AI success is no longer determined solely by the quality of the model. It increasingly depends on whether the underlying infrastructure can deliver low-latency, resilient, and scalable AI services close to where users and applications operate.
CISO Forum: How should Indian CIOs and CISOs think about the trade-off between performance, latency, and cost when architecting AI deployment? Is there a practical framework, or does it vary too much by industry?
Mitesh Jain: Most enterprises deploy AI, assuming it must balance performance, latency, and cost. The reality is that building the right infrastructure could help enterprises enhance all relevant elements by placing their workloads where they are best suited. The right approach is to view AI through a distributed-infrastructure lens. Large-scale work, such as data processing, should be handled from a centralised space, while AI inference should run on edge computing users. This will help reduce costs, improve latency performance, and enhance the user experience. Again, platforms like Kubernetes help enterprises deploy workloads in a decentralised manner and avoid data transfer costs, enabling CIOs to gain better visibility into performance.
What CISOs should be mindful of is building the security infrastructure in this manner from the first day itself. Protection against malware and other threats should be deployed closer to where AI workloads are embedded and carried on. Hence, the entire AI infrastructure must be built with performance, high security, and cost optimisation in equal measure.
CISO Forum: Data residency and DPDP compliance add another layer to infrastructure decisions in India. How much is regulation, versus latency or cost, actually shaping where enterprises choose to run inference?
Mitesh Jain: The terms regulation, latency, and cost are no longer viewed as separate entities; instead, they are shaping conversations in the industry. The DPDP Act has set non-negotiable guidelines for data flow and compliance. Enterprises cannot transfer data overseas for processing or inference purposes.
However, these guidelines and frameworks are also shaping the architecture of data. For instance, if an enterprise chooses to align with the DPDP Act and keep the flow of data within the boundaries for processing, the company starts looking at the decentralised cloud footprint in India. Enterprises can also run inference on edge nodes and distribute cloud computing across Indian states, thereby naturally addressing the compliance mandates of the DPDP Act. This will also reduce latency to single-digit milliseconds and remove the massive cost of routing traffic across clusters.
CISO Forum: Three years from now, do you expect India’s AI story to be told through its models and talent, or through its infrastructure? What needs to happen for the infrastructure layer to keep pace?
Mitesh Jain: India already has an incomparable pool of talent resources and optimised localised AI models and programs. However, three years down the line, India’s AI story will be defined by the execution of its infrastructure. India’s pool of talent, AI models, and applications have immense potential, but with a strong infrastructure backbone, enterprises can convert that potential into ROI.
Three vital things should be kept in mind to ensure the infrastructure layer keeps pace.
- Distributed GPU infrastructure: Enterprises must move away from keeping AI models in Tier-1 data centres. There is a need to expand decentralised GPU systems deployed across Indian states to meet the growing demand for edge computing.
- Dismissing the “Cloud Tax”: Enterprises should move towards greater pricing transparency. Unpredictable fees will obstruct mid-market AI applications if left unaudited.
- Security: As AI is transitioning from chatbots to agentic infrastructure, the security algorithm must evolve. At this point, establishing a robust edge security infrastructure is essential to protect high-frequency AI workloads.
