Jensen Huang's Visit to Japan: What Business Needs to Know About the Next GPU Cycle
Jensen Huang's visit to Japan launches a new GPU cycle. What this means for Ukrainian business and AI automation — read our detailed breakdown.

When NVIDIA CEO Jensen Huang flew to Tokyo in May 2025 and signed deals with SoftBank, NTT, and the Japanese government totaling over $200 billion in GPU infrastructure investments, most Ukrainian entrepreneurs barely noticed. And that was a mistake — because Jensen Huang's visit to Japan effectively kicked off the countdown to the next major hardware cycle, which will change the cost of AI compute access for businesses of any size within 12–18 months. Japan became more than just another stop on a corporate tour: it transformed into a testing ground for large-scale trials of the new generation of Blackwell Ultra chips — and the results of these tests will directly impact how much you'll pay for API requests to GPT, Claude, or Gemini.
What Actually Happened in Tokyo — and Why It's Not Just a PR Tour
The official chronicle of the visit reads like a standard set of corporate handshakes. But behind the scenes, a concrete architectural deal was taking place: SoftBank confirmed an order for 100,000 Blackwell Ultra GPUs for its new data center in Osaka, and NTT announced the construction of joint infrastructure with 1 gigawatt of power — more than the combined server capacity of all Japanese cloud providers five years ago.
Blackwell Ultra Architecture: What's Changing Under the Hood
Blackwell Ultra is not just the successor to Hopper H100. Here are the key parameters worth knowing even if you don't understand semiconductors: the chip delivers 2.2x higher performance for inference tasks compared to H100, and the new NVLink 5.0 technology allows connecting up to 576 GPUs into a single compute node without significant speed loss. For business, this means one thing: the cost of a single AI query will continue to drop — and the drop will be steep.
Let's recall the context: between the release of H100 (2022) and its mass availability in cloud services, about 14 months passed. Blackwell launched in late 2024 — so by mid-2026, these capabilities will become standard for Azure, AWS, and Google Cloud. And the Japanese deal signals that regional data center diversification is accelerating, meaning access to new chips will appear at more geographical locations simultaneously.
The Geopolitical Dimension: Japan as an Alternative to Taiwan
There's another aspect that's rarely discussed in business media. The Japanese deals are part of NVIDIA's strategy to reduce dependence on Taiwanese manufacturing capacity from TSMC. The Japanese government allocated ¥4 trillion (~$27 billion) to develop the semiconductor industry, and TSMC is already building a second plant in Kumamoto. For companies building long-term AI strategies, this is an important signal: the GPU supply chain is becoming less vulnerable to geopolitical risks related to the Taiwan Strait.
Similar logic for diversifying risks should be applied to choosing AI providers. If you're concerned about the stability of AI service providers, we recommend reading our breakdown of risks from market leader instability for your AI stack — it contains concrete steps to hedge dependence on a single vendor.
How GPU Cycles Impact Pricing and AI Accessibility for Small and Medium Business
Most SMBs in Ukraine don't buy GPUs directly. They use cloud APIs — and this is where the GPU cycle plays a decisive role. The mechanism is simple: new chips reduce operational inference costs for providers, and competition between them passes this reduction through cheaper rates to end users.
Let's look at concrete numbers. In 2023, GPT-4 cost $0.06 per 1,000 output tokens. By early 2025, GPT-4o mini dropped to $0.0006 — that's 100 times cheaper in two years. This isn't marketing, it's a direct function of increasing GPU capacity and competition between OpenAI, Anthropic, and Google. The Blackwell cycle promises another wave of price cuts — likely a 3–5x reduction in the cost of complex multimodal queries by the end of 2026.
What This Means for AI Project Budgeting Right Now
If you're currently postponing AI implementation because of "expense" — you're making a decision based on outdated data. Concrete example: a company with 50 employees automating incoming request processing through an AI agent based on Claude 3.5 Sonnet currently pays around $80–120 per month for 500,000 tokens of daily traffic. A year ago, the same configuration would have cost $400–600. By the end of 2026 — following GPU cycle logic — the same task could cost $20–40.
But there's a nuance: companies that automate processes now will have a 12–18 month advantage over those waiting for "things to get cheaper." This advantage is measured not just in money, but in accumulated data, tuned models, and a team that knows how to work with AI tools. For more details on how to calculate the real cost and ROI of AI agents, see our honest breakdown of investment in an AI agent for Ukrainian business.
Inference Chips as a New Frontier of Competition
It's worth separately noting a trend directly related to the Japanese deals: NVIDIA is no longer the only player in the GPU space. Amazon Trainium 2, Google TPU v5, and now Broadcom chips for OpenAI are forming an alternative ecosystem. NVIDIA's Japanese partners are betting on NVIDIA — which strengthens the company's position in the region, but at the same time means that competitors have an incentive to aggressively lower prices on their platforms.
The practical conclusion for Ukrainian SMBs is this: consider a multi-provider architecture for AI solutions. Don't lock yourself into a single API. Tools like LiteLLM or LangChain allow you to switch between providers by literally changing a single line of configuration — giving you the ability to automatically move to the cheaper option as new chips hit the market. For more on how this inference chip race is changing the rules of the game, see our breakdown of competition between OpenAI, Broadcom, and NVIDIA.
Practical Steps for Business: How to Prepare for the Next GPU Cycle
Knowing about GPU cycles is good. But knowledge without action is worth nothing. Here's a concrete plan for an SMB owner or manager who wants to end up on the right side of this transition.
Audit Current AI Spending and Growth Points
The first step is understanding where you are now. For most companies, this means:
- Making a registry of all AI subscriptions and API spending (ChatGPT Plus, Copilot, API keys for proprietary solutions)
- Measuring actual usage: how many queries per month, how many tokens, what's the average cost per task
- Identifying processes that are "ripe" for automation but have been postponed due to cost
Second — look at benchmarks. If your current AI task requires GPT-4 level, check if Claude Haiku or Gemini Flash can handle it — they cost 10–20 times less and for structured tasks (classification, data extraction, template response generation) often give equivalent results. For details on how to measure the real autonomy and quality of an agent before scaling, there's a separate practical guide.
What to Build Now to Win from Cost Reduction
The GPU cycle isn't just "everything will get cheaper." It's a shift in which tasks become economically viable for automation. Today, some multi-agent scenarios (when multiple AI agents work in parallel on different parts of a task) still look expensive. But at 5x cheaper token costs, they become standard even for small businesses.
So now it's worth investing not in specific models, but in architecture: building processes where AI agents can be easily replaced or updated. This means:
- Keeping business logic separate from the choice of a specific model
- Building systems based on standard protocols (for example, MCP from Anthropic or A2A from Google)
- Documenting prompts and settings so that migration between providers takes hours, not weeks
For those who want to understand parallel agent architectures — how to run multiple tasks simultaneously without chaos — there's detailed technical material on parallel programming for AI agents.
The Japanese Lesson for Ukrainian Business: Think Infrastructurally
Japanese corporations — SoftBank, Toyota, Fujitsu — are betting on GPU infrastructure not because they're "tech companies." They're doing it because they understand: AI compute is becoming as basic an infrastructure as electricity or the internet. And whoever controls the infrastructure dictates the terms.
For Ukrainian SMBs, "controlling infrastructure" means something different: not building your own servers, but having a clear strategy for which AI competencies remain inside the company. This could be your own RAG system on corporate documents, trained on your business specifics. Or a voice AI agent integrated with your CRM. Or automated HR onboarding.
For example, one Ukrainian recruiting service deployed a RAG system based on internal data — and cut initial candidate screening time from 4 hours to 25 minutes per 100 resumes. Not through magic, but because they built the infrastructure in advance — before it became an obvious solution for everyone. A similar case with RAG implementation in HR processes is detailed here.
FAQ
What is a GPU cycle and why does it matter for my business? A GPU cycle is a regular wave of hardware updates for AI computing that occurs every 2–3 years. Each new cycle reduces the cost of AI queries and makes new types of tasks accessible. For business, this means you need to review your AI usage strategy at least once a year.
Do I need to buy GPUs to benefit from the new cycle? No. For 99% of Ukrainian SMBs, GPUs remain infrastructure for cloud providers — AWS, Azure, Google Cloud, or specialized AI APIs. Your job is to choose the right providers and build a flexible solution architecture that allows you to quickly switch between them.
When exactly will the new Blackwell chips become available through standard cloud services? Mass deployment of Blackwell on cloud platforms is expected in the second half of 2025 — first half of 2026. The SoftBank and NTT data centers mentioned in this article plan to launch their first commercial Blackwell Ultra services in Q1 2026.
How does the GPU cycle affect open-weight models like Llama or Mistral? New GPUs make running large open models significantly cheaper even on proprietary equipment. This accelerates the divergence between leading open-weight and commercial models: in 2026, open models will be able to perform tasks that today require GPT-4 or Claude 3. More on this trend in the article about open-weight models in 2026.
Should I stop AI projects and wait for prices to drop? No, and here's why: companies implementing AI now accumulate data, experience, and processes. When cost reduction happens, they'll be able to scale with minimal overhead — while those who waited will spend time and money training their teams and setting up processes from scratch.
Huang's Japan visit is not just a corporate news item. It's a barometer for where the AI industry is heading over the next two years: more capacity, lower costs, broader geographic accessibility. For Ukrainian business, this is a window of opportunity that's opening right now — and the smartest move is not to wait, but to start building flexible AI infrastructure today while competitors are still asleep. If you want to understand exactly where to start in your specific situation — book a 15-minute consultation with our team: we'll help you develop a roadmap for AI automation that takes your budget and current process state into account.
Have questions? Ask the AI agent right now
Responds in seconds, knows everything about our services and will help with your situation
You might also like
Power AI: Why Energy Companies Are Snapping Up Land for Data Centers — and What It Changes
Power AI is rewriting the rules: energy companies are buying up land for data centers. What it means for your business, AI pricing, and the decisions you need to make right now.
NewsAnthropic Restores Access to Claude Fable 5 and Mythos 5: What It Means for Your Business After the Cybersecurity Block
Anthropic has restored access to Claude Fable 5 and Mythos 5 following a U.S. government review. What changed, which new classifiers now protect the models, and what it means for your business.
NewsWhy Christopher Nolan Is Right: AI as a "Trojan Horse" in Corporate Security
AI as a trojan horse in corporate security: real threats to your business, data breach cases, and practical advice for protection.
