Skip to Content

Claude Code Rakuten Case Study: 79% Faster Feature Delivery

What Rakuten reported about Claude Code, vLLM, 7 hours, and a 24-to-5-day delivery comparison
2026-04-27 20:57:21 Updated 2026-08-21 21:10:51.550722 — min read 333 views
Claude Code Rakuten Case Study: 79% Faster Feature Delivery
Claude Code Rakuten case study reports one engineer assigning a difficult vLLM implementation to an AI coding agent. Rakuten says the run lasted 7 hours, reached 99.9% numerical accuracy, and helped reduce reported feature time from 24 days to 5 days. The result is useful evidence, not a universal productivity guarantee.

What You'll Learn

  • What Rakuten says Claude Code did inside the vLLM project
  • How to read the 7-hour, 99.9%, and 24-to-5-day figures without overstating them
  • Why agentic coding still needs human direction, review, and security controls
  • How software teams can test a similar workflow with evidence instead of slogans

The Claude Code Rakuten case study is interesting because it describes a concrete engineering task rather than a vague claim that an AI assistant made developers faster. Rakuten says Kenta Naruse, a Machine Learning Engineer, asked Claude Code to implement an activation vector extraction method in vLLM. The project involved a codebase that Rakuten describes as having 12.5 million lines across multiple programming languages.

The published account says Claude Code completed the task in 7 hours of sustained agent-run work in a single run. Rakuten also says the result reached 99.9% numerical accuracy against the reference method. Separately, Rakuten reports that its time to market for new features fell from 24 days to 5 days, which it presents as a 79% reduction.

Those figures deserve attention. They also need boundaries. The official case study is one company account about one implementation and one reported delivery comparison. It is not a controlled trial across teams, repositories, models, or project types. The useful question is not whether every engineer will get the same result. It is what the case reveals about task selection, agent permissions, verification, and the human work that remains after code generation.

What Rakuten Actually Reported

Rakuten’s official account says its engineering teams use Anthropic’s Claude Code to automate coding tasks and accelerate time to market. The page identifies three headline results: 7 hours of sustained agent-run coding on a complex open-source refactoring project, a reported reduction from 24 days to 5 days for new feature time to market, and 99.9% accuracy on complex code modifications.

The wording matters. Rakuten reports the result as an example of how an engineering team used an agent. It does not say that Claude Code writes every production change without oversight. It does not publish a universal success rate. It also does not provide enough detail on the 24-day and 5-day measurements to treat the difference as a controlled causal estimate for all feature work.

A source-audited reading keeps the claim strong but narrow. The task was difficult, the run was long, the reported numerical comparison was high, and the company says its delivery process changed. Those are meaningful facts. The next step is to inspect what the task required and how the team could have checked the result.

Reported itemWhat the source saysSafe reading
TaskActivation vector extraction in vLLMA specific machine learning implementation
Repository12.5 million lines across multiple languagesRakuten’s description of the codebase
Agent run7 hours in a single agent-run sessionA reported case result, not a normal operating target
Accuracy99.9% numerical accuracy against a reference methodA task-specific numerical comparison

For a related account of the same topic, see our Claude Code 7-hour agent-run test analysis. It should be read alongside the official Rakuten source, not instead of it.

The vLLM Task and Why It Was Difficult

An activation vector extraction method is not a routine text edit. It sits inside a machine learning system where tensor shapes, numerical behavior, reference implementations, and runtime paths all matter. The task also involved vLLM, an open-source inference library that Rakuten describes as a large codebase written across multiple programming languages.

That combination creates several sources of failure. A coding agent can place a function in the wrong module, misunderstand a tensor convention, introduce a precision error, or produce code that passes a narrow test while failing under another model. Large repositories add another problem. The correct change may depend on documentation, call sites, build rules, test fixtures, and conventions that are not visible in one file.

The case therefore tests more than code completion. It tests repository search, context selection, planning, tool use, implementation, execution, and comparison against a reference. A long agent run can be useful when the work is decomposable and the repository has enough signals for verification. It can also create a large review burden if the agent changes too much without leaving a clear trail.

That is why the 12.5 million line figure should not be treated as a claim that the agent held every line in active context at once. It describes the repository scale reported by Rakuten. The engineering question is how the agent selected relevant files, commands, tests, and references while working inside that repository.

What Seven Hours Means in This Case

Rakuten says Claude Code completed the implementation in 7 hours of agent-run work in a single run. The case study reports that Naruse provided occasional guidance while the agent worked. This is a useful description of delegation. It is not proof that the agent required no engineering effort before or after the run.

Before an agent starts, someone must define the task, identify the reference method, choose the repository state, decide what access is allowed, and set the acceptance test. After the run, someone must inspect the diff, reproduce the numerical comparison, check security implications, and decide whether the change belongs in the product. The hours reported by Rakuten describe the agent’s sustained coding run, not the complete cost of responsible delivery.

That distinction changes how teams should measure agent work. Track setup time, agent time, human review time, rework, test coverage, rollback risk, and the value of the shipped change. A shorter command session can still produce a slower release if review becomes a bottleneck. A longer session can be worthwhile if it completes a task that previously required several handoffs.

How the 24 to 5 Day Comparison Was Calculated

Rakuten says the time to market for new features fell from 24 days to 5 days after its use of Claude Code. The arithmetic behind the reported 79% reduction is straightforward. The difference is 19 days, and 19 divided by 24 is about 79%. But the arithmetic is not the same thing as the study design.

The source does not provide a controlled experiment that holds feature size, team composition, review policy, release calendar, and dependency work constant. It also does not say that every feature moved from 24 days to 5 days. The report is best presented as Rakuten’s observed or reported comparison for its development process.

This is not a weakness that makes the case useless. Engineering teams make decisions with internal before-and-after evidence all the time. The discipline is to label it correctly. A reported process improvement can guide a pilot. It should not become a promise in a sales deck or a forecast for a different repository.

FigureSource contextEditorial limit
24 daysRakuten’s reported earlier feature time to marketDo not assume every feature had this duration
5 daysRakuten’s reported later feature time to marketDo not assume the same result elsewhere
79%Rakuten’s reported reduction from the two figuresReport it as a case comparison, not a universal effect
7 hoursOne sustained Claude Code implementation runDo not equate it with total project cost

Our AI cost versus human worker analysis provides a broader way to think about task economics. The Rakuten case adds a specific engineering example, but it does not replace workload-level measurement.

What 99.9% Numerical Accuracy Does and Does Not Mean

Rakuten says the implementation achieved 99.9% numerical accuracy compared with the reference method. That phrase is more precise than saying the code was 99.9% correct in every respect. Numerical accuracy usually refers to a comparison between outputs under a defined test or metric. It does not automatically cover API design, error handling, maintainability, security, or behavior outside the tested inputs.

A reference comparison is valuable for a method that operates on model data. Small numerical differences can accumulate or alter downstream behavior. Still, a team should ask what was compared, what tolerance was used, how many inputs were tested, and whether the reference itself was independent. The source summary available on Rakuten’s page does not answer every one of those questions.

Teams should separate at least three checks. First, numerical equivalence asks whether the new method matches the reference. Second, software correctness asks whether it works through the supported API and runtime paths. Third, production safety asks whether it can be deployed with acceptable latency, access controls, observability, and rollback procedures.

Review layerQuestionEvidence to keep
NumericalDoes the output match the reference method?Fixed inputs, tolerance, and comparison results
SoftwareDoes the change work through supported paths?Unit, integration, build, and regression tests
OperationalCan the team run and reverse it safely?Logs, alerts, permissions, rollout and rollback records
ReviewCan another engineer understand the change?Diff notes, assumptions, design choice, and sign-off

What Claude Code Can Do According to Its Documentation

Claude Code documentation describes an agentic coding tool that reads a codebase, edits files, runs commands, and integrates with development tools. It also describes work across multiple files and tools. Those capabilities line up with the kind of repository work described by Rakuten.

The vLLM Claude Code integration documentation explains how a vLLM server can expose an Anthropic Messages API backend. That is an integration path, not proof that every model or repository will reproduce the Rakuten result.

The documentation does not turn the tool into an independent software engineer with guaranteed judgment. Reading a codebase is not the same as understanding every architectural tradeoff. Running a command is not the same as choosing the right acceptance criterion. Editing files is not the same as owning the release decision.

This distinction is especially important when a team describes a run as agent-run. In the Rakuten account, the agent worked for 7 hours while Naruse provided occasional guidance. That is meaningful delegation. It still implies a human-defined task and a human-controlled environment.

For a technical comparison with other development agents, see our AI coding agents guide. The right choice depends on repository access, tool permissions, model behavior, review controls, and the type of work. A single case study cannot settle that comparison.

Why Human Guidance Still Matters

Human guidance is not just a fallback for a weak model. It is part of the control system. A developer chooses the problem, supplies constraints, points the agent toward the right reference, and decides which commands are safe to run. In a large repository, those choices can determine whether the agent spends its time on the right files.

Review also has a different role from prompting. A prompt can explain what the team wants. Review checks what the agent actually changed. That check should include the diff, generated files, test results, dependency changes, credentials exposure, and any new network or filesystem behavior.

The most useful workflow is neither full manual typing nor blind delegation. It is a loop. The engineer states the task, the agent proposes and implements a change, automated checks test the result, and a human reviews the evidence before merge. If the evidence is weak, the task returns to the loop instead of moving directly to production.

What This Case Does Not Prove

The Rakuten account does not prove that Claude Code can safely change every large codebase. It does not prove that every 12.5 million line repository will produce a similar result. It does not prove that a 7-hour run replaces the full cost of engineering work. And it does not prove that an agent can be left with unrestricted access to a production environment.

The account also does not prove that a 79% reduction is portable across companies. Delivery time can change because of release scheduling, staffing, review queues, test automation, project selection, or a change in what teams count as a feature. The source supports the Rakuten result as reported. It does not support a general percentage for the software industry.

Finally, the case does not settle the question of jobs. An agent can change task allocation without making a simple prediction about employment. Teams may ship more, review more, change roles, or reduce some kinds of work. The direction depends on management choices, demand, quality requirements, and the skills that remain scarce.

How to Evaluate a Similar Coding Agent Project

A team that wants to test a Claude Code workflow should begin with a bounded project. Pick a task with a clear reference implementation or acceptance test. Define the repository commit, allowed tools, sensitive paths, network permissions, and reviewer. Keep the agent’s access narrow until the team understands its behavior.

Measure more than elapsed agent time. Capture the human setup time, the time spent reviewing, the number of retries, the number of changed files, test coverage, defects found after review, and the final release time. If the goal is a numerical method, preserve the reference inputs and the comparison output. If the goal is a product feature, preserve the acceptance tests and user-facing behavior.

The result should be compared with the team’s normal process. One successful run is a useful signal. A decision needs repeated measurements across tasks that resemble the work the team actually does. The comparison also needs a quality threshold. Faster delivery with hidden regressions is not a productivity gain.

Test stageControlPass condition
ScopeFixed task, repository state, and permissionsThe agent cannot silently expand the assignment
ImplementationRecorded commands, changes, and guidanceThe team can reconstruct what happened
VerificationReference comparison and software testsOutputs meet the stated acceptance criteria
ReleaseHuman sign-off and rollback planThe change can be reversed if monitoring finds a problem

Our MCP server security checklist is relevant when coding tools connect to external services. The same principle applies here. Tool access should be treated as a production permission, not as a harmless convenience.

Enterprise Security and Review Controls

Agentic coding changes the security boundary because the tool can read files, edit files, and run commands. A team should decide which repositories, secrets, network endpoints, and deployment commands are in scope before the first run. Permission prompts are useful, but they do not replace repository policy or access review.

Logs matter too. Keep the task description, tool actions, changed files, test results, reviewer comments, and final commit together. If a defect appears later, the team should be able to identify what the agent saw and what a human approved. This is especially important for code that handles model weights, user data, credentials, or external requests.

The Anthropic report landing page frames agentic coding as a shift toward orchestrating agents while balancing productivity against oversight, quality, and security. That framing is more defensible than a replacement story. The Rakuten result shows what a well-scoped delegation can look like. It does not remove the need for governance.

For a broader discussion of workflow controls, see our AI coding agent cost analysis and compare its cost assumptions with the actual review burden in your repository.

What This Means for Software Teams in India and Asia

For software teams in India, Japan, and other Asian markets, the practical lesson is not simply to reduce typing. The work is to shorten the path from a well-defined request to a tested change while keeping architecture and release responsibility with engineers.

That opportunity comes with a practical warning. A large legacy system may have incomplete tests, undocumented interfaces, old build rules, and sensitive customer data. An agent can help search and modify such a system, but the team still needs local knowledge. The best early project is one where the expected behavior can be checked without exposing critical production access.

Teams should also measure whether the bottleneck has moved. If implementation gets faster but review queues grow, the release process may not improve. If agents create more parallel changes, reviewers need better ownership, test automation, and a clear way to reject weak work. The human role may shift, but it does not disappear.

Bottom Line on the Rakuten Case Study

The Claude Code Rakuten case study is a valuable primary-source example because it names a specific task, a specific repository, a reported 7-hour agent run, and a 99.9% numerical comparison. Rakuten also reports a 24-day to 5-day change in feature time to market. Those facts are worth testing against the needs of a real engineering team.

The right conclusion is narrower than the old article’s promise. Claude Code can read codebases, edit files, run commands, and support an agentic workflow. A capable team can delegate difficult work and review the result. But one case does not establish a universal 79% improvement, guaranteed accuracy, or a replacement for engineering judgment.

Start with a bounded task. Keep the permissions narrow. Preserve the reference method and tests. Record both agent time and human review time. Then compare the result with the team’s normal process. That is how the Rakuten example becomes useful evidence rather than another unsupported productivity slogan.

Frequently Asked Questions

Rakuten says Kenta Naruse asked Claude Code to implement an activation vector extraction method inside vLLM, an open-source inference library. The case study describes this as a difficult implementation in a large codebase.
Rakuten describes the vLLM codebase as having 12.5 million lines across multiple programming languages. That figure describes repository scale and does not mean the agent held every line in active context at once.
Rakuten reports 7 hours of sustained agent-run coding in a single run. The case study also reports that Naruse provided occasional guidance while the agent worked.
Rakuten says the implementation reached 99.9% numerical accuracy compared with the reference method. That is a task-specific numerical comparison, not a guarantee of universal software correctness or production safety.
Rakuten reports that time to market for new features fell from 24 days to 5 days and presents this as a 79% reduction. The figures should be read as a company-reported case comparison.
No. It is one company account about one implementation and one delivery comparison. Results can change with task size, repository quality, testing, review time, permissions, staffing, and release processes.
Use a bounded task with a reference method or clear acceptance tests. Fix the repository state and permissions, record agent and review time, preserve commands and diffs, run software and quality checks, and require human sign-off before release.
SK Jabedul Haque
Written by

SK Jabedul Haque

Founder & Chief Editor

Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.

Read full bio

Never miss an update

Get our clearest explainers on schemes, markets and money — read what matters, without the noise.

Explore more articles
In this article