📊 Key Data
  • 32-fold efficiency gain: SMITH framework reduces token output from 3,206 to 100, cutting compute costs by 97%. - Model size advantage: 4-billion-parameter model outperforms 30-billion-parameter baseline in tool creation. - Tool reuse: AI agents share a curated library of high-performing tools across multiple tasks.
🎯 Expert Consensus

Experts would likely conclude that Appier's SMITH framework represents a significant advancement in making Agentic AI economically viable for enterprise deployment by drastically reducing compute costs through self-tooling mechanisms.

about 11 hours ago
Slashing the AI Compute Bill: How Self-Tooling Agents Are Rewriting the Rules

Slashing the AI Compute Bill: How Self-Tooling Agents Are Rewriting the Rules

SINGAPORE – September 30, 2026 — For the past three years, the enterprise technology sector has been chasing the promise of Agentic AI. The vision was intoxicating: autonomous digital workers capable of managing complex workflows, from supply chain logistics to dynamic marketing campaigns, with minimal human oversight. Yet, as chief technology officers and enterprise architects quickly discovered, this autonomy came with a crippling price tag. Recursive reasoning—the step-by-step logic chains required for an AI to solve complex, multi-layered problems—consumes a staggering amount of compute power.

Today, a significant breakthrough in mitigating these prohibitive costs has emerged from an unexpected source. Appier, a Tokyo-listed company traditionally known for its AI-native marketing and advertising technology, announced that its latest research paper has been accepted at NeurIPS 2026, the world's premier artificial intelligence conference. The paper introduces a reinforcement learning framework dubbed SMITH (Schema-grounded Multi-task Iterative Tool Honing). By allowing AI agents to build, test, and refine their own software tools within a single training loop, SMITH effectively transforms expensive, repetitive reasoning into lightweight, reusable functions. The implications for enterprise AI deployment are profound.

Slashing the LLM Inference Bill

The most pressing hurdle in scaling Agentic AI has always been the economic reality of the Large Language Model (LLM) inference bill. When an AI agent encounters a complex problem, it typically relies on "chain-of-thought" reasoning. It thinks through the problem step-by-step, generating thousands of tokens in the process. When deployed across millions of daily enterprise operations—such as converting financial metrics, processing customer data, or querying disparate databases—these token costs spiral out of control.

Appier's SMITH framework tackles this inefficiency head-on by turning repeated reasoning into a reusable tool that can be called directly. According to the research findings, this method reduced the average token output from a bloated 3,206 tokens using conventional step-by-step reasoning down to a mere 100 tokens.

This represents an approximate 32-fold gain in efficiency. For an enterprise software architect, a 97 percent reduction in token usage is not merely an incremental improvement; it is the difference between an AI project being a costly experimental pilot and a viable, scalable production system. When facing similar problems, the AI no longer needs to run the full, expensive reasoning process every single time. It simply reaches into its library, pulls out the proven tool it previously engineered, and executes the task, drastically lowering compute costs while maintaining performance.

Beyond Pre-Configured APIs

The technical shift presented in the NeurIPS paper moves the industry away from a brittle reliance on human-engineered APIs. Historically, AI systems have depended on human developers to build APIs or configure fixed tools in advance. Whenever business needs, data sources, or tasks changed, these tools had to be manually rebuilt. Even newer methods that allowed AI to create its own tools typically assigned the creation of the tool and the use of the tool to entirely separate models.

This separation created a disconnect. The tool-building model received little to no feedback on how its creation performed in the real world. It could not easily determine if its tool was clearly described, if its parameters were designed correctly, or if other models could reliably call upon it.

SMITH fundamentally alters this dynamic by putting both skills into a single, closed training loop. The AI learns to build tools and use them simultaneously. If a tool fails to run or has a vague description, that failure feeds directly back into the model.

"When SMITH trains a model to use tools, it sees only the tool's description and parameter specifications, not the underlying code," noted Chieh-Yen Lin, Research Scientist at Appier. "This makes the clarity of each description, and whether the tool can be called correctly, direct feedback during training. Our experiments also confirmed that other models can use these tools to solve problems more effectively. Looking ahead, we hope to build models that can continuously interact with their environment and take on a wider range of tasks."

Remarkably, the research demonstrates that effective tool creation does not require massive, resource-heavy models. A model of approximately 4 billion parameters trained with SMITH successfully built tools that outperformed those generated by an on-the-fly 30-billion-parameter baseline model on unseen tasks. Furthermore, the tools built by the smaller model generalized so well that they could be utilized by a highly lightweight model of only 350 million parameters.

From MarTech Vendor to Deep-Tech Contender

Appier's presence at NeurIPS—often dubbed the "Olympics of AI"—signals a strategic evolution for the company. Founded in 2012, Appier built its reputation by delivering Agentic AI as a Service (AaaS) for AdTech and MarTech solutions. However, gaining acceptance at a tier-one machine learning conference validates the company's deep-tech credibility well beyond standard commercial marketing claims.

This transition from a regional advertising platform to a recognized contributor in global AI research highlights a broader industry trend. The companies that will dominate the next decade of enterprise software are those actively solving the foundational bottlenecks of machine learning, rather than just wrapping existing foundational models in a new user interface.

"Humans turn their problem-solving experience into tools, so they never have to start from scratch. AI agents are now evolving in the same way," stated Dr. Chih-Han Yu, CEO and Co-founder of Appier. "This research shows that agents can learn to build tools, continuously refine them, and share proven tools across models of all sizes, making multi-agent collaboration more efficient and scalable. NeurIPS's acceptance of this paper further recognizes Appier's forward-looking research and innovation. We will continue to bring Agentic AI into real-world applications and deliver measurable results for businesses."

The Enterprise Reality Check

The true test of the SMITH framework will be its integration into commercial environments. The research curriculum demonstrated that the AI could learn a method from just four simple examples and successfully generalize that knowledge to construct tools for sixteen harder, previously unseen problems. The system then curates a shared tool library, preserving high-performing tools and discarding weaker ones, which multiple AI agents can access.

In sectors like advertising and marketing, where Appier already operates, this capability could be transformative. Agents handling distinct but related tasks—such as customer data ingestion, personalization formatting, customer service routing, and programmatic ad buying—can now share a consistent, evolving library of proven tools. Whether a business is entering a new demographic market or onboarding a client with limited historical data, the AI can leverage verifiable, scalable capabilities rather than starting from zero.

As the industry pushes toward increasingly complex multi-agent systems, the ability to democratize tool creation across smaller, more efficient models will be paramount. By successfully merging tool creation and execution into a single, self-refining loop, Appier has provided a compelling blueprint for the future of scalable enterprise AI. The days of the bloated LLM inference bill may finally be numbered, replaced by a new era of highly specialized, self-tooling digital workforces.

Topics & Related

Sector:
AI & Machine Learning
Theme:
Agentic AI
Large Language Models
Event:
Scientific Publication

📝 This article is still being updated

Are you a relevant expert who could contribute your opinion or insights to this article? We'd love to hear from you. We will give you full credit for your contribution.

Contribute Your Expertise →
UAID: 51188