Connect with us

NEWS

Five Production Skills Now Filter Data Science Jobs

Retrieval, routing, guardrails, and evals now decide who ships LLM work, while agent loops are the skill most likely to waste a year.

Published

on

Gartner’s 2026 CIO survey found that only 17% of organizations have deployed AI agents, while more than 60% expect to within two years.

Anyone can still stand up an LLM demo with one API call. Getting that feature through real users, messy company data, and a bill a finance team will accept now depends on five production LLM skills, and that is the work data scientists are being hired to do.

Gartner’s 17 Percent Is the Job Filter

The same research firm already put a date on the wreckage. On June 25, 2025, Gartner said over 40% of agentic AI projects will be canceled by the end of 2027, blaming escalating costs, unclear business value, or weak risk controls. Model quality did not make that list.

THE ADOPTION GAP IN ONE LOOK

  • In production: Only 17% of organizations have deployed AI agents, according to Gartner’s 2026 CIO and Technology Executive Survey.
  • Still planning: More than 60% expect to deploy them within two years, the steepest adoption curve in that survey.
  • On the chopping block: Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027.
  • Where the money sits: A January 2025 Gartner poll of 3,412 webinar attendees found 19% with significant agentic bets, 42% with conservative bets, 8% with none, and 31% waiting or unsure.

Anushree Verma, a senior director analyst at Gartner, described the pile-up in the June 2025 note.

Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.

Anushree Verma, Senior Director Analyst, Gartner newsroom, June 25, 2025

She added that the hype blinds teams to the real cost and complexity of deploying agents at scale, which is how projects stall before production. Gartner still forecasts that by 2028, 15% of day-to-day work decisions will be made by agents, up from 0% in 2024, and that 33% of enterprise software will include agentic AI, up from less than 1% in 2024. Those 2028 figures are a destination. The 17% is the road.

Five Skills Answer What a Frontier Model Cannot

Sara Nóbrega, an AI engineer who deploys machine learning systems, recently grouped the missing work into five skills: retrieval, routing, guardrails, evals, and agent loops. Each one answers a question a frontier model will not answer on its own.

THE FIVE QUESTIONS A BASE MODEL CANNOT CLOSE

Skill Question it answers What fails without it
Retrieval (RAG) Where does company data come from? Answers that ignore internal docs
Routing Why is the bill this high? Every request hits a frontier model
Guardrails What happens on a hostile prompt? Injection and leaked personal data
Evals Did the last change help? Prompt edits shipped on feel
Agent loops How does multi-step work recover? One empty tool call restarts the task

Hiring pages already read like that table. Agoda’s staff LLM data scientist role asks for production RAG, tool orchestration, and evaluation. Socure’s staff role asks for offline benchmarks, safety checks, and guardrails on LLM tools. Gartner’s own data scientist listing in Gurgaon asks candidates to own evaluation, guardrails, and hallucination control on RAG and agent workflows. The title still says data scientist. The work is production LLM plumbing.

Retrieval Still Decides What the Model Knows

A base model does not know your ticket history, your policy PDF, or last week’s pricing sheet. Fine-tuning every time those files change is slow and expensive. Retrieval-augmented generation fetches the relevant chunks at query time and stuffs them into context, which is why RAG is still the first pattern most live systems use.

The interesting work is the fetch, not the generate step. Chunk size, hybrid search that mixes keywords with embeddings, and a reranker that reorders the top hits decide whether the model sees the right paragraph. A tiny CPU demo makes the failure obvious: a user asks the system to “send something back,” the matching doc says “returns,” keyword search misses it, and the embedding match catches it because the two phrases sit close in vector space.

Break your own notes on purpose. Ask a question the retriever gets wrong and watch the answer rot. Change chunk size. Add BM25 next to the vectors. Compare the top 3 hits before and after a cross-encoder rerank. When a numpy array runs out of room, FAISS is the speed move, and Qdrant or Pinecone add filtering. LlamaIndex and LangChain will store and search for you. RAGAS then scores faithfulness, answer relevancy, context precision, and context recall, so you stop eyeballing a chatbot and start measuring retrieval.

That is the part notebooks skip. A pretty demo with one PDF is not a system that survives a synonym, a table, and a policy update on the same afternoon.

Easy Requests Should Not Hit Frontier Models

Sending every prompt to a frontier model is the fastest way to a bill nobody can explain. Routing scores a request, sends the easy high-volume work to a small or local model, and keeps the expensive model for the hard cases. Caching repeated prompts sits next to that, as a separate lever on the tokens you already decided to send.

Tin Lo, founding AI product engineer at LiteLLM, wrote in August 2026 that teams can expect roughly 40% cost reductions from day one with the gateway’s Auto Router, with more after the tier maps are tuned. One production user shared the logs. From April 15 to August 9, 2026, the router ran 272,876 requests and 7.08 billion tokens for 450-plus people across dev, staging, and prod, and recorded $12,249 saved on flagship spend.

LITELLM AUTO ROUTER, THREE MODEL FAMILIES

Router Requests Spent Saved Savings
claude-auto-latest 109,675 $8,997 $7,951 46.9%
gemini-auto-latest 90,294 $1,327 $1,280 49.1%
gpt-auto-latest 72,906 $1,412 $3,019 68.1%

Actual spend across that window was $11,736 against a $23,985 flagship counterfactual, a published 51.1% cut. Blended cost landed at $1.66 per 1 million tokens against $3.39 at the flagship. 95% of requests never needed the top tier. The savings rate moved from 42.9% in May to 50.0% in June, 53.7% in July, and 60.7% in August through the 9th as the team retuned which prompts belonged in which tier. Users kept calling one model name. The router picked the tier.

Lo’s 40% day-one figure is the install number. The 51.1% is the four-month log. The 60.7% is what the same traffic did after the maps matured. A router only pays off if every call is logged with task type, tokens, and cost. RouteLLM is the research-shaped alternative when the decision itself is the product. LiteLLM is the gateway a lot of teams already put in front of many providers. Hand-rolling a cheap path for the single highest-volume simple task is still the right first week of work.

Excessive Agency Jumped to Third on the 2026 List

Guardrails moved from a bonus slide to an interview question because the model reads instructions and data on the same channel. An attacker can write input that the model treats as a new instruction and follows, because it cannot tell the two apart. Prompt injection was already LLM01 on the 2025 OWASP list. The 2026 Top 10 for LLM Applications, published August 4, 2026, kept it there.

HOW THE 2026 LLM TOP 10 MOVED

  • LLM01 Prompt injection: Still first, on the strength of the practitioner vote and the defense work already spent holding it off.
  • LLM02 Sensitive information disclosure: Still second, the top slot where the vote and the incident record agree.
  • LLM03 Excessive agency: Up from sixth, the sharpest move, because agentic deployments are where the damage is landing.
  • LLM06 Unbounded consumption: Up four places, as teams weigh runaway bills and resource exhaustion more heavily.
  • LLM10 Improper output handling: Down from fifth to tenth, the longest fall on the list.

The 2026 edition mixed a 75% practitioner vote with a 25% read of thousands of real incidents. Excessive agency is the entry that should change how a data scientist scopes an agent. Too many tools, too many permissions, or too much autonomy turns a jailbreak into an action outside the chat window. Regex is a first screen, not a defense. OWASP’s own guidance is stacked: input checks, output filters, least-privilege tools, human approval for sensitive actions, and regular adversarial tests. Production stacks reach for Microsoft Presidio on personal data, and for Guardrails AI or NeMo Guardrails on validation.

The interview version is blunt. Red-team your own system prompt. Name the injection paths. Write the attack and the patch from memory. If you cannot, you do not own the guardrail.

A Tiny Eval Set Beats Prompt Guesswork

Change a prompt with no eval and you are guessing. When a multi-step agent fails with no trace, you see the wrong final answer and none of the steps that produced it. Start with 20 to 50 real inputs and the answers you would accept. Even 8 examples shows the shape of the failures. The first time a prompt change holds a score, the work stops feeling like paperwork.

Exact match dies on open-ended answers. The common fix is a second model that grades the first. DeepEval’s docs push a binary pass or fail, including a strict mode that collapses a score to 1 or 0, because a 0-to-100 judge is a number you cannot defend. LLM judges also lean toward longer answers, the first answer they see, and a passing grade. Label about 20 outputs yourself and measure how often the judge agrees before you trust the dashboard.

DeepEval is pytest-native, MIT-licensed, and runs locally with no account, which is why it drops into CI for a solo engineer. promptfoo is YAML-configured and strong at side-by-side model compares and red-teaming. RAGAS stays the RAG-specific kit. For live traces, Langfuse, LangSmith, Arize Phoenix, and Braintrust sit on top of a shared schema. OpenTelemetry graduated from the CNCF in May 2026, with more than 12,000 contributions from over 2,800 companies, and the project’s generative AI semantic conventions are the long-term shape for those traces, including spans for model calls and agent steps. The conventions are still marked development. Graduation is why teams are willing to bet on the schema anyway.

Nóbrega would start here if nothing else is on fire. Building a test set is dull, and it is what makes senior engineers trust what you ship.

Agent Loops Carry the Cancellation Rate

This is the skill with the loudest tutorials and the worst completion rate. An agent is a model plus tools plus a loop that keeps calling tools until the task is done. Build that loop by hand once. The model will return prose instead of JSON. A tool will come back empty. The loop will spin. That fragility is why LangGraph and CrewAI exist: state machines, retries, and checkpoints so a run resumes from the failed step instead of restarting the whole job.

THE DATES BEHIND THE AGENT CAUTION

  1. January 2025: Gartner polls 3,412 webinar attendees and finds 19% with significant agentic investment and 31% still waiting or unsure.
  2. June 25, 2025: Gartner publishes the forecast that over 40% of agentic AI projects will be canceled by the end of 2027.
  3. April 15, 2026: The 2026 Hype Cycle for Agentic AI puts the category at the Peak of Inflated Expectations, with 17% deployed and more than 60% planning to follow.
  4. May 2026: OpenTelemetry graduates at the CNCF, and agent-span conventions become the default way to see a failing loop.
  5. August 4, 2026: OWASP moves excessive agency to third on the LLM Top 10 because agentic deployments are where the damage is landing.

Force a failure and make the agent recover. Give it a real tool and handle junk output. Then read a postmortem of a canceled agent program. Cost and governance are the lessons no tutorial teaches, and they are the questions that keep you credible: does this task need an agent at all, or would a function call do? The hype is loud. The cancellation rate is the part of the forecast that is already coming true in slow motion.

Architecture Calls Go to People With Proof

A working demo is cheap now. Proof that it deserves production is the scarce part. The people who can ground a model in company data, cut a large share off the bill, stop a jailbreak, and show a score before and after a prompt change are the people who get to say a task should run local, or should not be an agent at all.

You do not need all five skills in one quarter. Trying to learn them in parallel is how people burn out and keep none of them. Pick the hole in the system you already run. Retrieval if it cannot see your files. Routing if the bill is a mystery. Guardrails if you have never tried to break your own prompt. Evals if you are still editing on feel. Agent loops last, and only after you are willing to conclude that a plain function call would have done the job.

Gartner’s cancellation window runs through 2027. The 17% who already shipped agents did not get there with a bigger model. They got there with retrieval that fails on purpose, a router that logs every dollar, a guardrail you can defend, and an eval set boring enough that a senior engineer will trust it.

Harry is the editor of BUDGY APP, an independent title he owns and runs after ten years in journalism that began on a reporter's desk and ended up at the editor's. Numbers get particular attention here. A percentage in a business story is recomputed from the underlying figures before it goes live, a benchmark in a technology or gaming review is quoted with the conditions it was measured under, and a transfer fee or a lap time in the sports and auto pages is traced back to the club, the league or the timing sheet that published it. The same rule covers news, science, entertainment, lifestyle and travel: if a figure cannot be tied to a filing, a dataset, a transcript or a test Harry ran himself, it does not appear. Readers around the world see prices in the original currency with a conversion alongside. Errors are corrected in the open under a published corrections policy, with the change noted on the article. Questions about any figure reach him at support@budgyapp.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending