Agentic Flows
Multi-step business processes rebuilt as agent pipelines with branching, retries and checkpoints. Claims intake, procurement, onboarding, month-end close.
Applied AI research lab · Agents, models, machines
KResLab builds AI that works as a team: autonomous agents, custom models and connected devices that plan, act and improve together. We pick the right model for every job, from frontier LLMs to tiny models on the edge, and move your software from waiting for clicks to pursuing intent.
Thesis
Most AI in production is a function call with good manners. A prompt goes in, text comes out, and a person stitches the results together. That person becomes the bottleneck. We build systems where agents do the stitching: they understand the goal, split it into work, check each other, and come back to a human only when a decision really needs one.
Fixed rules. Breaks on the first input nobody planned for.
A model answers when asked. A person does everything else.
One agent calls tools in a loop to finish a single task.
A planner hands work to specialists and merges what they return.
Peer agents find each other by capability, negotiate and share memory without a central script.
The mesh scores its own runs and rewrites its playbooks, inside guardrails you set.
Architecture
Every system we ship follows the same stack. A model layer sits in the middle, routed by difficulty: frontier LLMs plan, fast models execute, and small or specialist models (vision, forecasting, speech) handle the volume. We stay model-agnostic, so you can change providers without rebuilding. Above the models, a mesh of agents divides the work. Below them, memory and tools connect everything to your data and your machines.
Tracing runs through every tier. Each hop between agents is recorded, replayable and scored, which is how the mesh improves without drifting.
Web and mobile apps, chat, voice lines, dashboards, physical devices.
Planners, specialists, critics and sentinels exchanging messages over a shared blackboard.
Reasoning, perception and prediction, routed by difficulty and cost.
Vector and graph memory, MCP servers, a tool registry with scoped credentials.
Cloud regions, on-prem clusters, edge gateways, sensor and actuator fleets.
Practice areas
Sixteen elements, six groups, covering the full AI stack from data to devices. Most engagements combine three or four. Select an element to read what it covers.
Multi-step business processes rebuilt as agent pipelines with branching, retries and checkpoints. Claims intake, procurement, onboarding, month-end close.
Instruments
Mesh runtime and orchestration
Declare agents, their models, tools and limits in one file. Lattice handles discovery, message routing, retries, spend caps and the human gate. Every hop lands in a trace you can replay.
mesh: cold-chain-ops
router: by-difficulty # any provider
agents:
- id: sentinel
model: small-llm # 40k sensor events / min
watches: [reefer.temp, door.state]
- id: diagnostician
model: fast-llm
tools: [mcp://fleet-telemetry, mcp://maintenance-db]
- id: inspector
model: vision/cargo-v3 # trained in-house
watches: [cargo.cam]
- id: planner
model: frontier-llm
escalates_to: human.dispatch
policies:
irreversible_actions: require_approval
max_spend_per_run: 4.00 USD
trace: every_hop
Agent sandbox and eval bench
Grow a mesh in an isolated culture, replay real history through it, and score the result before anything touches production. Petri runs on every change, so a regression shows up as a number, not a complaint.
$ petri run suites/claims-triage --seeds 500
▸ spawning 6 agents in isolated culture
▸ replaying 500 historical claims
routing accuracy 96.4% +2.1 vs main
policy violations 0
human escalations 11.8%
p95 time to decision 38 s
cost per claim $0.031
✓ suite passed · promoted build 2026.10.04-r3
Agent runtime for IoT gateways
Small local models sit on the gateway, read the sensor stream and throw away the noise. Only events that need real reasoning go up to the mesh, which keeps bandwidth and API spend low and response times short.
14:02:11 reefer-17 temp 6.8 °C ↑ limit 5.0
14:02:11 edge anomaly score 0.91 → escalate
14:02:13 diag compressor duty 100%, door shut 3 h
14:02:14 planner reroute to depot B, +22 min
14:02:14 gate waiting for dispatcher approval
14:02:40 gate approved by dispatcher-02
14:02:41 actuator route pushed, driver notified
Protocol
Four phases, each with a result you can judge before paying for the next one.
We map the workflow, find where people act as glue between systems, and write down what a successful agent would measurably change.
A working prototype in Petri against your anonymized data. You see scores, not slides.
A pilot on live traffic behind human gates. Every decision traced, every failure replayed and fixed.
We ship to your infrastructure, train your team on Lattice, and keep improving the mesh on a monthly cycle.
Experiment log
Client names are withheld. Numbers are measured against the client's own baseline over the trial period.
Edge sentinels on 140 refrigerated trucks feed an LLM diagnostician and a route planner. Dispatchers approve reroutes from their phones.
IoEdMx −31%spoiled loads−74%alerts needing a personA six-agent intake crew reads documents, checks policy terms, flags fraud signals and drafts the first decision for an adjuster.
AfEvGr 38 sto first decision, was 2 days0policy violations in 500-claim benchVoice agents take bookings and in-room requests across 30 venues, while a demand forecast sets staffing and stock for each night.
VxPrAg 62%of calls resolved without staff0.8 smedian time to answerA custom vision model inspects welds on the line while planning agents rehearse shift changes in a digital twin before a supervisor signs off.
CvDtMx 99.2%defect recall at line speed−18%changeover timeLab rules
If we can't score it, we can't ship it. Every agent arrives with its own test suite.
Every message, tool call and decision is recorded and can be replayed step by step.
Payments, deletions and physical actuation stop at a human gate until you decide otherwise.
We route each step to the cheapest model that clears the bench, and reserve heavy reasoning for planning.
Meshes run in your environment. Traces and memory never leave it without your say.
Open a lab session
Send the workflow, who touches it today, and what done looks like. We reply within two working days with a first hypothesis.