Applied AI research lab · Agents, models, machines

Software that thinks in teams.

KResLab builds AI that works as a team: autonomous agents, custom models and connected devices that plan, act and improve together. We pick the right model for every job, from frontier LLMs to tiny models on the edge, and move your software from waiting for clicks to pursuing intent.

Agentic flowsAI agentsGenerative AIComputer visionPredictive MLFine-tuningAI for IoTEdge inferenceDigital twinsVoice agentsMLOpsEvals & red-teaming

Thesis

From tool calls to intent.

Most AI in production is a function call with good manners. A prompt goes in, text comes out, and a person stitches the results together. That person becomes the bottleneck. We build systems where agents do the stitching: they understand the goal, split it into work, check each other, and come back to a human only when a decision really needs one.

  1. L0

    Scripted

    Fixed rules. Breaks on the first input nobody planned for.

  2. L1

    Assisted

    A model answers when asked. A person does everything else.

  3. L2

    Tool-using

    One agent calls tools in a loop to finish a single task.

  4. L3

    Orchestrated

    A planner hands work to specialists and merges what they return.

  5. L4

    Mesh

    Peer agents find each other by capability, negotiate and share memory without a central script.

  6. L5

    Adaptive

    The mesh scores its own runs and rewrites its playbooks, inside guardrails you set.

Architecture

Any model. Five tiers.

Every system we ship follows the same stack. A model layer sits in the middle, routed by difficulty: frontier LLMs plan, fast models execute, and small or specialist models (vision, forecasting, speech) handle the volume. We stay model-agnostic, so you can change providers without rebuilding. Above the models, a mesh of agents divides the work. Below them, memory and tools connect everything to your data and your machines.

Tracing runs through every tier. Each hop between agents is recorded, replayable and scored, which is how the mesh improves without drifting.

  • ModelsCommercial APIs, open-weight, or trained by us
  • RoutingFrontier plans · fast executes · small triages
  • ProtocolModel Context Protocol for every tool
  • DeploymentYour cloud, on-prem, or edge gateway
  • Trace retentionEvery decision, 400 days by default
T5

Surfaces

Web and mobile apps, chat, voice lines, dashboards, physical devices.

intent enters
Human gateIrreversible actions wait for a signature
T4

Agent mesh

Planners, specialists, critics and sentinels exchanging messages over a shared blackboard.

work is divided
T3

Model layer · provider-agnostic

Reasoning, perception and prediction, routed by difficulty and cost.

Frontier LLM · planFast LLM · executeVision · perceiveForecast · predictSmall · triage
decisions are made
T2

Memory & tools

Vector and graph memory, MCP servers, a tool registry with scoped credentials.

context is kept
T1

Substrate

Cloud regions, on-prem clusters, edge gateways, sensor and actuator fleets.

things happen

Practice areas

The periodic table of what we build.

Sixteen elements, six groups, covering the full AI stack from data to devices. Most engagements combine three or four. Select an element to read what it covers.

01 Orchestration
Af

Agentic Flows

Multi-step business processes rebuilt as agent pipelines with branching, retries and checkpoints. Claims intake, procurement, onboarding, month-end close.

Typical bench time
3 weeks
Instruments
Lattice, Petri

Instruments

Built in the lab, used on every engagement.

Lattice

Mesh runtime and orchestration

Declare agents, their models, tools and limits in one file. Lattice handles discovery, message routing, retries, spend caps and the human gate. Every hop lands in a trace you can replay.

  • Agents per mesh2 to 400
  • Handoff overheadunder 50 ms
  • Runs onKubernetes, VMs, edge gateways
cold-chain.mesh.yamlLattice
mesh: cold-chain-ops
router: by-difficulty    # any provider
agents:
  - id: sentinel
    model: small-llm    # 40k sensor events / min
    watches: [reefer.temp, door.state]
  - id: diagnostician
    model: fast-llm
    tools: [mcp://fleet-telemetry, mcp://maintenance-db]
  - id: inspector
    model: vision/cargo-v3  # trained in-house
    watches: [cargo.cam]
  - id: planner
    model: frontier-llm
    escalates_to: human.dispatch
policies:
  irreversible_actions: require_approval
  max_spend_per_run: 4.00 USD
  trace: every_hop

Petri

Agent sandbox and eval bench

Grow a mesh in an isolated culture, replay real history through it, and score the result before anything touches production. Petri runs on every change, so a regression shows up as a number, not a complaint.

  • Scenario replayhistorical, synthetic, adversarial
  • Scored onaccuracy, policy, cost, latency
  • Promotionautomatic when the suite passes
terminalPetri
$ petri run suites/claims-triage --seeds 500
▸ spawning 6 agents in isolated culture
▸ replaying 500 historical claims
  routing accuracy        96.4%   +2.1 vs main
  policy violations       0
  human escalations       11.8%
  p95 time to decision    38 s
  cost per claim          $0.031
✓ suite passed · promoted build 2026.10.04-r3

Synapse Edge

Agent runtime for IoT gateways

Small local models sit on the gateway, read the sensor stream and throw away the noise. Only events that need real reasoning go up to the mesh, which keeps bandwidth and API spend low and response times short.

  • HardwareARM64 and x86 gateways, 4 GB RAM
  • Offline modelocal rules until the link returns
  • Escalation ratetypically 1 in 2,000 events
event log · reefer fleetSynapse Edge
14:02:11  reefer-17   temp 6.8 °C ↑   limit 5.0
14:02:11  edge        anomaly score 0.91 → escalate
14:02:13  diag        compressor duty 100%, door shut 3 h
14:02:14  planner     reroute to depot B, +22 min
14:02:14  gate        waiting for dispatcher approval
14:02:40  gate        approved by dispatcher-02
14:02:41  actuator    route pushed, driver notified

Protocol

How an engagement runs.

Four phases, each with a result you can judge before paying for the next one.

  1. 011 week

    Hypothesis

    We map the workflow, find where people act as glue between systems, and write down what a successful agent would measurably change.

  2. 022–3 weeks

    Bench

    A working prototype in Petri against your anonymized data. You see scores, not slides.

  3. 034–6 weeks

    Trial

    A pilot on live traffic behind human gates. Every decision traced, every failure replayed and fixed.

  4. 04ongoing

    Production

    We ship to your infrastructure, train your team on Lattice, and keep improving the mesh on a monthly cycle.

Experiment log

Recent entries from the notebook.

Client names are withheld. Numbers are measured against the client's own baseline over the trial period.

Entry Setup Elements Result
EXP-0217Cold-chain logistics

Edge sentinels on 140 refrigerated trucks feed an LLM diagnostician and a route planner. Dispatchers approve reroutes from their phones.

IoEdMx −31%spoiled loads−74%alerts needing a person
EXP-0198Insurance claims

A six-agent intake crew reads documents, checks policy terms, flags fraud signals and drafts the first decision for an adjuster.

AfEvGr 38 sto first decision, was 2 days0policy violations in 500-claim bench
EXP-0183Hospitality venues

Voice agents take bookings and in-room requests across 30 venues, while a demand forecast sets staffing and stock for each night.

VxPrAg 62%of calls resolved without staff0.8 smedian time to answer
EXP-0171Light manufacturing

A custom vision model inspects welds on the line while planning agents rehearse shift changes in a digital twin before a supervisor signs off.

CvDtMx 99.2%defect recall at line speed−18%changeover time

Lab rules

Autonomy is earned, one eval at a time.

  1. No agent without an eval.

    If we can't score it, we can't ship it. Every agent arrives with its own test suite.

  2. No action without a trace.

    Every message, tool call and decision is recorded and can be replayed step by step.

  3. Irreversible means a person signs.

    Payments, deletions and physical actuation stop at a human gate until you decide otherwise.

  4. Smallest model that passes.

    We route each step to the cheapest model that clears the bench, and reserve heavy reasoning for planning.

  5. Your data stays yours.

    Meshes run in your environment. Traces and memory never leave it without your say.

Open a lab session

Bring us a problem that needs more than one mind.

Send the workflow, who touches it today, and what done looks like. We reply within two working days with a first hypothesis.

info@kreslab.dev Email the lab