2026-08-20

Daily AI Digest: Agents trade, prevent IT tickets and answer site questions

This is the GolemWorkers daily AI digest for people who want to know what practical agents can now do. Today’s edition follows agents moving money, building automations, answering construction questions and proving which skills actually help—then gives you a trusted-source monitoring recipe to use yourself.

An operations lead places a human approval control at the center of finance, IT, construction, and software work.
An IT administrator reviews a physical automation runbook while a technician traces a network fault.

Serval’s Catalyst finds IT work and drafts automations

Serval made Catalyst generally available and enabled it by default for customers. IT and operations administrators can give it ticket history, a written procedure or plain-language instructions; it then proposes workflows, access policies, forms and background checks as drafts for human review before they go live.

Source: Serval
A trader places a limited set of tokens into a separate lockbox while keeping the larger reserve secured.

Binance opens trading and payments to authorized AI agents

Binance launched Agent OS so developers can connect agents to market data, account information, trading and payments. Users can confine an agent to a sub-account, block withdrawals and require approval for every order or permit autonomous trades; Binance does not add a separate loss cap, so the funded sub-account remains the practical limit.

Source: TechCrunch
A site manager and project engineer compare a marked request with construction plans while work continues nearby.

PlanRadar agents answer site questions from project records

PlanRadar added agents that can complete routine construction work inside its project platform. Its Response agent reads an incoming request for information, searches the project’s documents and prepares a sourced reply; teams can also describe their own agents in plain language, test them before activation and audit every action afterward.

Source: AEC Magazine
Two engineers compare modular task paths and resource counters on a software-lab workbench.

TrueForge opens its agent loop for local and team use

TrueFoundry released TrueForge under the MIT License for teams that want to run multi-step agents with their choice of model and tools. Developers can start locally with SQLite, then move the same harness to shared infrastructure; the company says context controls and on-demand tool loading cut costs in its limited benchmark, a claim teams can test in their own workloads.

Source: TrueFoundry
A lab team evaluates two identical task runs, one with an organized instruction binder and one without.

NVIDIA tests whether agent skills improve real task runs

NVIDIA released SkillEvaluator as an open-source way to compare an agent running with and without a packaged skill under the same prompt, model and grading criteria. Teams using Codex, Claude Code, Cursor and other harnesses can run isolated checks for safety, duplication and measurable performance lift before trusting a skill in production.

Source: NVIDIA Developer

Agent idea of the day

Run a trusted-source change monitor every morning

An operations lead reviews verified source changes while an assistant files current notices and discards stale material.

What this agent does

Monitor a short approved list of official sources and deliver only new, decision-relevant changes from a defined time window.

Best for: Founders, operations leads, buyers and compliance owners who need fresh vendor, policy, pricing or supply updates without an open-ended web search.

Give it

  • Five to twenty trusted domains and the exact pages or topics that matter
  • A freshness window, such as the previous 24 hours
  • Rules for what counts as a material change and what should be ignored
  • A destination for the brief and a named reviewer

Tell it to

  1. Search only the approved domains and reject any result whose publication date cannot be verified inside the requested window.
  2. Open each candidate source, compare it with the last accepted version and discard reposts, commentary and unchanged pages.
  3. Extract the concrete change, effective date, affected product or process and the action the reviewer may need to take.
  4. Group related updates, attach the canonical links and label any missing or conflicting evidence instead of guessing.
  5. Deliver a short morning brief and wait for human approval before changing purchases, policies, accounts or production systems.

Run it: Run every weekday morning and on demand before a purchase, renewal, launch or compliance review.

You get

A dated brief of verified material changes, each with its canonical source, effective date, affected owner and recommended next action.

Keep a human in control

  • Never widen the domain allowlist or date window without the owner’s approval.
  • Treat undated, inaccessible or conflicting material as unresolved and keep it out of the accepted-change list.
  • Do not log in, accept terms, place orders or modify external systems during monitoring.
  • Keep the prior accepted brief so every reported change can be audited against the last known state.

Feasibility: Amazon Bedrock AgentCore Web Search now supports per-request domain include/exclude lists and publication-date bounds, enforced server-side. Runtime filters can narrow an administrator’s policy but cannot expand it, and unverifiable results are excluded when filters are active. Source: AWS Machine Learning Blog →

GolemWorkers — OpenClaw ready to work: launch in seconds, no setup, online 24/7.