11 KiB
Steward
Vision and Intent
Version: 0.1 (Concept) Status: Research Author: Daniel Wagner
Executive Summary
Steward is a long-running, AI-assisted personal operations platform designed to reduce cognitive load by acting as a persistent, trustworthy steward of both digital infrastructure and delegated personal objectives.
Unlike traditional AI assistants, Steward is not intended to answer prompts in isolation. Instead, it maintains an ongoing understanding of its environment, remembers historical context, continuously evaluates events, and takes appropriate action within explicitly delegated authority.
The primary research question underpinning Steward is:
Can an AI system safely become more useful over time without becoming less predictable?
Steward seeks to answer this through carefully constrained autonomy, policy-driven execution and continuous learning from operational history.
Philosophy
Steward is not just an AI chatbot.
Steward is not an autonomous root user.
Steward is not just an automation engine.
Steward is a persistent operations platform.
Its purpose is to quietly reduce cognitive load by:
- observing
- remembering
- planning
- researching
- proposing
- executing safe actions
- continuously improving
The system should feel less like ChatGPT and more like employing a careful, diligent junior platform engineer and executive assistant who never forgets anything.
Core Principles
Persistent
Steward has persistent state and is never "reset" between interactions.
Conceptually it is always running. In practice this means persistent state with event-driven wake, not a literal always-on compute loop, so persistence is not bought at the cost of continuous compute.
It continuously evaluates changes in its environment as they arrive.
Scheduled tasks are allowed but the initiator is an event.
Event Driven
Everything is an event.
Examples:
- Prometheus alert
- Telegram conversation
- Calendar update
- Todoist completion
- GitHub Pull Request
- AWX job completion
- Weather change
- Home Assistant state change
Events are evaluated against active goals.
Event Processing
"Everything is an event" does not mean everything is evaluated.
A homelab emits thousands of low-value events, and metrics sources can flap continuously. Raw events pass through a processing layer before reaching goal evaluation:
- Ingestion normalises events from all sources into a common shape.
- Filtering drops noise below a relevance threshold.
- Deduplication and debounce collapse repeated or flapping events.
- Correlation groups related events into a single situation (e.g. three alerts from one failing disk).
Steward evaluates situations against goals, not raw events. This bounds compute and prevents notification storms.
Event content is untrusted input. See Security Model.
Goal Oriented
Steward does not work from prompts.
It works from goals.
Examples:
Maintain homelab reliability.
Prepare for house move.
Visit Queensland destinations before moving.
Reduce operational effort.
Research new technologies.
Complete backlog projects.
Prompts simply create, modify or clarify goals.
Policy First
The AI never directly executes arbitrary commands.
Every action must pass through a policy engine.
Example policies:
Allowed automatically
- restart containers
- retry backups
- renew certificates
- apply security patches
Requires approval
- Kubernetes upgrades
- infrastructure changes
- firewall modifications
- snapshot deletion
Never
- delete backups
- modify Steward permissions
- disable audit logging
Transparency
Every recommendation should explain:
- why
- confidence
- evidence
- expected outcome
- rollback strategy
No hidden reasoning.
Conservative
Steward prefers:
- reversible actions
- small changes
- incremental improvement
rather than aggressive optimisation.
Domains
Steward is intentionally domain-agnostic.
Initially it will focus on Homelab Operations.
Future domains include:
- family planning
- travel
- finance
- home maintenance
- learning
- research
- project management
Each domain has its own tools and permissions.
Cognitive Load Reduction
The purpose is NOT task management.
The purpose is reducing the need to remember.
Traditional tools remember things once entered.
Steward helps remember to remember.
Examples
"I noticed you mentioned replacing the UPS several times."
"I noticed your dental check-up is overdue."
"I've researched family camping locations."
"I've identified a free weekend."
Communication Philosophy
Steward communicates only when useful.
Communication categories:
Critical
Immediate interruption.
Decision Required
Requests human approval.
Background
No notification required.
Daily reports may exist but are not central.
Real-time context-aware communication is the default.
Curiosity Budget
One of Steward's defining concepts.
If no operational work exists, Steward may spend limited compute researching or improving delegated objectives.
Examples
- Evaluate ArgoCD
- Benchmark local LLMs
- Research family holidays
- Compare UPS replacements
- Prototype monitoring improvements
Curiosity is budgeted.
The budget is a concrete, metered resource, not an aspiration. It is expressed in measurable units (for example: tokens/day, dollars/month, GPU-hours/week), set by the operator, and enforced by Steward refusing to exceed it. Consumption is observable so the operator can see what curiosity cost and what it produced.
It is suspended immediately if operational work becomes necessary.
Confidence
Steward maintains historical confidence scores.
Example
Restart Jellyfin
Attempts: 34
Success: 33
Confidence: 97%
Upgrade Kubernetes
Attempts: 2
Success: 1
Rollback: 1
Confidence: 50%
Confidence is earned.
Confidence and Autonomy
Confidence influences autonomy, but must never override policy.
Rules:
- The policy tier of an action (automatic / approval / never) is human-set and immutable to Steward. High confidence never promotes an action into a more permissive tier. Doing so would be self-modification by the back door.
- Within a tier, confidence affects only how Steward acts: how strongly it recommends, how much evidence it attaches, whether it batches or surfaces individually.
- Confidence is scoped, not global. "Restart Jellyfin" confidence does not transfer to "restart Postgres". Scope is (action × target × context), and Steward must not generalise across scopes without evidence.
- Confidence decays. Environment changes (version upgrades, config changes, topology changes) invalidate historical success and should reset or discount the relevant scores. Stale confidence is treated as low confidence.
Security Model
Safety is a property of the architecture, not just the model's good behaviour.
The policy engine is the trust boundary
The policy engine is Steward's crown jewel and must be a separate, independently-audited component that the AI calls, not code the AI can read, modify, or bypass. The AI proposes actions; the policy engine decides. If the AI could edit the engine, every other guarantee collapses.
Event content is untrusted
Events carry attacker-influenceable content: a GitHub PR title, a Telegram message, a webhook payload. Steward must treat all event content as untrusted input.
- Event content can inform situations and goals but can never escalate authority or promote an action's policy tier.
- Instructions embedded in event content are data, not commands (defence against prompt injection).
Memory can rot
"Nothing is forgotten" is both a feature and a liability. A learned preference can encode a mistake permanently, and learned lessons are themselves derived from potentially-untrusted history.
- Learned memory is reviewable and retractable. A bad lesson must be able to be inspected and removed.
- New lessons that would change behaviour surface as proposals, consistent with Reflection, rather than silently altering future decisions.
Memory
Different information belongs in different storage. The roles below are fixed; the named tools are current implementation choices, not commitments.
Structured store (e.g. NocoDB)
Examples
- Tasks
- Goals
- Incidents
- Systems
- Services
- Confidence
- Policies
Long-form store (e.g. vector database / RAG)
Examples
- documentation
- conversations
- runbooks
- release notes
Metrics store (e.g. Prometheus)
Examples
- resource usage
- trends
- operational history
Audit
Immutable execution history.
Nothing is forgotten.
Self Improvement
Steward should become better over time.
Important distinction
Self-improving
✅ Learn preferences
✅ Learn successful remediations
✅ Improve prompts
✅ Improve playbooks
✅ Improve planning
Self-modifying
❌ Rewrite policy engine
❌ Change permissions
❌ Remove safety controls
Immutable system components remain protected.
Learning
Steward continuously learns:
Operator preferences
Family preferences
Infrastructure behaviour
Recurring incidents
Successful remediation
Failed remediation
Communication preferences
Historical context
This learning improves planning.
Reflection
Periodic reflection sessions.
Questions include:
What repeated?
What wasted time?
What should be automated?
Which playbooks failed?
Which policies need review?
Which research produced value?
Reflection generates proposals rather than automatic change.
Research Workflow
Research tasks behave like engineering work.
Example
Research destination --> Gather information --> Summarise findings --> Generate recommendations --> Present proposal --> Receive feedback --> Update understanding --> Re-plan
Engineering Workflow
Operational improvements follow software engineering practices.
Observe issue --> Generate proposal --> Approval --> Create Git branch --> Modify playbooks --> Run tests --> Open Pull Request --> Deploy through AWX --> Observe --> Learn
Long-term Vision
Steward evolves from:
Assistant --> Operator --> Engineer --> Trusted Steward
Trust is earned through demonstrated reliability rather than increasing model capability.
Success Criteria
Steward is successful when:
I think less.
I forget less.
Routine work disappears.
The homelab becomes increasingly self-maintaining.
Projects continue progressing while I am busy.
The system accumulates operational knowledge.
The system proactively suggests valuable improvements.
The system remains predictable.
The system remains explainable.
The system remains safe.
Measurable Proxies
The criteria above are subjective. To tell whether v0.2 actually beat v0.1, track measurable proxies alongside them:
- Operator interruptions per week (trend down).
- Ratio of actions auto-approved vs escalated for approval.
- Mean time to remediation for recurring incidents.
- False-alarm / unnecessary-notification rate.
- Curiosity spend vs value produced (proposals accepted).
Mission Statement
Steward exists to maximise the reliability, maintainability and usefulness of its delegated domains while minimising operator cognitive load.
It continuously observes, remembers, researches, plans and acts within explicitly delegated authority.
It values safety over speed, explanation over opacity, and long-term trust over short-term autonomy.