# Steward ## Vision and Intent Version: 0.1 (Concept) Status: Research Author: Daniel Wagner --- # Executive Summary Steward is a long-running, AI-assisted personal operations platform designed to reduce cognitive load by acting as a persistent, trustworthy steward of both digital infrastructure and delegated personal objectives. Unlike traditional AI assistants, Steward is not intended to answer prompts in isolation. Instead, it maintains an ongoing understanding of its environment, remembers historical context, continuously evaluates events, and takes appropriate action within explicitly delegated authority. The primary research question underpinning Steward is: > Can an AI system safely become more useful over time without becoming less predictable? Steward seeks to answer this through carefully constrained autonomy, policy-driven execution and continuous learning from operational history. --- # Philosophy Steward is not *just* an AI chatbot. Steward is not an autonomous root user. Steward is not *just* an automation engine. Steward is a persistent operations platform. Its purpose is to quietly reduce cognitive load by: - observing - remembering - planning - researching - proposing - executing safe actions - continuously improving The system should feel less like ChatGPT and more like employing a careful, diligent junior platform engineer and executive assistant who never forgets anything. --- # Core Principles ## Persistent Steward has persistent state and is never "reset" between interactions. Conceptually it is always running. In practice this means persistent state with event-driven wake, not a literal always-on compute loop, so persistence is not bought at the cost of continuous compute. It continuously evaluates changes in its environment as they arrive. Scheduled tasks are allowed but the initiator is an event. --- ## Event Driven Everything is an event. Examples: - Prometheus alert - Telegram conversation - Calendar update - Todoist completion - GitHub Pull Request - AWX job completion - Weather change - Home Assistant state change Events are evaluated against active goals. --- ## Event Processing "Everything is an event" does not mean everything is evaluated. A homelab emits thousands of low-value events, and metrics sources can flap continuously. Raw events pass through a processing layer before reaching goal evaluation: - **Ingestion** normalises events from all sources into a common shape. - **Filtering** drops noise below a relevance threshold. - **Deduplication and debounce** collapse repeated or flapping events. - **Correlation** groups related events into a single *situation* (e.g. three alerts from one failing disk). Steward evaluates *situations* against goals, not raw events. This bounds compute and prevents notification storms. Event content is untrusted input. See [Security Model](#security-model). --- ## Goal Oriented Steward does not work from prompts. It works from goals. Examples: Maintain homelab reliability. Prepare for house move. Visit Queensland destinations before moving. Reduce operational effort. Research new technologies. Complete backlog projects. Prompts simply create, modify or clarify goals. --- ## Policy First The AI never directly executes arbitrary commands. Every action must pass through a policy engine. Example policies: Allowed automatically - restart containers - retry backups - renew certificates - apply security patches Requires approval - Kubernetes upgrades - infrastructure changes - firewall modifications - snapshot deletion Never - delete backups - modify Steward permissions - disable audit logging --- ## Transparency Every recommendation should explain: - why - confidence - evidence - expected outcome - rollback strategy No hidden reasoning. --- ## Conservative Steward prefers: - reversible actions - small changes - incremental improvement rather than aggressive optimisation. --- # Domains Steward is intentionally domain-agnostic. Initially it will focus on Homelab Operations. Future domains include: - family planning - travel - finance - home maintenance - learning - research - project management Each domain has its own tools and permissions. --- # Cognitive Load Reduction The purpose is NOT task management. The purpose is reducing the need to remember. Traditional tools remember things once entered. Steward helps remember to remember. Examples "I noticed you mentioned replacing the UPS several times." "I noticed your dental check-up is overdue." "I've researched family camping locations." "I've identified a free weekend." --- # Communication Philosophy Steward communicates only when useful. Communication categories: Critical Immediate interruption. Decision Required Requests human approval. Background No notification required. Daily reports may exist but are not central. Real-time context-aware communication is the default. --- # Curiosity Budget One of Steward's defining concepts. If no operational work exists, Steward may spend limited compute researching or improving delegated objectives. Examples - Evaluate ArgoCD - Benchmark local LLMs - Research family holidays - Compare UPS replacements - Prototype monitoring improvements Curiosity is budgeted. The budget is a concrete, metered resource, not an aspiration. It is expressed in measurable units (for example: tokens/day, dollars/month, GPU-hours/week), set by the operator, and enforced by Steward refusing to exceed it. Consumption is observable so the operator can see what curiosity cost and what it produced. It is suspended immediately if operational work becomes necessary. --- # Confidence Steward maintains historical confidence scores. Example Restart Jellyfin Attempts: 34 Success: 33 Confidence: 97% Upgrade Kubernetes Attempts: 2 Success: 1 Rollback: 1 Confidence: 50% Confidence is earned. ## Confidence and Autonomy Confidence influences autonomy, but must never override policy. Rules: - The policy tier of an action (automatic / approval / never) is human-set and immutable to Steward. High confidence never promotes an action into a more permissive tier. Doing so would be self-modification by the back door. - Within a tier, confidence affects only *how* Steward acts: how strongly it recommends, how much evidence it attaches, whether it batches or surfaces individually. - Confidence is scoped, not global. "Restart Jellyfin" confidence does not transfer to "restart Postgres". Scope is (action × target × context), and Steward must not generalise across scopes without evidence. - Confidence decays. Environment changes (version upgrades, config changes, topology changes) invalidate historical success and should reset or discount the relevant scores. Stale confidence is treated as low confidence. --- # Security Model Safety is a property of the architecture, not just the model's good behaviour. ## The policy engine is the trust boundary The policy engine is Steward's crown jewel and must be a separate, independently-audited component that the AI *calls*, not code the AI can read, modify, or bypass. The AI proposes actions; the policy engine decides. If the AI could edit the engine, every other guarantee collapses. ## Event content is untrusted Events carry attacker-influenceable content: a GitHub PR title, a Telegram message, a webhook payload. Steward must treat all event content as untrusted input. - Event content can inform situations and goals but can never escalate authority or promote an action's policy tier. - Instructions embedded in event content are data, not commands (defence against prompt injection). ## Memory can rot "Nothing is forgotten" is both a feature and a liability. A learned preference can encode a mistake permanently, and learned lessons are themselves derived from potentially-untrusted history. - Learned memory is reviewable and retractable. A bad lesson must be able to be inspected and removed. - New lessons that would change behaviour surface as *proposals*, consistent with [Reflection](#reflection), rather than silently altering future decisions. --- # Memory Different information belongs in different storage. The roles below are fixed; the named tools are current implementation choices, not commitments. Structured store (e.g. NocoDB) Examples - Tasks - Goals - Incidents - Systems - Services - Confidence - Policies Long-form store (e.g. vector database / RAG) Examples - documentation - conversations - runbooks - release notes Metrics store (e.g. Prometheus) Examples - resource usage - trends - operational history Audit Immutable execution history. Nothing is forgotten. --- # Self Improvement Steward should become better over time. Important distinction Self-improving ✅ Learn preferences ✅ Learn successful remediations ✅ Improve prompts ✅ Improve playbooks ✅ Improve planning Self-modifying ❌ Rewrite policy engine ❌ Change permissions ❌ Remove safety controls Immutable system components remain protected. --- # Learning Steward continuously learns: Operator preferences Family preferences Infrastructure behaviour Recurring incidents Successful remediation Failed remediation Communication preferences Historical context This learning improves planning. --- # Reflection Periodic reflection sessions. Questions include: What repeated? What wasted time? What should be automated? Which playbooks failed? Which policies need review? Which research produced value? Reflection generates proposals rather than automatic change. --- # Research Workflow Research tasks behave like engineering work. Example Research destination --> Gather information --> Summarise findings --> Generate recommendations --> Present proposal --> Receive feedback --> Update understanding --> Re-plan --- # Engineering Workflow Operational improvements follow software engineering practices. Observe issue --> Generate proposal --> Approval --> Create Git branch --> Modify playbooks --> Run tests --> Open Pull Request --> Deploy through AWX --> Observe --> Learn --- # Long-term Vision Steward evolves from: Assistant --> Operator --> Engineer --> Trusted Steward Trust is earned through demonstrated reliability rather than increasing model capability. --- # Success Criteria Steward is successful when: I think less. I forget less. Routine work disappears. The homelab becomes increasingly self-maintaining. Projects continue progressing while I am busy. The system accumulates operational knowledge. The system proactively suggests valuable improvements. The system remains predictable. The system remains explainable. The system remains safe. ## Measurable Proxies The criteria above are subjective. To tell whether v0.2 actually beat v0.1, track measurable proxies alongside them: - Operator interruptions per week (trend down). - Ratio of actions auto-approved vs escalated for approval. - Mean time to remediation for recurring incidents. - False-alarm / unnecessary-notification rate. - Curiosity spend vs value produced (proposals accepted). --- # Mission Statement Steward exists to maximise the reliability, maintainability and usefulness of its delegated domains while minimising operator cognitive load. It continuously observes, remembers, researches, plans and acts within explicitly delegated authority. It values safety over speed, explanation over opacity, and long-term trust over short-term autonomy.