steward_mirror/AGENTS.md
Andrew Ridgway 0667adbf72
Revert "docs: document pr_reviewer manual-trigger workflow in AGENTS.md"
This reverts commit e2ba3dd0990d31167f48bfd0f33f273f0e426514.
2026-08-18 22:08:16 +10:00

11 KiB

AI Agent Guidelines for Steward

This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.

Project Overview

Steward is a long-running, AI-assisted personal operations platform that:

  • Runs as a Telegram bot with group/channel support
  • Uses OmegaConf for flexible configuration management
  • Supports message threads for organized conversations
  • Stores and retrieves conversation summaries via knowledge base
  • Integrates with OpenAI-compatible LLM APIs
  • Optionally integrates with MCP/OpenAPI tool servers for function calling

Key Technologies:

  • Python 3.12+
  • python-telegram-bot (Telegram bot framework)
  • OmegaConf (configuration management)
  • Pydantic (data validation)
  • OpenAI API (LLM integration)
  • AsyncIO (async/await patterns throughout)

Directory Structure

steward/
├── steward/                    # Main package
│   ├── bot/                   # Telegram bot implementation
│   │   └── telegram.py        # All bot handlers and logic
│   ├── config.py              # OmegaConf-based configuration
│   ├── config_schema.yaml     # Default configuration schema
│   ├── llm/                   # LLM client
│   ├── memory/                # Thread memory storage
│   ├── proposals/             # Proposal generation
│   ├── tools/                 # MCP tool integration
│   └── main.py                # Application entry point
├── tests/                     # Test suite
│   ├── conftest.py            # Pytest fixtures and helpers
│   ├── test_bot.py            # Bot handler tests
│   ├── test_llm_client.py     # LLM client tests
│   ├── test_proposals.py      # Proposal generation tests
│   ├── test_thread_memory.py  # Memory storage tests
│   └── test_tools.py          # Tool integration tests
├── .github/workflows/         # CI/CD workflows
├── docker-compose.yml         # Production compose (uses ghcr.io image)
├── docker-compose.dev.yml     # Development compose (builds locally)
├── Dockerfile                 # Container image definition
├── pyproject.toml             # Project metadata and dependencies
├── CONFIGURATION.md           # Configuration guide
└── AGENTS.md                  # This file

Development Workflow

Initial Setup

  1. Install uv (faster Python package manager):
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env  # Add to PATH
  1. Create virtual environment and install dependencies:
uv venv
source .venv/bin/activate
uv pip install -e ".[dev]"
  1. Set up pre-commit hooks (runs ruff + tests on commit):
pre-commit install

From now on, git commit will automatically:

  • Run ruff check and ruff format (with auto-fix)
  • Run pytest to ensure tests pass

You can skip hooks with git commit --no-verify if needed.

Before Making Changes

  1. Understand the Context

    • Read CONFIGURATION.md for config system details
    • Review the async/await patterns in existing code
    • Note that the bot uses Telegram's python-telegram-bot library
  2. Check Existing Tests

    • Run tests locally or via CI before proposing changes
    • Tests use tests/conftest.py fixtures for setup
    • The _make_settings() helper maps legacy config to new structure
  3. Code Style

    • Follow the ruff linting rules in pyproject.toml
    • Line length: 100 characters
    • Use type hints throughout (mypy strict mode)
    • Import order: standard lib → third party → local
    • Pre-commit will auto-fix most linting issues

Making Changes

Configuration Changes

  • Update steward/config_schema.yaml for schema changes
  • Update steward/config.py for new config sections
  • Update .env.example to show all available settings
  • Document in CONFIGURATION.md

Bot Handler Changes

  • All handlers in steward/bot/telegram.py are async
  • Handlers receive Update and ContextTypes.DEFAULT_TYPE parameters
  • Use update.message.reply_text() for responses
  • Check user/group authorization early in handlers
  • Test with tests/test_bot.py

Adding Tests

  • Use tests/conftest.py for shared fixtures
  • Use _make_settings() helper to create test settings
  • Use _make_update() to create mock Telegram updates
  • All bot tests should be async (use pytest-asyncio)
  • Mock external APIs with respx or pytest-mock

Creating New Features

For Configuration:

  1. Add section to config_schema.yaml
  2. Create pydantic model in config.py
  3. Add properties to Settings class for backward compatibility
  4. Document in CONFIGURATION.md

For Bot Handlers:

  1. Add async handler function in steward/bot/telegram.py
  2. Add to application in build_application()
  3. Add authorization checks (_is_allowed(), _is_group_enabled())
  4. Write tests in tests/test_bot.py

For LLM Integration:

  1. Update steward/llm/client.py
  2. Add history context handling if needed
  3. Test with tests/test_llm_client.py
  4. Consider impact on token usage

Docker & Deployment

Local Development:

docker compose -f docker-compose.dev.yml up --build

Production:

docker compose up
# Uses prebuilt image from ghcr.io/djw4/steward:latest

Key Docker Notes:

  • Uses uv instead of pip for faster installation
  • Non-root user (UID 1000) for security
  • Thread memory stored in /data volume
  • Environment variables set via .env file

Common Tasks

Adding a New Bot Command

  1. Create handler function in steward/bot/telegram.py:
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
    settings: Settings = context.bot_data["settings"]
    user = update.effective_user
    chat = update.effective_chat
    
    # Authorization checks
    if user is None or not _is_allowed(user.id, settings):
        return
    if chat is None or not _is_group_enabled(chat.id, settings):
        return
    
    # Handler logic here
    await update.message.reply_text("Response")
  1. Register in build_application():
app.add_handler(CommandHandler("mycommand", mycommand_handler))
  1. Add test in tests/test_bot.py:
@pytest.mark.asyncio
async def test_mycommand_handler_does_something():
    # Test implementation

Adding a New Configuration Section

  1. Update steward/config_schema.yaml:
mysection:
  my_setting: "default_value"
  my_number: 42
  1. Create pydantic model in steward/config.py:
class MyConfig(BaseModel):
    my_setting: str = Field(default="default_value")
    my_number: int = Field(default=42)
  1. Add to Settings class and add property for backward compatibility:
class Settings(BaseModel):
    mysection: MyConfig = Field(default_factory=MyConfig)
    
    @property
    def my_setting(self) -> str:
        return self.mysection.my_setting

Running Tests Locally

After setting up your venv (see Initial Setup), you can run tests:

# Activate venv if not already activated
source .venv/bin/activate

# Run all tests
pytest tests/ -v

# Run specific test file
pytest tests/test_bot.py -v

# Run specific test
pytest tests/test_bot.py::test_mytest -v

# Run with coverage
pytest tests/ --cov=steward --cov-report=html

# Run linting only (pre-commit does this automatically)
ruff check steward/ tests/

# Auto-fix linting issues
ruff check steward/ tests/ --fix
ruff format steward/ tests/

Note: Pre-commit will run these checks automatically on commit, so you usually don't need to run them manually unless you want to check before committing.

Debugging Issues

Bot not responding:

  1. Check .env file has all required settings
  2. Verify bot has message access in BotFather
  3. Check logs: docker compose logs steward -f
  4. Verify group ID format (negative for groups: -1005308306472)

Tests failing:

  1. Run ruff check steward/ tests/ to check linting
  2. Run mypy steward/ for type checking
  3. Check test fixture setup in conftest.py
  4. Review error messages carefully for import/config issues

Configuration issues:

  1. Check .env file syntax (should be KEY=value)
  2. For lists, prefer CSV format: STEWARD__TELEGRAM__ALLOWED_USER_IDS=123,456 (JSON lists are also supported)
  3. New format variables use STEWARD__SECTION__KEY pattern
  4. Check CONFIGURATION.md for examples

Testing Strategy

Test Organization

  • test_bot.py: Handler logic, authorization, history management
  • test_llm_client.py: LLM API integration, error handling
  • test_proposals.py: Proposal generation, API calls
  • test_thread_memory.py: Memory storage, search, serialization
  • test_tools.py: Tool calling, MCP integration

Test Patterns

Handler Tests:

@pytest.mark.asyncio
async def test_handler_name():
    settings = _make_settings()
    update = _make_update(user_id=123, text="hello")
    context = _make_context(settings)
    
    await handler_name(update, context)
    
    # Assert on reply_text mock calls

API Tests:

def test_api_call():
    with respx.mock:
        respx.post("https://api.example.com/endpoint").mock(
            return_value=httpx.Response(200, json={"result": "ok"})
        )
        
        # Call function that makes API request
        result = function_under_test()
        
        assert result == expected_value

CI/CD Pipeline

The CI workflow (.github/workflows/ci.yml) runs:

  1. Linting (ruff): python -m ruff check steward/ tests/
  2. Tests (pytest): python -m pytest tests/ -v
  3. Docker Build: Builds and pushes image to GHCR

All three must pass for merges to main.

Common Pitfalls

❌ Don't:

  • Use synchronous code instead of async/await
  • Hardcode secrets in config files (use env vars)
  • Forget to add authorization checks to new handlers
  • Write tests without using conftest fixtures
  • Modify .env in version control (update .env.example instead)
  • Change config structure without updating all layers (schema, pydantic, docs)

✅ Do:

  • Use async/await consistently throughout
  • Store secrets in environment variables
  • Check both user and group authorization
  • Use pytest.mark.asyncio for async tests
  • Keep .env out of git (it's in .gitignore)
  • Update schema → pydantic → properties → documentation together

Questions? Issues?

  • Check CONFIGURATION.md for config-related questions
  • Review existing handlers in steward/bot/telegram.py for patterns
  • Look at test examples in tests/ for testing patterns
  • Check .github/workflows/ci.yml for what CI expects

Key Files to Know

File Purpose
steward/config.py Configuration system - OmegaConf + pydantic
steward/bot/telegram.py All bot handlers and core logic
steward/main.py Application entry point and setup
tests/conftest.py Pytest fixtures and helpers
pyproject.toml Dependencies and tool configuration
CONFIGURATION.md User-facing config guide
.github/workflows/ci.yml CI/CD pipeline definition