steward_mirror/AGENTS.md
Andrew Ridgway 0667adbf72
Revert "docs: document pr_reviewer manual-trigger workflow in AGENTS.md"
This reverts commit e2ba3dd0990d31167f48bfd0f33f273f0e426514.
2026-08-18 22:08:16 +10:00

355 lines
11 KiB
Markdown

# AI Agent Guidelines for Steward
This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.
## Project Overview
**Steward** is a long-running, AI-assisted personal operations platform that:
- Runs as a Telegram bot with group/channel support
- Uses OmegaConf for flexible configuration management
- Supports message threads for organized conversations
- Stores and retrieves conversation summaries via knowledge base
- Integrates with OpenAI-compatible LLM APIs
- Optionally integrates with MCP/OpenAPI tool servers for function calling
**Key Technologies:**
- Python 3.12+
- `python-telegram-bot` (Telegram bot framework)
- OmegaConf (configuration management)
- Pydantic (data validation)
- OpenAI API (LLM integration)
- AsyncIO (async/await patterns throughout)
## Directory Structure
```
steward/
├── steward/ # Main package
│ ├── bot/ # Telegram bot implementation
│ │ └── telegram.py # All bot handlers and logic
│ ├── config.py # OmegaConf-based configuration
│ ├── config_schema.yaml # Default configuration schema
│ ├── llm/ # LLM client
│ ├── memory/ # Thread memory storage
│ ├── proposals/ # Proposal generation
│ ├── tools/ # MCP tool integration
│ └── main.py # Application entry point
├── tests/ # Test suite
│ ├── conftest.py # Pytest fixtures and helpers
│ ├── test_bot.py # Bot handler tests
│ ├── test_llm_client.py # LLM client tests
│ ├── test_proposals.py # Proposal generation tests
│ ├── test_thread_memory.py # Memory storage tests
│ └── test_tools.py # Tool integration tests
├── .github/workflows/ # CI/CD workflows
├── docker-compose.yml # Production compose (uses ghcr.io image)
├── docker-compose.dev.yml # Development compose (builds locally)
├── Dockerfile # Container image definition
├── pyproject.toml # Project metadata and dependencies
├── CONFIGURATION.md # Configuration guide
└── AGENTS.md # This file
```
## Development Workflow
### Initial Setup
1. **Install uv** (faster Python package manager):
```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env # Add to PATH
```
2. **Create virtual environment and install dependencies**:
```bash
uv venv
source .venv/bin/activate
uv pip install -e ".[dev]"
```
3. **Set up pre-commit hooks** (runs ruff + tests on commit):
```bash
pre-commit install
```
From now on, `git commit` will automatically:
- Run `ruff check` and `ruff format` (with auto-fix)
- Run `pytest` to ensure tests pass
You can skip hooks with `git commit --no-verify` if needed.
### Before Making Changes
1. **Understand the Context**
- Read CONFIGURATION.md for config system details
- Review the async/await patterns in existing code
- Note that the bot uses Telegram's `python-telegram-bot` library
2. **Check Existing Tests**
- Run tests locally or via CI before proposing changes
- Tests use `tests/conftest.py` fixtures for setup
- The `_make_settings()` helper maps legacy config to new structure
3. **Code Style**
- Follow the ruff linting rules in `pyproject.toml`
- Line length: 100 characters
- Use type hints throughout (mypy strict mode)
- Import order: standard lib → third party → local
- Pre-commit will auto-fix most linting issues
### Making Changes
#### Configuration Changes
- Update `steward/config_schema.yaml` for schema changes
- Update `steward/config.py` for new config sections
- Update `.env.example` to show all available settings
- Document in `CONFIGURATION.md`
#### Bot Handler Changes
- All handlers in `steward/bot/telegram.py` are async
- Handlers receive `Update` and `ContextTypes.DEFAULT_TYPE` parameters
- Use `update.message.reply_text()` for responses
- Check user/group authorization early in handlers
- Test with `tests/test_bot.py`
#### Adding Tests
- Use `tests/conftest.py` for shared fixtures
- Use `_make_settings()` helper to create test settings
- Use `_make_update()` to create mock Telegram updates
- All bot tests should be async (use `pytest-asyncio`)
- Mock external APIs with `respx` or `pytest-mock`
#### Creating New Features
**For Configuration:**
1. Add section to `config_schema.yaml`
2. Create pydantic model in `config.py`
3. Add properties to `Settings` class for backward compatibility
4. Document in `CONFIGURATION.md`
**For Bot Handlers:**
1. Add async handler function in `steward/bot/telegram.py`
2. Add to application in `build_application()`
3. Add authorization checks (`_is_allowed()`, `_is_group_enabled()`)
4. Write tests in `tests/test_bot.py`
**For LLM Integration:**
1. Update `steward/llm/client.py`
2. Add history context handling if needed
3. Test with `tests/test_llm_client.py`
4. Consider impact on token usage
### Docker & Deployment
**Local Development:**
```bash
docker compose -f docker-compose.dev.yml up --build
```
**Production:**
```bash
docker compose up
# Uses prebuilt image from ghcr.io/djw4/steward:latest
```
**Key Docker Notes:**
- Uses `uv` instead of pip for faster installation
- Non-root user (UID 1000) for security
- Thread memory stored in `/data` volume
- Environment variables set via `.env` file
## Common Tasks
### Adding a New Bot Command
1. Create handler function in `steward/bot/telegram.py`:
```python
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
settings: Settings = context.bot_data["settings"]
user = update.effective_user
chat = update.effective_chat
# Authorization checks
if user is None or not _is_allowed(user.id, settings):
return
if chat is None or not _is_group_enabled(chat.id, settings):
return
# Handler logic here
await update.message.reply_text("Response")
```
2. Register in `build_application()`:
```python
app.add_handler(CommandHandler("mycommand", mycommand_handler))
```
3. Add test in `tests/test_bot.py`:
```python
@pytest.mark.asyncio
async def test_mycommand_handler_does_something():
# Test implementation
```
### Adding a New Configuration Section
1. Update `steward/config_schema.yaml`:
```yaml
mysection:
my_setting: "default_value"
my_number: 42
```
2. Create pydantic model in `steward/config.py`:
```python
class MyConfig(BaseModel):
my_setting: str = Field(default="default_value")
my_number: int = Field(default=42)
```
3. Add to `Settings` class and add property for backward compatibility:
```python
class Settings(BaseModel):
mysection: MyConfig = Field(default_factory=MyConfig)
@property
def my_setting(self) -> str:
return self.mysection.my_setting
```
### Running Tests Locally
After setting up your venv (see Initial Setup), you can run tests:
```bash
# Activate venv if not already activated
source .venv/bin/activate
# Run all tests
pytest tests/ -v
# Run specific test file
pytest tests/test_bot.py -v
# Run specific test
pytest tests/test_bot.py::test_mytest -v
# Run with coverage
pytest tests/ --cov=steward --cov-report=html
# Run linting only (pre-commit does this automatically)
ruff check steward/ tests/
# Auto-fix linting issues
ruff check steward/ tests/ --fix
ruff format steward/ tests/
```
Note: Pre-commit will run these checks automatically on commit, so you usually don't need to run them manually unless you want to check before committing.
### Debugging Issues
**Bot not responding:**
1. Check `.env` file has all required settings
2. Verify bot has message access in BotFather
3. Check logs: `docker compose logs steward -f`
4. Verify group ID format (negative for groups: `-1005308306472`)
**Tests failing:**
1. Run `ruff check steward/ tests/` to check linting
2. Run `mypy steward/` for type checking
3. Check test fixture setup in `conftest.py`
4. Review error messages carefully for import/config issues
**Configuration issues:**
1. Check `.env` file syntax (should be `KEY=value`)
2. For lists, prefer CSV format: `STEWARD__TELEGRAM__ALLOWED_USER_IDS=123,456` (JSON lists are also supported)
3. New format variables use `STEWARD__SECTION__KEY` pattern
4. Check `CONFIGURATION.md` for examples
## Testing Strategy
### Test Organization
- **`test_bot.py`**: Handler logic, authorization, history management
- **`test_llm_client.py`**: LLM API integration, error handling
- **`test_proposals.py`**: Proposal generation, API calls
- **`test_thread_memory.py`**: Memory storage, search, serialization
- **`test_tools.py`**: Tool calling, MCP integration
### Test Patterns
**Handler Tests:**
```python
@pytest.mark.asyncio
async def test_handler_name():
settings = _make_settings()
update = _make_update(user_id=123, text="hello")
context = _make_context(settings)
await handler_name(update, context)
# Assert on reply_text mock calls
```
**API Tests:**
```python
def test_api_call():
with respx.mock:
respx.post("https://api.example.com/endpoint").mock(
return_value=httpx.Response(200, json={"result": "ok"})
)
# Call function that makes API request
result = function_under_test()
assert result == expected_value
```
### CI/CD Pipeline
The CI workflow (`.github/workflows/ci.yml`) runs:
1. **Linting** (ruff): `python -m ruff check steward/ tests/`
2. **Tests** (pytest): `python -m pytest tests/ -v`
3. **Docker Build**: Builds and pushes image to GHCR
All three must pass for merges to main.
## Common Pitfalls
### ❌ Don't:
- Use synchronous code instead of async/await
- Hardcode secrets in config files (use env vars)
- Forget to add authorization checks to new handlers
- Write tests without using conftest fixtures
- Modify `.env` in version control (update `.env.example` instead)
- Change config structure without updating all layers (schema, pydantic, docs)
### ✅ Do:
- Use `async/await` consistently throughout
- Store secrets in environment variables
- Check both user and group authorization
- Use `pytest.mark.asyncio` for async tests
- Keep `.env` out of git (it's in `.gitignore`)
- Update schema → pydantic → properties → documentation together
## Questions? Issues?
- Check `CONFIGURATION.md` for config-related questions
- Review existing handlers in `steward/bot/telegram.py` for patterns
- Look at test examples in `tests/` for testing patterns
- Check `.github/workflows/ci.yml` for what CI expects
## Key Files to Know
| File | Purpose |
|------|---------|
| `steward/config.py` | Configuration system - OmegaConf + pydantic |
| `steward/bot/telegram.py` | All bot handlers and core logic |
| `steward/main.py` | Application entry point and setup |
| `tests/conftest.py` | Pytest fixtures and helpers |
| `pyproject.toml` | Dependencies and tool configuration |
| `CONFIGURATION.md` | User-facing config guide |
| `.github/workflows/ci.yml` | CI/CD pipeline definition |