steward_mirror/AGENTS.md
Daniel Wagner 070e967833 test: update test suite for OmegaConf config system
- Add tests/conftest.py with make_settings() helper that maps legacy config
  parameter names to new nested OmegaConf structure
- Update all test files to use the conftest fixture
- All 73 tests now pass with new config system
- Maintains backward compatibility via Settings class properties

Also add AGENTS.md with comprehensive AI agent guidelines:
- Project overview and key technologies
- Directory structure reference
- Development workflow and common tasks
- Testing strategy and patterns
- CI/CD pipeline overview
- Common pitfalls and best practices
- Debugging guide for agents working on the project
2026-07-26 12:48:40 +10:00

317 lines
9.9 KiB
Markdown

# AI Agent Guidelines for Steward
This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.
## Project Overview
**Steward** is a long-running, AI-assisted personal operations platform that:
- Runs as a Telegram bot with group/channel support
- Uses OmegaConf for flexible configuration management
- Supports message threads for organized conversations
- Stores and retrieves conversation summaries via knowledge base
- Integrates with OpenAI-compatible LLM APIs
- Optionally integrates with MCP/OpenAPI tool servers for function calling
**Key Technologies:**
- Python 3.12+
- `python-telegram-bot` (Telegram bot framework)
- OmegaConf (configuration management)
- Pydantic (data validation)
- OpenAI API (LLM integration)
- AsyncIO (async/await patterns throughout)
## Directory Structure
```
steward/
├── steward/ # Main package
│ ├── bot/ # Telegram bot implementation
│ │ └── telegram.py # All bot handlers and logic
│ ├── config.py # OmegaConf-based configuration
│ ├── config_schema.yaml # Default configuration schema
│ ├── llm/ # LLM client
│ ├── memory/ # Thread memory storage
│ ├── proposals/ # Proposal generation
│ ├── tools/ # MCP tool integration
│ └── main.py # Application entry point
├── tests/ # Test suite
│ ├── conftest.py # Pytest fixtures and helpers
│ ├── test_bot.py # Bot handler tests
│ ├── test_llm_client.py # LLM client tests
│ ├── test_proposals.py # Proposal generation tests
│ ├── test_thread_memory.py # Memory storage tests
│ └── test_tools.py # Tool integration tests
├── .github/workflows/ # CI/CD workflows
├── docker-compose.yml # Production compose (uses ghcr.io image)
├── docker-compose.dev.yml # Development compose (builds locally)
├── Dockerfile # Container image definition
├── pyproject.toml # Project metadata and dependencies
├── CONFIGURATION.md # Configuration guide
└── AGENTS.md # This file
```
## Development Workflow
### Before Making Changes
1. **Understand the Context**
- Read CONFIGURATION.md for config system details
- Review the async/await patterns in existing code
- Note that the bot uses Telegram's `python-telegram-bot` library
2. **Check Existing Tests**
- Run tests locally or via CI before proposing changes
- Tests use `tests/conftest.py` fixtures for setup
- The `_make_settings()` helper maps legacy config to new structure
3. **Code Style**
- Follow the ruff linting rules in `pyproject.toml`
- Line length: 100 characters
- Use type hints throughout (mypy strict mode)
- Import order: standard lib → third party → local
### Making Changes
#### Configuration Changes
- Update `steward/config_schema.yaml` for schema changes
- Update `steward/config.py` for new config sections
- Update `.env.example` to show all available settings
- Document in `CONFIGURATION.md`
#### Bot Handler Changes
- All handlers in `steward/bot/telegram.py` are async
- Handlers receive `Update` and `ContextTypes.DEFAULT_TYPE` parameters
- Use `update.message.reply_text()` for responses
- Check user/group authorization early in handlers
- Test with `tests/test_bot.py`
#### Adding Tests
- Use `tests/conftest.py` for shared fixtures
- Use `_make_settings()` helper to create test settings
- Use `_make_update()` to create mock Telegram updates
- All bot tests should be async (use `pytest-asyncio`)
- Mock external APIs with `respx` or `pytest-mock`
#### Creating New Features
**For Configuration:**
1. Add section to `config_schema.yaml`
2. Create pydantic model in `config.py`
3. Add properties to `Settings` class for backward compatibility
4. Document in `CONFIGURATION.md`
**For Bot Handlers:**
1. Add async handler function in `steward/bot/telegram.py`
2. Add to application in `build_application()`
3. Add authorization checks (`_is_allowed()`, `_is_group_enabled()`)
4. Write tests in `tests/test_bot.py`
**For LLM Integration:**
1. Update `steward/llm/client.py`
2. Add history context handling if needed
3. Test with `tests/test_llm_client.py`
4. Consider impact on token usage
### Docker & Deployment
**Local Development:**
```bash
docker compose -f docker-compose.dev.yml up --build
```
**Production:**
```bash
docker compose up
# Uses prebuilt image from ghcr.io/djw4/steward:latest
```
**Key Docker Notes:**
- Uses `uv` instead of pip for faster installation
- Non-root user (UID 1000) for security
- Thread memory stored in `/data` volume
- Environment variables set via `.env` file
## Common Tasks
### Adding a New Bot Command
1. Create handler function in `steward/bot/telegram.py`:
```python
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
settings: Settings = context.bot_data["settings"]
user = update.effective_user
chat = update.effective_chat
# Authorization checks
if user is None or not _is_allowed(user.id, settings):
return
if chat is None or not _is_group_enabled(chat.id, settings):
return
# Handler logic here
await update.message.reply_text("Response")
```
2. Register in `build_application()`:
```python
app.add_handler(CommandHandler("mycommand", mycommand_handler))
```
3. Add test in `tests/test_bot.py`:
```python
@pytest.mark.asyncio
async def test_mycommand_handler_does_something():
# Test implementation
```
### Adding a New Configuration Section
1. Update `steward/config_schema.yaml`:
```yaml
mysection:
my_setting: "default_value"
my_number: 42
```
2. Create pydantic model in `steward/config.py`:
```python
class MyConfig(BaseModel):
my_setting: str = Field(default="default_value")
my_number: int = Field(default=42)
```
3. Add to `Settings` class and add property for backward compatibility:
```python
class Settings(BaseModel):
mysection: MyConfig = Field(default_factory=MyConfig)
@property
def my_setting(self) -> str:
return self.mysection.my_setting
```
### Running Tests Locally
```bash
# Install dev dependencies
pip install -e ".[dev]"
# Run all tests
pytest tests/ -v
# Run specific test file
pytest tests/test_bot.py -v
# Run specific test
pytest tests/test_bot.py::test_mytest -v
# Run with coverage
pytest tests/ --cov=steward --cov-report=html
```
### Debugging Issues
**Bot not responding:**
1. Check `.env` file has all required settings
2. Verify bot has message access in BotFather
3. Check logs: `docker compose logs steward -f`
4. Verify group ID format (negative for groups: `-1005308306472`)
**Tests failing:**
1. Run `ruff check steward/ tests/` to check linting
2. Run `mypy steward/` for type checking
3. Check test fixture setup in `conftest.py`
4. Review error messages carefully for import/config issues
**Configuration issues:**
1. Check `.env` file syntax (should be `KEY=value`)
2. For lists, use JSON format: `STEWARD__TELEGRAM__ALLOWED_USER_IDS='[123, 456]'`
3. New format variables use `STEWARD__SECTION__KEY` pattern
4. Check `CONFIGURATION.md` for examples
## Testing Strategy
### Test Organization
- **`test_bot.py`**: Handler logic, authorization, history management
- **`test_llm_client.py`**: LLM API integration, error handling
- **`test_proposals.py`**: Proposal generation, API calls
- **`test_thread_memory.py`**: Memory storage, search, serialization
- **`test_tools.py`**: Tool calling, MCP integration
### Test Patterns
**Handler Tests:**
```python
@pytest.mark.asyncio
async def test_handler_name():
settings = _make_settings()
update = _make_update(user_id=123, text="hello")
context = _make_context(settings)
await handler_name(update, context)
# Assert on reply_text mock calls
```
**API Tests:**
```python
def test_api_call():
with respx.mock:
respx.post("https://api.example.com/endpoint").mock(
return_value=httpx.Response(200, json={"result": "ok"})
)
# Call function that makes API request
result = function_under_test()
assert result == expected_value
```
### CI/CD Pipeline
The CI workflow (`.github/workflows/ci.yml`) runs:
1. **Linting** (ruff): `python -m ruff check steward/ tests/`
2. **Tests** (pytest): `python -m pytest tests/ -v`
3. **Docker Build**: Builds and pushes image to GHCR
All three must pass for merges to main.
## Common Pitfalls
### ❌ Don't:
- Use synchronous code instead of async/await
- Hardcode secrets in config files (use env vars)
- Forget to add authorization checks to new handlers
- Write tests without using conftest fixtures
- Modify `.env` in version control (update `.env.example` instead)
- Change config structure without updating all layers (schema, pydantic, docs)
### ✅ Do:
- Use `async/await` consistently throughout
- Store secrets in environment variables
- Check both user and group authorization
- Use `pytest.mark.asyncio` for async tests
- Keep `.env` out of git (it's in `.gitignore`)
- Update schema → pydantic → properties → documentation together
## Questions? Issues?
- Check `CONFIGURATION.md` for config-related questions
- Review existing handlers in `steward/bot/telegram.py` for patterns
- Look at test examples in `tests/` for testing patterns
- Check `.github/workflows/ci.yml` for what CI expects
## Key Files to Know
| File | Purpose |
|------|---------|
| `steward/config.py` | Configuration system - OmegaConf + pydantic |
| `steward/bot/telegram.py` | All bot handlers and core logic |
| `steward/main.py` | Application entry point and setup |
| `tests/conftest.py` | Pytest fixtures and helpers |
| `pyproject.toml` | Dependencies and tool configuration |
| `CONFIGURATION.md` | User-facing config guide |
| `.github/workflows/ci.yml` | CI/CD pipeline definition |