355 lines
11 KiB
Markdown
355 lines
11 KiB
Markdown
# AI Agent Guidelines for Steward
|
|
|
|
This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.
|
|
|
|
## Project Overview
|
|
|
|
**Steward** is a long-running, AI-assisted personal operations platform that:
|
|
- Runs as a Telegram bot with group/channel support
|
|
- Uses OmegaConf for flexible configuration management
|
|
- Supports message threads for organized conversations
|
|
- Stores and retrieves conversation summaries via knowledge base
|
|
- Integrates with OpenAI-compatible LLM APIs
|
|
- Optionally integrates with MCP/OpenAPI tool servers for function calling
|
|
|
|
**Key Technologies:**
|
|
- Python 3.12+
|
|
- `python-telegram-bot` (Telegram bot framework)
|
|
- OmegaConf (configuration management)
|
|
- Pydantic (data validation)
|
|
- OpenAI API (LLM integration)
|
|
- AsyncIO (async/await patterns throughout)
|
|
|
|
## Directory Structure
|
|
|
|
```
|
|
steward/
|
|
├── steward/ # Main package
|
|
│ ├── bot/ # Telegram bot implementation
|
|
│ │ └── telegram.py # All bot handlers and logic
|
|
│ ├── config.py # OmegaConf-based configuration
|
|
│ ├── config_schema.yaml # Default configuration schema
|
|
│ ├── llm/ # LLM client
|
|
│ ├── memory/ # Thread memory storage
|
|
│ ├── proposals/ # Proposal generation
|
|
│ ├── tools/ # MCP tool integration
|
|
│ └── main.py # Application entry point
|
|
├── tests/ # Test suite
|
|
│ ├── conftest.py # Pytest fixtures and helpers
|
|
│ ├── test_bot.py # Bot handler tests
|
|
│ ├── test_llm_client.py # LLM client tests
|
|
│ ├── test_proposals.py # Proposal generation tests
|
|
│ ├── test_thread_memory.py # Memory storage tests
|
|
│ └── test_tools.py # Tool integration tests
|
|
├── .github/workflows/ # CI/CD workflows
|
|
├── docker-compose.yml # Production compose (uses ghcr.io image)
|
|
├── docker-compose.dev.yml # Development compose (builds locally)
|
|
├── Dockerfile # Container image definition
|
|
├── pyproject.toml # Project metadata and dependencies
|
|
├── CONFIGURATION.md # Configuration guide
|
|
└── AGENTS.md # This file
|
|
```
|
|
|
|
## Development Workflow
|
|
|
|
### Initial Setup
|
|
|
|
1. **Install uv** (faster Python package manager):
|
|
```bash
|
|
curl -LsSf https://astral.sh/uv/install.sh | sh
|
|
source $HOME/.local/bin/env # Add to PATH
|
|
```
|
|
|
|
2. **Create virtual environment and install dependencies**:
|
|
```bash
|
|
uv venv
|
|
source .venv/bin/activate
|
|
uv pip install -e ".[dev]"
|
|
```
|
|
|
|
3. **Set up pre-commit hooks** (runs ruff + tests on commit):
|
|
```bash
|
|
pre-commit install
|
|
```
|
|
|
|
From now on, `git commit` will automatically:
|
|
- Run `ruff check` and `ruff format` (with auto-fix)
|
|
- Run `pytest` to ensure tests pass
|
|
|
|
You can skip hooks with `git commit --no-verify` if needed.
|
|
|
|
### Before Making Changes
|
|
|
|
1. **Understand the Context**
|
|
- Read CONFIGURATION.md for config system details
|
|
- Review the async/await patterns in existing code
|
|
- Note that the bot uses Telegram's `python-telegram-bot` library
|
|
|
|
2. **Check Existing Tests**
|
|
- Run tests locally or via CI before proposing changes
|
|
- Tests use `tests/conftest.py` fixtures for setup
|
|
- The `_make_settings()` helper maps legacy config to new structure
|
|
|
|
3. **Code Style**
|
|
- Follow the ruff linting rules in `pyproject.toml`
|
|
- Line length: 100 characters
|
|
- Use type hints throughout (mypy strict mode)
|
|
- Import order: standard lib → third party → local
|
|
- Pre-commit will auto-fix most linting issues
|
|
|
|
### Making Changes
|
|
|
|
#### Configuration Changes
|
|
- Update `steward/config_schema.yaml` for schema changes
|
|
- Update `steward/config.py` for new config sections
|
|
- Update `.env.example` to show all available settings
|
|
- Document in `CONFIGURATION.md`
|
|
|
|
#### Bot Handler Changes
|
|
- All handlers in `steward/bot/telegram.py` are async
|
|
- Handlers receive `Update` and `ContextTypes.DEFAULT_TYPE` parameters
|
|
- Use `update.message.reply_text()` for responses
|
|
- Check user/group authorization early in handlers
|
|
- Test with `tests/test_bot.py`
|
|
|
|
#### Adding Tests
|
|
- Use `tests/conftest.py` for shared fixtures
|
|
- Use `_make_settings()` helper to create test settings
|
|
- Use `_make_update()` to create mock Telegram updates
|
|
- All bot tests should be async (use `pytest-asyncio`)
|
|
- Mock external APIs with `respx` or `pytest-mock`
|
|
|
|
#### Creating New Features
|
|
|
|
**For Configuration:**
|
|
1. Add section to `config_schema.yaml`
|
|
2. Create pydantic model in `config.py`
|
|
3. Add properties to `Settings` class for backward compatibility
|
|
4. Document in `CONFIGURATION.md`
|
|
|
|
**For Bot Handlers:**
|
|
1. Add async handler function in `steward/bot/telegram.py`
|
|
2. Add to application in `build_application()`
|
|
3. Add authorization checks (`_is_allowed()`, `_is_group_enabled()`)
|
|
4. Write tests in `tests/test_bot.py`
|
|
|
|
**For LLM Integration:**
|
|
1. Update `steward/llm/client.py`
|
|
2. Add history context handling if needed
|
|
3. Test with `tests/test_llm_client.py`
|
|
4. Consider impact on token usage
|
|
|
|
### Docker & Deployment
|
|
|
|
**Local Development:**
|
|
```bash
|
|
docker compose -f docker-compose.dev.yml up --build
|
|
```
|
|
|
|
**Production:**
|
|
```bash
|
|
docker compose up
|
|
# Uses prebuilt image from ghcr.io/djw4/steward:latest
|
|
```
|
|
|
|
**Key Docker Notes:**
|
|
- Uses `uv` instead of pip for faster installation
|
|
- Non-root user (UID 1000) for security
|
|
- Thread memory stored in `/data` volume
|
|
- Environment variables set via `.env` file
|
|
|
|
## Common Tasks
|
|
|
|
### Adding a New Bot Command
|
|
|
|
1. Create handler function in `steward/bot/telegram.py`:
|
|
```python
|
|
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
|
|
settings: Settings = context.bot_data["settings"]
|
|
user = update.effective_user
|
|
chat = update.effective_chat
|
|
|
|
# Authorization checks
|
|
if user is None or not _is_allowed(user.id, settings):
|
|
return
|
|
if chat is None or not _is_group_enabled(chat.id, settings):
|
|
return
|
|
|
|
# Handler logic here
|
|
await update.message.reply_text("Response")
|
|
```
|
|
|
|
2. Register in `build_application()`:
|
|
```python
|
|
app.add_handler(CommandHandler("mycommand", mycommand_handler))
|
|
```
|
|
|
|
3. Add test in `tests/test_bot.py`:
|
|
```python
|
|
@pytest.mark.asyncio
|
|
async def test_mycommand_handler_does_something():
|
|
# Test implementation
|
|
```
|
|
|
|
### Adding a New Configuration Section
|
|
|
|
1. Update `steward/config_schema.yaml`:
|
|
```yaml
|
|
mysection:
|
|
my_setting: "default_value"
|
|
my_number: 42
|
|
```
|
|
|
|
2. Create pydantic model in `steward/config.py`:
|
|
```python
|
|
class MyConfig(BaseModel):
|
|
my_setting: str = Field(default="default_value")
|
|
my_number: int = Field(default=42)
|
|
```
|
|
|
|
3. Add to `Settings` class and add property for backward compatibility:
|
|
```python
|
|
class Settings(BaseModel):
|
|
mysection: MyConfig = Field(default_factory=MyConfig)
|
|
|
|
@property
|
|
def my_setting(self) -> str:
|
|
return self.mysection.my_setting
|
|
```
|
|
|
|
### Running Tests Locally
|
|
|
|
After setting up your venv (see Initial Setup), you can run tests:
|
|
|
|
```bash
|
|
# Activate venv if not already activated
|
|
source .venv/bin/activate
|
|
|
|
# Run all tests
|
|
pytest tests/ -v
|
|
|
|
# Run specific test file
|
|
pytest tests/test_bot.py -v
|
|
|
|
# Run specific test
|
|
pytest tests/test_bot.py::test_mytest -v
|
|
|
|
# Run with coverage
|
|
pytest tests/ --cov=steward --cov-report=html
|
|
|
|
# Run linting only (pre-commit does this automatically)
|
|
ruff check steward/ tests/
|
|
|
|
# Auto-fix linting issues
|
|
ruff check steward/ tests/ --fix
|
|
ruff format steward/ tests/
|
|
```
|
|
|
|
Note: Pre-commit will run these checks automatically on commit, so you usually don't need to run them manually unless you want to check before committing.
|
|
|
|
### Debugging Issues
|
|
|
|
**Bot not responding:**
|
|
1. Check `.env` file has all required settings
|
|
2. Verify bot has message access in BotFather
|
|
3. Check logs: `docker compose logs steward -f`
|
|
4. Verify group ID format (negative for groups: `-1005308306472`)
|
|
|
|
**Tests failing:**
|
|
1. Run `ruff check steward/ tests/` to check linting
|
|
2. Run `mypy steward/` for type checking
|
|
3. Check test fixture setup in `conftest.py`
|
|
4. Review error messages carefully for import/config issues
|
|
|
|
**Configuration issues:**
|
|
1. Check `.env` file syntax (should be `KEY=value`)
|
|
2. For lists, prefer CSV format: `STEWARD__TELEGRAM__ALLOWED_USER_IDS=123,456` (JSON lists are also supported)
|
|
3. New format variables use `STEWARD__SECTION__KEY` pattern
|
|
4. Check `CONFIGURATION.md` for examples
|
|
|
|
## Testing Strategy
|
|
|
|
### Test Organization
|
|
|
|
- **`test_bot.py`**: Handler logic, authorization, history management
|
|
- **`test_llm_client.py`**: LLM API integration, error handling
|
|
- **`test_proposals.py`**: Proposal generation, API calls
|
|
- **`test_thread_memory.py`**: Memory storage, search, serialization
|
|
- **`test_tools.py`**: Tool calling, MCP integration
|
|
|
|
### Test Patterns
|
|
|
|
**Handler Tests:**
|
|
```python
|
|
@pytest.mark.asyncio
|
|
async def test_handler_name():
|
|
settings = _make_settings()
|
|
update = _make_update(user_id=123, text="hello")
|
|
context = _make_context(settings)
|
|
|
|
await handler_name(update, context)
|
|
|
|
# Assert on reply_text mock calls
|
|
```
|
|
|
|
**API Tests:**
|
|
```python
|
|
def test_api_call():
|
|
with respx.mock:
|
|
respx.post("https://api.example.com/endpoint").mock(
|
|
return_value=httpx.Response(200, json={"result": "ok"})
|
|
)
|
|
|
|
# Call function that makes API request
|
|
result = function_under_test()
|
|
|
|
assert result == expected_value
|
|
```
|
|
|
|
### CI/CD Pipeline
|
|
|
|
The CI workflow (`.github/workflows/ci.yml`) runs:
|
|
|
|
1. **Linting** (ruff): `python -m ruff check steward/ tests/`
|
|
2. **Tests** (pytest): `python -m pytest tests/ -v`
|
|
3. **Docker Build**: Builds and pushes image to GHCR
|
|
|
|
All three must pass for merges to main.
|
|
|
|
## Common Pitfalls
|
|
|
|
### ❌ Don't:
|
|
- Use synchronous code instead of async/await
|
|
- Hardcode secrets in config files (use env vars)
|
|
- Forget to add authorization checks to new handlers
|
|
- Write tests without using conftest fixtures
|
|
- Modify `.env` in version control (update `.env.example` instead)
|
|
- Change config structure without updating all layers (schema, pydantic, docs)
|
|
|
|
### ✅ Do:
|
|
- Use `async/await` consistently throughout
|
|
- Store secrets in environment variables
|
|
- Check both user and group authorization
|
|
- Use `pytest.mark.asyncio` for async tests
|
|
- Keep `.env` out of git (it's in `.gitignore`)
|
|
- Update schema → pydantic → properties → documentation together
|
|
|
|
## Questions? Issues?
|
|
|
|
- Check `CONFIGURATION.md` for config-related questions
|
|
- Review existing handlers in `steward/bot/telegram.py` for patterns
|
|
- Look at test examples in `tests/` for testing patterns
|
|
- Check `.github/workflows/ci.yml` for what CI expects
|
|
|
|
## Key Files to Know
|
|
|
|
| File | Purpose |
|
|
|------|---------|
|
|
| `steward/config.py` | Configuration system - OmegaConf + pydantic |
|
|
| `steward/bot/telegram.py` | All bot handlers and core logic |
|
|
| `steward/main.py` | Application entry point and setup |
|
|
| `tests/conftest.py` | Pytest fixtures and helpers |
|
|
| `pyproject.toml` | Dependencies and tool configuration |
|
|
| `CONFIGURATION.md` | User-facing config guide |
|
|
| `.github/workflows/ci.yml` | CI/CD pipeline definition |
|