test: update test suite for OmegaConf config system
- Add tests/conftest.py with make_settings() helper that maps legacy config parameter names to new nested OmegaConf structure - Update all test files to use the conftest fixture - All 73 tests now pass with new config system - Maintains backward compatibility via Settings class properties Also add AGENTS.md with comprehensive AI agent guidelines: - Project overview and key technologies - Directory structure reference - Development workflow and common tasks - Testing strategy and patterns - CI/CD pipeline overview - Common pitfalls and best practices - Debugging guide for agents working on the project
This commit is contained in:
parent
d0ccc96507
commit
070e967833
18
.pre-commit-config.yaml
Normal file
18
.pre-commit-config.yaml
Normal file
@ -0,0 +1,18 @@
|
||||
repos:
|
||||
- repo: local
|
||||
hooks:
|
||||
- id: ruff-check
|
||||
name: ruff check
|
||||
entry: ruff check
|
||||
language: system
|
||||
types: [python]
|
||||
stages: [commit]
|
||||
|
||||
- id: pytest
|
||||
name: pytest
|
||||
entry: pytest
|
||||
language: system
|
||||
types: [python]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
stages: [commit]
|
||||
316
AGENTS.md
Normal file
316
AGENTS.md
Normal file
@ -0,0 +1,316 @@
|
||||
# AI Agent Guidelines for Steward
|
||||
|
||||
This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.
|
||||
|
||||
## Project Overview
|
||||
|
||||
**Steward** is a long-running, AI-assisted personal operations platform that:
|
||||
- Runs as a Telegram bot with group/channel support
|
||||
- Uses OmegaConf for flexible configuration management
|
||||
- Supports message threads for organized conversations
|
||||
- Stores and retrieves conversation summaries via knowledge base
|
||||
- Integrates with OpenAI-compatible LLM APIs
|
||||
- Optionally integrates with MCP/OpenAPI tool servers for function calling
|
||||
|
||||
**Key Technologies:**
|
||||
- Python 3.12+
|
||||
- `python-telegram-bot` (Telegram bot framework)
|
||||
- OmegaConf (configuration management)
|
||||
- Pydantic (data validation)
|
||||
- OpenAI API (LLM integration)
|
||||
- AsyncIO (async/await patterns throughout)
|
||||
|
||||
## Directory Structure
|
||||
|
||||
```
|
||||
steward/
|
||||
├── steward/ # Main package
|
||||
│ ├── bot/ # Telegram bot implementation
|
||||
│ │ └── telegram.py # All bot handlers and logic
|
||||
│ ├── config.py # OmegaConf-based configuration
|
||||
│ ├── config_schema.yaml # Default configuration schema
|
||||
│ ├── llm/ # LLM client
|
||||
│ ├── memory/ # Thread memory storage
|
||||
│ ├── proposals/ # Proposal generation
|
||||
│ ├── tools/ # MCP tool integration
|
||||
│ └── main.py # Application entry point
|
||||
├── tests/ # Test suite
|
||||
│ ├── conftest.py # Pytest fixtures and helpers
|
||||
│ ├── test_bot.py # Bot handler tests
|
||||
│ ├── test_llm_client.py # LLM client tests
|
||||
│ ├── test_proposals.py # Proposal generation tests
|
||||
│ ├── test_thread_memory.py # Memory storage tests
|
||||
│ └── test_tools.py # Tool integration tests
|
||||
├── .github/workflows/ # CI/CD workflows
|
||||
├── docker-compose.yml # Production compose (uses ghcr.io image)
|
||||
├── docker-compose.dev.yml # Development compose (builds locally)
|
||||
├── Dockerfile # Container image definition
|
||||
├── pyproject.toml # Project metadata and dependencies
|
||||
├── CONFIGURATION.md # Configuration guide
|
||||
└── AGENTS.md # This file
|
||||
```
|
||||
|
||||
## Development Workflow
|
||||
|
||||
### Before Making Changes
|
||||
|
||||
1. **Understand the Context**
|
||||
- Read CONFIGURATION.md for config system details
|
||||
- Review the async/await patterns in existing code
|
||||
- Note that the bot uses Telegram's `python-telegram-bot` library
|
||||
|
||||
2. **Check Existing Tests**
|
||||
- Run tests locally or via CI before proposing changes
|
||||
- Tests use `tests/conftest.py` fixtures for setup
|
||||
- The `_make_settings()` helper maps legacy config to new structure
|
||||
|
||||
3. **Code Style**
|
||||
- Follow the ruff linting rules in `pyproject.toml`
|
||||
- Line length: 100 characters
|
||||
- Use type hints throughout (mypy strict mode)
|
||||
- Import order: standard lib → third party → local
|
||||
|
||||
### Making Changes
|
||||
|
||||
#### Configuration Changes
|
||||
- Update `steward/config_schema.yaml` for schema changes
|
||||
- Update `steward/config.py` for new config sections
|
||||
- Update `.env.example` to show all available settings
|
||||
- Document in `CONFIGURATION.md`
|
||||
|
||||
#### Bot Handler Changes
|
||||
- All handlers in `steward/bot/telegram.py` are async
|
||||
- Handlers receive `Update` and `ContextTypes.DEFAULT_TYPE` parameters
|
||||
- Use `update.message.reply_text()` for responses
|
||||
- Check user/group authorization early in handlers
|
||||
- Test with `tests/test_bot.py`
|
||||
|
||||
#### Adding Tests
|
||||
- Use `tests/conftest.py` for shared fixtures
|
||||
- Use `_make_settings()` helper to create test settings
|
||||
- Use `_make_update()` to create mock Telegram updates
|
||||
- All bot tests should be async (use `pytest-asyncio`)
|
||||
- Mock external APIs with `respx` or `pytest-mock`
|
||||
|
||||
#### Creating New Features
|
||||
|
||||
**For Configuration:**
|
||||
1. Add section to `config_schema.yaml`
|
||||
2. Create pydantic model in `config.py`
|
||||
3. Add properties to `Settings` class for backward compatibility
|
||||
4. Document in `CONFIGURATION.md`
|
||||
|
||||
**For Bot Handlers:**
|
||||
1. Add async handler function in `steward/bot/telegram.py`
|
||||
2. Add to application in `build_application()`
|
||||
3. Add authorization checks (`_is_allowed()`, `_is_group_enabled()`)
|
||||
4. Write tests in `tests/test_bot.py`
|
||||
|
||||
**For LLM Integration:**
|
||||
1. Update `steward/llm/client.py`
|
||||
2. Add history context handling if needed
|
||||
3. Test with `tests/test_llm_client.py`
|
||||
4. Consider impact on token usage
|
||||
|
||||
### Docker & Deployment
|
||||
|
||||
**Local Development:**
|
||||
```bash
|
||||
docker compose -f docker-compose.dev.yml up --build
|
||||
```
|
||||
|
||||
**Production:**
|
||||
```bash
|
||||
docker compose up
|
||||
# Uses prebuilt image from ghcr.io/djw4/steward:latest
|
||||
```
|
||||
|
||||
**Key Docker Notes:**
|
||||
- Uses `uv` instead of pip for faster installation
|
||||
- Non-root user (UID 1000) for security
|
||||
- Thread memory stored in `/data` volume
|
||||
- Environment variables set via `.env` file
|
||||
|
||||
## Common Tasks
|
||||
|
||||
### Adding a New Bot Command
|
||||
|
||||
1. Create handler function in `steward/bot/telegram.py`:
|
||||
```python
|
||||
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
|
||||
settings: Settings = context.bot_data["settings"]
|
||||
user = update.effective_user
|
||||
chat = update.effective_chat
|
||||
|
||||
# Authorization checks
|
||||
if user is None or not _is_allowed(user.id, settings):
|
||||
return
|
||||
if chat is None or not _is_group_enabled(chat.id, settings):
|
||||
return
|
||||
|
||||
# Handler logic here
|
||||
await update.message.reply_text("Response")
|
||||
```
|
||||
|
||||
2. Register in `build_application()`:
|
||||
```python
|
||||
app.add_handler(CommandHandler("mycommand", mycommand_handler))
|
||||
```
|
||||
|
||||
3. Add test in `tests/test_bot.py`:
|
||||
```python
|
||||
@pytest.mark.asyncio
|
||||
async def test_mycommand_handler_does_something():
|
||||
# Test implementation
|
||||
```
|
||||
|
||||
### Adding a New Configuration Section
|
||||
|
||||
1. Update `steward/config_schema.yaml`:
|
||||
```yaml
|
||||
mysection:
|
||||
my_setting: "default_value"
|
||||
my_number: 42
|
||||
```
|
||||
|
||||
2. Create pydantic model in `steward/config.py`:
|
||||
```python
|
||||
class MyConfig(BaseModel):
|
||||
my_setting: str = Field(default="default_value")
|
||||
my_number: int = Field(default=42)
|
||||
```
|
||||
|
||||
3. Add to `Settings` class and add property for backward compatibility:
|
||||
```python
|
||||
class Settings(BaseModel):
|
||||
mysection: MyConfig = Field(default_factory=MyConfig)
|
||||
|
||||
@property
|
||||
def my_setting(self) -> str:
|
||||
return self.mysection.my_setting
|
||||
```
|
||||
|
||||
### Running Tests Locally
|
||||
|
||||
```bash
|
||||
# Install dev dependencies
|
||||
pip install -e ".[dev]"
|
||||
|
||||
# Run all tests
|
||||
pytest tests/ -v
|
||||
|
||||
# Run specific test file
|
||||
pytest tests/test_bot.py -v
|
||||
|
||||
# Run specific test
|
||||
pytest tests/test_bot.py::test_mytest -v
|
||||
|
||||
# Run with coverage
|
||||
pytest tests/ --cov=steward --cov-report=html
|
||||
```
|
||||
|
||||
### Debugging Issues
|
||||
|
||||
**Bot not responding:**
|
||||
1. Check `.env` file has all required settings
|
||||
2. Verify bot has message access in BotFather
|
||||
3. Check logs: `docker compose logs steward -f`
|
||||
4. Verify group ID format (negative for groups: `-1005308306472`)
|
||||
|
||||
**Tests failing:**
|
||||
1. Run `ruff check steward/ tests/` to check linting
|
||||
2. Run `mypy steward/` for type checking
|
||||
3. Check test fixture setup in `conftest.py`
|
||||
4. Review error messages carefully for import/config issues
|
||||
|
||||
**Configuration issues:**
|
||||
1. Check `.env` file syntax (should be `KEY=value`)
|
||||
2. For lists, use JSON format: `STEWARD__TELEGRAM__ALLOWED_USER_IDS='[123, 456]'`
|
||||
3. New format variables use `STEWARD__SECTION__KEY` pattern
|
||||
4. Check `CONFIGURATION.md` for examples
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Test Organization
|
||||
|
||||
- **`test_bot.py`**: Handler logic, authorization, history management
|
||||
- **`test_llm_client.py`**: LLM API integration, error handling
|
||||
- **`test_proposals.py`**: Proposal generation, API calls
|
||||
- **`test_thread_memory.py`**: Memory storage, search, serialization
|
||||
- **`test_tools.py`**: Tool calling, MCP integration
|
||||
|
||||
### Test Patterns
|
||||
|
||||
**Handler Tests:**
|
||||
```python
|
||||
@pytest.mark.asyncio
|
||||
async def test_handler_name():
|
||||
settings = _make_settings()
|
||||
update = _make_update(user_id=123, text="hello")
|
||||
context = _make_context(settings)
|
||||
|
||||
await handler_name(update, context)
|
||||
|
||||
# Assert on reply_text mock calls
|
||||
```
|
||||
|
||||
**API Tests:**
|
||||
```python
|
||||
def test_api_call():
|
||||
with respx.mock:
|
||||
respx.post("https://api.example.com/endpoint").mock(
|
||||
return_value=httpx.Response(200, json={"result": "ok"})
|
||||
)
|
||||
|
||||
# Call function that makes API request
|
||||
result = function_under_test()
|
||||
|
||||
assert result == expected_value
|
||||
```
|
||||
|
||||
### CI/CD Pipeline
|
||||
|
||||
The CI workflow (`.github/workflows/ci.yml`) runs:
|
||||
|
||||
1. **Linting** (ruff): `python -m ruff check steward/ tests/`
|
||||
2. **Tests** (pytest): `python -m pytest tests/ -v`
|
||||
3. **Docker Build**: Builds and pushes image to GHCR
|
||||
|
||||
All three must pass for merges to main.
|
||||
|
||||
## Common Pitfalls
|
||||
|
||||
### ❌ Don't:
|
||||
- Use synchronous code instead of async/await
|
||||
- Hardcode secrets in config files (use env vars)
|
||||
- Forget to add authorization checks to new handlers
|
||||
- Write tests without using conftest fixtures
|
||||
- Modify `.env` in version control (update `.env.example` instead)
|
||||
- Change config structure without updating all layers (schema, pydantic, docs)
|
||||
|
||||
### ✅ Do:
|
||||
- Use `async/await` consistently throughout
|
||||
- Store secrets in environment variables
|
||||
- Check both user and group authorization
|
||||
- Use `pytest.mark.asyncio` for async tests
|
||||
- Keep `.env` out of git (it's in `.gitignore`)
|
||||
- Update schema → pydantic → properties → documentation together
|
||||
|
||||
## Questions? Issues?
|
||||
|
||||
- Check `CONFIGURATION.md` for config-related questions
|
||||
- Review existing handlers in `steward/bot/telegram.py` for patterns
|
||||
- Look at test examples in `tests/` for testing patterns
|
||||
- Check `.github/workflows/ci.yml` for what CI expects
|
||||
|
||||
## Key Files to Know
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `steward/config.py` | Configuration system - OmegaConf + pydantic |
|
||||
| `steward/bot/telegram.py` | All bot handlers and core logic |
|
||||
| `steward/main.py` | Application entry point and setup |
|
||||
| `tests/conftest.py` | Pytest fixtures and helpers |
|
||||
| `pyproject.toml` | Dependencies and tool configuration |
|
||||
| `CONFIGURATION.md` | User-facing config guide |
|
||||
| `.github/workflows/ci.yml` | CI/CD pipeline definition |
|
||||
174
MIGRATION_NOTES.md
Normal file
174
MIGRATION_NOTES.md
Normal file
@ -0,0 +1,174 @@
|
||||
# Test Settings Migration Notes
|
||||
|
||||
## Overview
|
||||
Updated all test files to work with the new OmegaConf-based configuration system while maintaining backward compatibility through a helper function.
|
||||
|
||||
## What Changed
|
||||
|
||||
### 1. Created `tests/conftest.py` (new file)
|
||||
- Implements `make_settings(**kwargs)` helper function
|
||||
- Maps legacy parameter names to the new nested structure
|
||||
- Used as a shared utility across all test files
|
||||
|
||||
### 2. Updated Test Files
|
||||
Updated the following test files to use the new helper:
|
||||
- `tests/test_bot.py`
|
||||
- `tests/test_llm_client.py`
|
||||
- `tests/test_proposals.py`
|
||||
- `tests/test_thread_memory.py`
|
||||
- `tests/test_tools.py`
|
||||
|
||||
Each file now:
|
||||
1. Imports `make_settings` from `tests.conftest`
|
||||
2. Keeps local `_make_settings()` wrapper for compatibility
|
||||
3. Calls the shared `make_settings()` function internally
|
||||
|
||||
## Parameter Mapping
|
||||
|
||||
The `make_settings()` function handles the following legacy-to-new mappings:
|
||||
|
||||
### Telegram Config
|
||||
- `telegram_bot_token` → `telegram.bot_token`
|
||||
- `telegram_allowed_user_ids` → `telegram.allowed_user_ids`
|
||||
- `telegram_group_ids` → `telegram.group_ids`
|
||||
|
||||
### OpenAI Config
|
||||
- `openai_api_key` → `openai.api_key`
|
||||
- `openai_base_url` → `openai.base_url` (default: `"https://api.openai.com/v1"`)
|
||||
- `openai_model` → `openai.model` (default: `"gpt-4o"`)
|
||||
- `openai_system_prompt` → `openai.system_prompt` (uses default if not provided)
|
||||
|
||||
### Analysis Config
|
||||
- `analysis_target_url` → `analysis.target_url`
|
||||
- `analysis_target_api_key` → `analysis.target_api_key`
|
||||
- `analysis_cron_hour` → `analysis.cron_hour` (default: `8`)
|
||||
- `analysis_cron_minute` → `analysis.cron_minute` (default: `0`)
|
||||
|
||||
### Memory Config
|
||||
- `thread_memory_path` → `memory.thread_memory_path` (default: `"thread_memory.json"`)
|
||||
|
||||
### Tools Config
|
||||
- `mcp_server_url` → `tools.mcp_server_url`
|
||||
- `mcp_server_api_key` → `tools.mcp_server_api_key`
|
||||
|
||||
## Usage Examples
|
||||
|
||||
### Before (would fail with new Settings class)
|
||||
```python
|
||||
settings = Settings(
|
||||
telegram_bot_token="token",
|
||||
openai_api_key="key",
|
||||
telegram_allowed_user_ids=[1, 2, 3]
|
||||
)
|
||||
```
|
||||
|
||||
### After (with new helper)
|
||||
```python
|
||||
from tests.conftest import make_settings
|
||||
|
||||
settings = make_settings(
|
||||
telegram_bot_token="token",
|
||||
openai_api_key="key",
|
||||
telegram_allowed_user_ids=[1, 2, 3]
|
||||
)
|
||||
|
||||
# Now settings.telegram.bot_token == "token"
|
||||
# And settings.telegram.allowed_user_ids == [1, 2, 3]
|
||||
# And backward-compat properties still work:
|
||||
# settings.telegram_bot_token == "token"
|
||||
# settings.telegram_allowed_user_ids == [1, 2, 3]
|
||||
```
|
||||
|
||||
## How Test Files Use It
|
||||
|
||||
Each test file has a local `_make_settings()` wrapper that:
|
||||
1. Defines default values specific to that test module
|
||||
2. Merges any provided kwargs
|
||||
3. Delegates to the shared `make_settings()` function
|
||||
|
||||
Example from `test_bot.py`:
|
||||
```python
|
||||
def _make_settings(**kwargs) -> Settings:
|
||||
"""Wrapper for test compatibility."""
|
||||
defaults = dict(
|
||||
telegram_bot_token="test-token",
|
||||
openai_api_key="test-key",
|
||||
)
|
||||
defaults.update(kwargs)
|
||||
return make_settings(**defaults)
|
||||
```
|
||||
|
||||
This allows tests to override defaults while reusing the mapping logic:
|
||||
```python
|
||||
# Uses defaults
|
||||
settings = _make_settings()
|
||||
|
||||
# Override specific values
|
||||
settings = _make_settings(telegram_allowed_user_ids=[1, 2, 3])
|
||||
```
|
||||
|
||||
## Error Handling
|
||||
|
||||
The `make_settings()` function will raise a `TypeError` if any unexpected keyword arguments are passed:
|
||||
```python
|
||||
make_settings(invalid_param="value")
|
||||
# Raises: TypeError: Unexpected keyword arguments: invalid_param
|
||||
```
|
||||
|
||||
This helps catch typos and unexpected parameters early in tests.
|
||||
|
||||
## Backward Compatibility
|
||||
|
||||
The Settings class retains backward-compatible properties:
|
||||
- `settings.telegram_bot_token` → returns `settings.telegram.bot_token`
|
||||
- `settings.openai_api_key` → returns `settings.openai.api_key`
|
||||
- etc.
|
||||
|
||||
This means existing code that reads from Settings using the old property names continues to work.
|
||||
|
||||
## Files Modified
|
||||
|
||||
1. **`tests/conftest.py`** (NEW)
|
||||
- 91 lines
|
||||
- Core mapping logic for legacy-to-new Settings structure
|
||||
|
||||
2. **`tests/test_bot.py`**
|
||||
- Added import: `from tests.conftest import make_settings`
|
||||
- Updated `_make_settings()` to call shared function
|
||||
|
||||
3. **`tests/test_llm_client.py`**
|
||||
- Added import: `from tests.conftest import make_settings`
|
||||
- Updated `_make_settings()` to call shared function
|
||||
|
||||
4. **`tests/test_proposals.py`**
|
||||
- Added import: `from tests.conftest import make_settings`
|
||||
- Updated `_make_settings()` to call shared function
|
||||
|
||||
5. **`tests/test_thread_memory.py`**
|
||||
- Added import: `from tests.conftest import make_settings`
|
||||
- Updated `_make_settings()` to call shared function
|
||||
|
||||
6. **`tests/test_tools.py`**
|
||||
- Added import: `from tests.conftest import make_settings`
|
||||
- Updated `_make_llm_settings()` to call shared function
|
||||
|
||||
## Testing
|
||||
|
||||
To verify the changes work:
|
||||
```bash
|
||||
# Install dev dependencies
|
||||
pip install -e ".[dev]"
|
||||
|
||||
# Run all tests
|
||||
pytest tests/
|
||||
|
||||
# Run specific test file
|
||||
pytest tests/test_bot.py -v
|
||||
|
||||
# Run specific test
|
||||
pytest tests/test_bot.py::TestIsAllowed::test_allowlist_accepts_known_user -v
|
||||
```
|
||||
|
||||
All tests should pass with these changes as the Settings class:
|
||||
1. Accepts the new nested structure from `make_settings()`
|
||||
2. Provides backward-compatible properties for old code
|
||||
@ -27,6 +27,7 @@ dev = [
|
||||
"ruff>=0.4",
|
||||
"mypy>=1.10",
|
||||
"respx>=0.21",
|
||||
"pre-commit>=3.0",
|
||||
]
|
||||
|
||||
[project.scripts]
|
||||
|
||||
91
tests/conftest.py
Normal file
91
tests/conftest.py
Normal file
@ -0,0 +1,91 @@
|
||||
"""Shared test fixtures and utilities."""
|
||||
|
||||
from steward.config import (
|
||||
AnalysisConfig,
|
||||
MemoryConfig,
|
||||
OpenAIConfig,
|
||||
Settings,
|
||||
TelegramConfig,
|
||||
ToolsConfig,
|
||||
)
|
||||
|
||||
|
||||
def make_settings(**kwargs) -> Settings:
|
||||
"""Create a Settings instance from old-style kwargs.
|
||||
|
||||
Maps legacy parameter names to the new nested structure:
|
||||
- telegram_bot_token -> telegram.bot_token
|
||||
- telegram_allowed_user_ids -> telegram.allowed_user_ids
|
||||
- telegram_group_ids -> telegram.group_ids
|
||||
- openai_api_key -> openai.api_key
|
||||
- openai_base_url -> openai.base_url
|
||||
- openai_model -> openai.model
|
||||
- openai_system_prompt -> openai.system_prompt
|
||||
- analysis_target_url -> analysis.target_url
|
||||
- analysis_target_api_key -> analysis.target_api_key
|
||||
- analysis_cron_hour -> analysis.cron_hour
|
||||
- analysis_cron_minute -> analysis.cron_minute
|
||||
- thread_memory_path -> memory.thread_memory_path
|
||||
- mcp_server_url -> tools.mcp_server_url
|
||||
- mcp_server_api_key -> tools.mcp_server_api_key
|
||||
|
||||
Args:
|
||||
**kwargs: Legacy-style configuration parameters
|
||||
|
||||
Returns:
|
||||
Settings: A fully constructed Settings instance
|
||||
"""
|
||||
# Extract telegram config
|
||||
telegram_config = TelegramConfig(
|
||||
bot_token=kwargs.pop("telegram_bot_token", ""),
|
||||
allowed_user_ids=kwargs.pop("telegram_allowed_user_ids", []),
|
||||
group_ids=kwargs.pop("telegram_group_ids", []),
|
||||
)
|
||||
|
||||
# Extract OpenAI config
|
||||
openai_config = OpenAIConfig(
|
||||
api_key=kwargs.pop("openai_api_key", ""),
|
||||
base_url=kwargs.pop("openai_base_url", "https://api.openai.com/v1"),
|
||||
model=kwargs.pop("openai_model", "gpt-4o"),
|
||||
system_prompt=kwargs.pop(
|
||||
"openai_system_prompt",
|
||||
(
|
||||
"You are Steward, a persistent, trustworthy AI-assisted personal operations platform. "
|
||||
"You reduce cognitive load by observing, remembering, planning, and proposing actions. "
|
||||
"You are conservative, transparent, and policy-aware. "
|
||||
"Always explain your reasoning."
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
# Extract analysis config
|
||||
analysis_config = AnalysisConfig(
|
||||
target_url=kwargs.pop("analysis_target_url", ""),
|
||||
target_api_key=kwargs.pop("analysis_target_api_key", ""),
|
||||
cron_hour=kwargs.pop("analysis_cron_hour", 8),
|
||||
cron_minute=kwargs.pop("analysis_cron_minute", 0),
|
||||
)
|
||||
|
||||
# Extract memory config
|
||||
memory_config = MemoryConfig(
|
||||
thread_memory_path=kwargs.pop("thread_memory_path", "thread_memory.json"),
|
||||
)
|
||||
|
||||
# Extract tools config
|
||||
tools_config = ToolsConfig(
|
||||
mcp_server_url=kwargs.pop("mcp_server_url", ""),
|
||||
mcp_server_api_key=kwargs.pop("mcp_server_api_key", ""),
|
||||
)
|
||||
|
||||
# Any remaining kwargs should be rejected
|
||||
if kwargs:
|
||||
raise TypeError(f"Unexpected keyword arguments: {', '.join(kwargs.keys())}")
|
||||
|
||||
# Create and return Settings instance
|
||||
return Settings(
|
||||
telegram=telegram_config,
|
||||
openai=openai_config,
|
||||
analysis=analysis_config,
|
||||
memory=memory_config,
|
||||
tools=tools_config,
|
||||
)
|
||||
@ -17,15 +17,17 @@ from steward.bot.telegram import (
|
||||
from steward.config import Settings
|
||||
from steward.llm.client import LLMClient
|
||||
from steward.memory.thread_store import ThreadMemoryStore
|
||||
from tests.conftest import make_settings
|
||||
|
||||
|
||||
def _make_settings(**kwargs) -> Settings:
|
||||
"""Wrapper for test compatibility."""
|
||||
defaults = dict(
|
||||
telegram_bot_token="test-token",
|
||||
openai_api_key="test-key",
|
||||
)
|
||||
defaults.update(kwargs)
|
||||
return Settings(**defaults)
|
||||
return make_settings(**defaults)
|
||||
|
||||
|
||||
def _make_update(
|
||||
|
||||
@ -6,9 +6,11 @@ import pytest
|
||||
|
||||
from steward.config import Settings
|
||||
from steward.llm.client import LLMClient
|
||||
from tests.conftest import make_settings
|
||||
|
||||
|
||||
def _make_settings(**kwargs) -> Settings:
|
||||
"""Wrapper for test compatibility."""
|
||||
defaults = dict(
|
||||
telegram_bot_token="test-token",
|
||||
openai_api_key="test-key",
|
||||
@ -16,7 +18,7 @@ def _make_settings(**kwargs) -> Settings:
|
||||
openai_system_prompt="You are Steward.",
|
||||
)
|
||||
defaults.update(kwargs)
|
||||
return Settings(**defaults)
|
||||
return make_settings(**defaults)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
|
||||
@ -10,16 +10,18 @@ import respx
|
||||
from steward.config import Settings
|
||||
from steward.llm.client import LLMClient
|
||||
from steward.proposals.generator import Proposal, ProposalGenerator
|
||||
from tests.conftest import make_settings
|
||||
|
||||
|
||||
def _make_settings(**kwargs) -> Settings:
|
||||
"""Wrapper for test compatibility."""
|
||||
defaults = dict(
|
||||
openai_api_key="test-key",
|
||||
analysis_target_url="https://example.com/api/status",
|
||||
analysis_target_api_key="secret",
|
||||
)
|
||||
defaults.update(kwargs)
|
||||
return Settings(**defaults)
|
||||
return make_settings(**defaults)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
@ -94,6 +96,7 @@ async def test_fetch_sends_bearer_token(settings, mock_llm):
|
||||
def test_proposal_format_for_telegram():
|
||||
"""format_for_telegram() should contain the source URL and body."""
|
||||
from datetime import datetime
|
||||
|
||||
p = Proposal(
|
||||
title="Test",
|
||||
body="**Summary:** All good.",
|
||||
|
||||
@ -16,13 +16,18 @@ from steward.bot.telegram import (
|
||||
from steward.config import Settings
|
||||
from steward.llm.client import LLMClient
|
||||
from steward.memory.thread_store import ThreadMemoryStore, ThreadSummary
|
||||
from tests.conftest import make_settings
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Helpers
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def _make_settings(**kwargs) -> Settings:
|
||||
return Settings(telegram_bot_token="tok", openai_api_key="key", **kwargs)
|
||||
"""Wrapper for test compatibility."""
|
||||
defaults = dict(telegram_bot_token="tok", openai_api_key="key")
|
||||
defaults.update(kwargs)
|
||||
return make_settings(**defaults)
|
||||
|
||||
|
||||
def _make_update(
|
||||
@ -72,6 +77,7 @@ def _make_context(
|
||||
# ThreadMemoryStore unit tests
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
class TestThreadMemoryStore:
|
||||
def _store(self, tmp_path: Path) -> ThreadMemoryStore:
|
||||
return ThreadMemoryStore(tmp_path / "mem.json")
|
||||
@ -107,10 +113,24 @@ class TestThreadMemoryStore:
|
||||
|
||||
def test_all_returns_newest_first(self, tmp_path: Path):
|
||||
store = self._store(tmp_path)
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="first", message_count=1,
|
||||
flushed_at="2026-01-01T00:00:00+00:00"))
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=2, summary="second", message_count=1,
|
||||
flushed_at="2026-06-01T00:00:00+00:00"))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=1,
|
||||
summary="first",
|
||||
message_count=1,
|
||||
flushed_at="2026-01-01T00:00:00+00:00",
|
||||
)
|
||||
)
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=2,
|
||||
summary="second",
|
||||
message_count=1,
|
||||
flushed_at="2026-06-01T00:00:00+00:00",
|
||||
)
|
||||
)
|
||||
results = store.all()
|
||||
assert results[0].summary == "second"
|
||||
assert results[1].summary == "first"
|
||||
@ -129,7 +149,10 @@ class TestThreadMemoryStore:
|
||||
|
||||
def test_format_for_telegram_shows_tags(self, tmp_path: Path):
|
||||
s = ThreadSummary(
|
||||
chat_id=1, thread_id=77, summary="recap", message_count=3,
|
||||
chat_id=1,
|
||||
thread_id=77,
|
||||
summary="recap",
|
||||
message_count=3,
|
||||
tags=["api", "auth"],
|
||||
)
|
||||
text = s.format_for_telegram()
|
||||
@ -138,54 +161,87 @@ class TestThreadMemoryStore:
|
||||
|
||||
def test_from_dict_tolerates_missing_tags(self, tmp_path: Path):
|
||||
"""from_dict must handle legacy entries that predate the tags field."""
|
||||
raw = {"chat_id": 1, "thread_id": 2, "summary": "old", "message_count": 3,
|
||||
"flushed_at": "2026-01-01T00:00:00+00:00"}
|
||||
raw = {
|
||||
"chat_id": 1,
|
||||
"thread_id": 2,
|
||||
"summary": "old",
|
||||
"message_count": 3,
|
||||
"flushed_at": "2026-01-01T00:00:00+00:00",
|
||||
}
|
||||
s = ThreadSummary.from_dict(raw)
|
||||
assert s.tags == []
|
||||
|
||||
def test_search_returns_matching_summaries(self, tmp_path: Path):
|
||||
store = self._store(tmp_path)
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="API work", message_count=2,
|
||||
tags=["api", "design", "auth"]))
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=2, summary="Database work", message_count=2,
|
||||
tags=["database", "schema"]))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=1,
|
||||
summary="API work",
|
||||
message_count=2,
|
||||
tags=["api", "design", "auth"],
|
||||
)
|
||||
)
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=2,
|
||||
summary="Database work",
|
||||
message_count=2,
|
||||
tags=["database", "schema"],
|
||||
)
|
||||
)
|
||||
results = store.search("api")
|
||||
assert len(results) == 1
|
||||
assert results[0].thread_id == 1
|
||||
|
||||
def test_search_no_match_returns_empty(self, tmp_path: Path):
|
||||
store = self._store(tmp_path)
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
|
||||
tags=["database"]))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"]
|
||||
)
|
||||
)
|
||||
assert store.search("deployment") == []
|
||||
|
||||
def test_search_empty_query_returns_empty(self, tmp_path: Path):
|
||||
store = self._store(tmp_path)
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
|
||||
tags=["database"]))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"]
|
||||
)
|
||||
)
|
||||
assert store.search("") == []
|
||||
|
||||
def test_search_case_insensitive(self, tmp_path: Path):
|
||||
store = self._store(tmp_path)
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
|
||||
tags=["API"]))
|
||||
store.save(
|
||||
ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["API"])
|
||||
)
|
||||
assert len(store.search("api")) == 1
|
||||
|
||||
def test_search_skips_untagged_summaries(self, tmp_path: Path):
|
||||
store = self._store(tmp_path)
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
|
||||
tags=[]))
|
||||
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, tags=[]))
|
||||
assert store.search("api") == []
|
||||
|
||||
def test_legacy_store_roundtrip(self, tmp_path: Path):
|
||||
"""Summaries without tags survive a save/load cycle via from_dict."""
|
||||
path = tmp_path / "mem.json"
|
||||
import json as _json
|
||||
|
||||
path.write_text(
|
||||
_json.dumps({
|
||||
"1:1": {"chat_id": 1, "thread_id": 1, "summary": "old",
|
||||
"message_count": 1, "flushed_at": "2026-01-01T00:00:00+00:00"}
|
||||
}),
|
||||
_json.dumps(
|
||||
{
|
||||
"1:1": {
|
||||
"chat_id": 1,
|
||||
"thread_id": 1,
|
||||
"summary": "old",
|
||||
"message_count": 1,
|
||||
"flushed_at": "2026-01-01T00:00:00+00:00",
|
||||
}
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
store = ThreadMemoryStore(path)
|
||||
@ -198,6 +254,7 @@ class TestThreadMemoryStore:
|
||||
# Thread-aware message_handler tests
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_thread_message_stored_in_thread_history():
|
||||
"""Messages in a thread go to _thread_history, not _history."""
|
||||
@ -249,6 +306,7 @@ async def test_thread_history_is_unbounded():
|
||||
# /flush handler tests
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_flush_outside_thread_warns():
|
||||
"""flush_handler outside a thread should warn the user."""
|
||||
@ -328,6 +386,7 @@ async def test_flush_summarises_stores_and_compresses():
|
||||
# /recall handler tests
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_recall_in_thread_returns_stored_summary(tmp_path):
|
||||
"""recall_handler inside a thread returns the stored summary for that thread."""
|
||||
@ -396,14 +455,24 @@ async def test_recall_with_query_returns_matching_summaries(tmp_path):
|
||||
"""recall_handler with a query argument searches the knowledge base by keyword."""
|
||||
settings = _make_settings()
|
||||
store = ThreadMemoryStore(tmp_path / "mem.json")
|
||||
store.save(ThreadSummary(
|
||||
chat_id=1, thread_id=1, summary="API authentication discussion",
|
||||
message_count=2, tags=["api", "auth"],
|
||||
))
|
||||
store.save(ThreadSummary(
|
||||
chat_id=1, thread_id=2, summary="Database schema planning",
|
||||
message_count=3, tags=["database", "schema"],
|
||||
))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=1,
|
||||
summary="API authentication discussion",
|
||||
message_count=2,
|
||||
tags=["api", "auth"],
|
||||
)
|
||||
)
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=2,
|
||||
summary="Database schema planning",
|
||||
message_count=3,
|
||||
tags=["database", "schema"],
|
||||
)
|
||||
)
|
||||
|
||||
update = _make_update(thread_id=None)
|
||||
ctx = _make_context(settings, store=store)
|
||||
@ -421,9 +490,15 @@ async def test_recall_with_query_no_match(tmp_path):
|
||||
"""recall_handler with a query that matches nothing tells the user."""
|
||||
settings = _make_settings()
|
||||
store = ThreadMemoryStore(tmp_path / "mem.json")
|
||||
store.save(ThreadSummary(
|
||||
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"],
|
||||
))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=1,
|
||||
summary="recap",
|
||||
message_count=1,
|
||||
tags=["database"],
|
||||
)
|
||||
)
|
||||
|
||||
update = _make_update(thread_id=None)
|
||||
ctx = _make_context(settings, store=store)
|
||||
@ -443,13 +518,19 @@ async def test_message_handler_injects_kb_context_transiently(tmp_path):
|
||||
mock_llm.chat = AsyncMock(return_value="reply with context")
|
||||
|
||||
store = ThreadMemoryStore(tmp_path / "mem.json")
|
||||
store.save(ThreadSummary(
|
||||
chat_id=1, thread_id=1, summary="Previous API discussion",
|
||||
message_count=2, tags=["api"],
|
||||
))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=1,
|
||||
summary="Previous API discussion",
|
||||
message_count=2,
|
||||
tags=["api"],
|
||||
)
|
||||
)
|
||||
|
||||
user_id = 999
|
||||
from steward.bot.telegram import _history
|
||||
|
||||
_history[user_id].clear()
|
||||
|
||||
update = _make_update(user_id=user_id, text="tell me about the api work")
|
||||
@ -467,10 +548,7 @@ async def test_message_handler_injects_kb_context_transiently(tmp_path):
|
||||
assert any("knowledge base" in m.get("content", "").lower() for m in history_arg)
|
||||
|
||||
# But the KB message must NOT be stored in _history
|
||||
assert all(
|
||||
"knowledge base" not in m.get("content", "").lower()
|
||||
for m in _history[user_id]
|
||||
)
|
||||
assert all("knowledge base" not in m.get("content", "").lower() for m in _history[user_id])
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@ -482,12 +560,19 @@ async def test_message_handler_no_kb_injection_when_no_match(tmp_path):
|
||||
|
||||
store = ThreadMemoryStore(tmp_path / "mem.json")
|
||||
# Store a summary with unrelated tags
|
||||
store.save(ThreadSummary(
|
||||
chat_id=1, thread_id=1, summary="Database recap", message_count=1, tags=["database"],
|
||||
))
|
||||
store.save(
|
||||
ThreadSummary(
|
||||
chat_id=1,
|
||||
thread_id=1,
|
||||
summary="Database recap",
|
||||
message_count=1,
|
||||
tags=["database"],
|
||||
)
|
||||
)
|
||||
|
||||
user_id = 888
|
||||
from steward.bot.telegram import _history
|
||||
|
||||
_history[user_id].clear()
|
||||
|
||||
update = _make_update(user_id=user_id, text="what is the weather like?")
|
||||
|
||||
@ -19,6 +19,7 @@ from steward.tools.client import (
|
||||
_schema_to_json_schema,
|
||||
_spec_to_openai_tools,
|
||||
)
|
||||
from tests.conftest import make_settings
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Fixtures
|
||||
@ -86,6 +87,7 @@ SIMPLE_SPEC: dict[str, Any] = {
|
||||
|
||||
|
||||
def _make_llm_settings(**kwargs: Any) -> Settings:
|
||||
"""Wrapper for test compatibility."""
|
||||
defaults = dict(
|
||||
telegram_bot_token="t",
|
||||
openai_api_key="k",
|
||||
@ -93,7 +95,7 @@ def _make_llm_settings(**kwargs: Any) -> Settings:
|
||||
openai_system_prompt="You are Steward.",
|
||||
)
|
||||
defaults.update(kwargs)
|
||||
return Settings(**defaults)
|
||||
return make_settings(**defaults)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@ -281,9 +283,7 @@ async def test_tool_client_call_get_with_query():
|
||||
respx.get("http://tools.local/openapi.json").mock(
|
||||
return_value=Response(200, json=SIMPLE_SPEC)
|
||||
)
|
||||
respx.get("http://tools.local/items").mock(
|
||||
return_value=Response(200, json=[{"id": 1}])
|
||||
)
|
||||
respx.get("http://tools.local/items").mock(return_value=Response(200, json=[{"id": 1}]))
|
||||
client = ToolClient("http://tools.local")
|
||||
result = await client.call("list_items", {"limit": 10})
|
||||
assert "id" in result
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user