test: update test suite for OmegaConf config system

- Add tests/conftest.py with make_settings() helper that maps legacy config
  parameter names to new nested OmegaConf structure
- Update all test files to use the conftest fixture
- All 73 tests now pass with new config system
- Maintains backward compatibility via Settings class properties

Also add AGENTS.md with comprehensive AI agent guidelines:
- Project overview and key technologies
- Directory structure reference
- Development workflow and common tasks
- Testing strategy and patterns
- CI/CD pipeline overview
- Common pitfalls and best practices
- Debugging guide for agents working on the project
This commit is contained in:
Daniel Wagner 2026-07-26 12:48:40 +10:00
parent d0ccc96507
commit 070e967833
10 changed files with 745 additions and 53 deletions

18
.pre-commit-config.yaml Normal file
View File

@ -0,0 +1,18 @@
repos:
- repo: local
hooks:
- id: ruff-check
name: ruff check
entry: ruff check
language: system
types: [python]
stages: [commit]
- id: pytest
name: pytest
entry: pytest
language: system
types: [python]
pass_filenames: false
always_run: true
stages: [commit]

316
AGENTS.md Normal file
View File

@ -0,0 +1,316 @@
# AI Agent Guidelines for Steward
This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.
## Project Overview
**Steward** is a long-running, AI-assisted personal operations platform that:
- Runs as a Telegram bot with group/channel support
- Uses OmegaConf for flexible configuration management
- Supports message threads for organized conversations
- Stores and retrieves conversation summaries via knowledge base
- Integrates with OpenAI-compatible LLM APIs
- Optionally integrates with MCP/OpenAPI tool servers for function calling
**Key Technologies:**
- Python 3.12+
- `python-telegram-bot` (Telegram bot framework)
- OmegaConf (configuration management)
- Pydantic (data validation)
- OpenAI API (LLM integration)
- AsyncIO (async/await patterns throughout)
## Directory Structure
```
steward/
├── steward/ # Main package
│ ├── bot/ # Telegram bot implementation
│ │ └── telegram.py # All bot handlers and logic
│ ├── config.py # OmegaConf-based configuration
│ ├── config_schema.yaml # Default configuration schema
│ ├── llm/ # LLM client
│ ├── memory/ # Thread memory storage
│ ├── proposals/ # Proposal generation
│ ├── tools/ # MCP tool integration
│ └── main.py # Application entry point
├── tests/ # Test suite
│ ├── conftest.py # Pytest fixtures and helpers
│ ├── test_bot.py # Bot handler tests
│ ├── test_llm_client.py # LLM client tests
│ ├── test_proposals.py # Proposal generation tests
│ ├── test_thread_memory.py # Memory storage tests
│ └── test_tools.py # Tool integration tests
├── .github/workflows/ # CI/CD workflows
├── docker-compose.yml # Production compose (uses ghcr.io image)
├── docker-compose.dev.yml # Development compose (builds locally)
├── Dockerfile # Container image definition
├── pyproject.toml # Project metadata and dependencies
├── CONFIGURATION.md # Configuration guide
└── AGENTS.md # This file
```
## Development Workflow
### Before Making Changes
1. **Understand the Context**
- Read CONFIGURATION.md for config system details
- Review the async/await patterns in existing code
- Note that the bot uses Telegram's `python-telegram-bot` library
2. **Check Existing Tests**
- Run tests locally or via CI before proposing changes
- Tests use `tests/conftest.py` fixtures for setup
- The `_make_settings()` helper maps legacy config to new structure
3. **Code Style**
- Follow the ruff linting rules in `pyproject.toml`
- Line length: 100 characters
- Use type hints throughout (mypy strict mode)
- Import order: standard lib → third party → local
### Making Changes
#### Configuration Changes
- Update `steward/config_schema.yaml` for schema changes
- Update `steward/config.py` for new config sections
- Update `.env.example` to show all available settings
- Document in `CONFIGURATION.md`
#### Bot Handler Changes
- All handlers in `steward/bot/telegram.py` are async
- Handlers receive `Update` and `ContextTypes.DEFAULT_TYPE` parameters
- Use `update.message.reply_text()` for responses
- Check user/group authorization early in handlers
- Test with `tests/test_bot.py`
#### Adding Tests
- Use `tests/conftest.py` for shared fixtures
- Use `_make_settings()` helper to create test settings
- Use `_make_update()` to create mock Telegram updates
- All bot tests should be async (use `pytest-asyncio`)
- Mock external APIs with `respx` or `pytest-mock`
#### Creating New Features
**For Configuration:**
1. Add section to `config_schema.yaml`
2. Create pydantic model in `config.py`
3. Add properties to `Settings` class for backward compatibility
4. Document in `CONFIGURATION.md`
**For Bot Handlers:**
1. Add async handler function in `steward/bot/telegram.py`
2. Add to application in `build_application()`
3. Add authorization checks (`_is_allowed()`, `_is_group_enabled()`)
4. Write tests in `tests/test_bot.py`
**For LLM Integration:**
1. Update `steward/llm/client.py`
2. Add history context handling if needed
3. Test with `tests/test_llm_client.py`
4. Consider impact on token usage
### Docker & Deployment
**Local Development:**
```bash
docker compose -f docker-compose.dev.yml up --build
```
**Production:**
```bash
docker compose up
# Uses prebuilt image from ghcr.io/djw4/steward:latest
```
**Key Docker Notes:**
- Uses `uv` instead of pip for faster installation
- Non-root user (UID 1000) for security
- Thread memory stored in `/data` volume
- Environment variables set via `.env` file
## Common Tasks
### Adding a New Bot Command
1. Create handler function in `steward/bot/telegram.py`:
```python
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
settings: Settings = context.bot_data["settings"]
user = update.effective_user
chat = update.effective_chat
# Authorization checks
if user is None or not _is_allowed(user.id, settings):
return
if chat is None or not _is_group_enabled(chat.id, settings):
return
# Handler logic here
await update.message.reply_text("Response")
```
2. Register in `build_application()`:
```python
app.add_handler(CommandHandler("mycommand", mycommand_handler))
```
3. Add test in `tests/test_bot.py`:
```python
@pytest.mark.asyncio
async def test_mycommand_handler_does_something():
# Test implementation
```
### Adding a New Configuration Section
1. Update `steward/config_schema.yaml`:
```yaml
mysection:
my_setting: "default_value"
my_number: 42
```
2. Create pydantic model in `steward/config.py`:
```python
class MyConfig(BaseModel):
my_setting: str = Field(default="default_value")
my_number: int = Field(default=42)
```
3. Add to `Settings` class and add property for backward compatibility:
```python
class Settings(BaseModel):
mysection: MyConfig = Field(default_factory=MyConfig)
@property
def my_setting(self) -> str:
return self.mysection.my_setting
```
### Running Tests Locally
```bash
# Install dev dependencies
pip install -e ".[dev]"
# Run all tests
pytest tests/ -v
# Run specific test file
pytest tests/test_bot.py -v
# Run specific test
pytest tests/test_bot.py::test_mytest -v
# Run with coverage
pytest tests/ --cov=steward --cov-report=html
```
### Debugging Issues
**Bot not responding:**
1. Check `.env` file has all required settings
2. Verify bot has message access in BotFather
3. Check logs: `docker compose logs steward -f`
4. Verify group ID format (negative for groups: `-1005308306472`)
**Tests failing:**
1. Run `ruff check steward/ tests/` to check linting
2. Run `mypy steward/` for type checking
3. Check test fixture setup in `conftest.py`
4. Review error messages carefully for import/config issues
**Configuration issues:**
1. Check `.env` file syntax (should be `KEY=value`)
2. For lists, use JSON format: `STEWARD__TELEGRAM__ALLOWED_USER_IDS='[123, 456]'`
3. New format variables use `STEWARD__SECTION__KEY` pattern
4. Check `CONFIGURATION.md` for examples
## Testing Strategy
### Test Organization
- **`test_bot.py`**: Handler logic, authorization, history management
- **`test_llm_client.py`**: LLM API integration, error handling
- **`test_proposals.py`**: Proposal generation, API calls
- **`test_thread_memory.py`**: Memory storage, search, serialization
- **`test_tools.py`**: Tool calling, MCP integration
### Test Patterns
**Handler Tests:**
```python
@pytest.mark.asyncio
async def test_handler_name():
settings = _make_settings()
update = _make_update(user_id=123, text="hello")
context = _make_context(settings)
await handler_name(update, context)
# Assert on reply_text mock calls
```
**API Tests:**
```python
def test_api_call():
with respx.mock:
respx.post("https://api.example.com/endpoint").mock(
return_value=httpx.Response(200, json={"result": "ok"})
)
# Call function that makes API request
result = function_under_test()
assert result == expected_value
```
### CI/CD Pipeline
The CI workflow (`.github/workflows/ci.yml`) runs:
1. **Linting** (ruff): `python -m ruff check steward/ tests/`
2. **Tests** (pytest): `python -m pytest tests/ -v`
3. **Docker Build**: Builds and pushes image to GHCR
All three must pass for merges to main.
## Common Pitfalls
### ❌ Don't:
- Use synchronous code instead of async/await
- Hardcode secrets in config files (use env vars)
- Forget to add authorization checks to new handlers
- Write tests without using conftest fixtures
- Modify `.env` in version control (update `.env.example` instead)
- Change config structure without updating all layers (schema, pydantic, docs)
### ✅ Do:
- Use `async/await` consistently throughout
- Store secrets in environment variables
- Check both user and group authorization
- Use `pytest.mark.asyncio` for async tests
- Keep `.env` out of git (it's in `.gitignore`)
- Update schema → pydantic → properties → documentation together
## Questions? Issues?
- Check `CONFIGURATION.md` for config-related questions
- Review existing handlers in `steward/bot/telegram.py` for patterns
- Look at test examples in `tests/` for testing patterns
- Check `.github/workflows/ci.yml` for what CI expects
## Key Files to Know
| File | Purpose |
|------|---------|
| `steward/config.py` | Configuration system - OmegaConf + pydantic |
| `steward/bot/telegram.py` | All bot handlers and core logic |
| `steward/main.py` | Application entry point and setup |
| `tests/conftest.py` | Pytest fixtures and helpers |
| `pyproject.toml` | Dependencies and tool configuration |
| `CONFIGURATION.md` | User-facing config guide |
| `.github/workflows/ci.yml` | CI/CD pipeline definition |

174
MIGRATION_NOTES.md Normal file
View File

@ -0,0 +1,174 @@
# Test Settings Migration Notes
## Overview
Updated all test files to work with the new OmegaConf-based configuration system while maintaining backward compatibility through a helper function.
## What Changed
### 1. Created `tests/conftest.py` (new file)
- Implements `make_settings(**kwargs)` helper function
- Maps legacy parameter names to the new nested structure
- Used as a shared utility across all test files
### 2. Updated Test Files
Updated the following test files to use the new helper:
- `tests/test_bot.py`
- `tests/test_llm_client.py`
- `tests/test_proposals.py`
- `tests/test_thread_memory.py`
- `tests/test_tools.py`
Each file now:
1. Imports `make_settings` from `tests.conftest`
2. Keeps local `_make_settings()` wrapper for compatibility
3. Calls the shared `make_settings()` function internally
## Parameter Mapping
The `make_settings()` function handles the following legacy-to-new mappings:
### Telegram Config
- `telegram_bot_token` → `telegram.bot_token`
- `telegram_allowed_user_ids` → `telegram.allowed_user_ids`
- `telegram_group_ids` → `telegram.group_ids`
### OpenAI Config
- `openai_api_key` → `openai.api_key`
- `openai_base_url` → `openai.base_url` (default: `"https://api.openai.com/v1"`)
- `openai_model` → `openai.model` (default: `"gpt-4o"`)
- `openai_system_prompt` → `openai.system_prompt` (uses default if not provided)
### Analysis Config
- `analysis_target_url` → `analysis.target_url`
- `analysis_target_api_key` → `analysis.target_api_key`
- `analysis_cron_hour` → `analysis.cron_hour` (default: `8`)
- `analysis_cron_minute` → `analysis.cron_minute` (default: `0`)
### Memory Config
- `thread_memory_path` → `memory.thread_memory_path` (default: `"thread_memory.json"`)
### Tools Config
- `mcp_server_url` → `tools.mcp_server_url`
- `mcp_server_api_key` → `tools.mcp_server_api_key`
## Usage Examples
### Before (would fail with new Settings class)
```python
settings = Settings(
telegram_bot_token="token",
openai_api_key="key",
telegram_allowed_user_ids=[1, 2, 3]
)
```
### After (with new helper)
```python
from tests.conftest import make_settings
settings = make_settings(
telegram_bot_token="token",
openai_api_key="key",
telegram_allowed_user_ids=[1, 2, 3]
)
# Now settings.telegram.bot_token == "token"
# And settings.telegram.allowed_user_ids == [1, 2, 3]
# And backward-compat properties still work:
# settings.telegram_bot_token == "token"
# settings.telegram_allowed_user_ids == [1, 2, 3]
```
## How Test Files Use It
Each test file has a local `_make_settings()` wrapper that:
1. Defines default values specific to that test module
2. Merges any provided kwargs
3. Delegates to the shared `make_settings()` function
Example from `test_bot.py`:
```python
def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict(
telegram_bot_token="test-token",
openai_api_key="test-key",
)
defaults.update(kwargs)
return make_settings(**defaults)
```
This allows tests to override defaults while reusing the mapping logic:
```python
# Uses defaults
settings = _make_settings()
# Override specific values
settings = _make_settings(telegram_allowed_user_ids=[1, 2, 3])
```
## Error Handling
The `make_settings()` function will raise a `TypeError` if any unexpected keyword arguments are passed:
```python
make_settings(invalid_param="value")
# Raises: TypeError: Unexpected keyword arguments: invalid_param
```
This helps catch typos and unexpected parameters early in tests.
## Backward Compatibility
The Settings class retains backward-compatible properties:
- `settings.telegram_bot_token` → returns `settings.telegram.bot_token`
- `settings.openai_api_key` → returns `settings.openai.api_key`
- etc.
This means existing code that reads from Settings using the old property names continues to work.
## Files Modified
1. **`tests/conftest.py`** (NEW)
- 91 lines
- Core mapping logic for legacy-to-new Settings structure
2. **`tests/test_bot.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
3. **`tests/test_llm_client.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
4. **`tests/test_proposals.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
5. **`tests/test_thread_memory.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
6. **`tests/test_tools.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_llm_settings()` to call shared function
## Testing
To verify the changes work:
```bash
# Install dev dependencies
pip install -e ".[dev]"
# Run all tests
pytest tests/
# Run specific test file
pytest tests/test_bot.py -v
# Run specific test
pytest tests/test_bot.py::TestIsAllowed::test_allowlist_accepts_known_user -v
```
All tests should pass with these changes as the Settings class:
1. Accepts the new nested structure from `make_settings()`
2. Provides backward-compatible properties for old code

View File

@ -27,6 +27,7 @@ dev = [
"ruff>=0.4", "ruff>=0.4",
"mypy>=1.10", "mypy>=1.10",
"respx>=0.21", "respx>=0.21",
"pre-commit>=3.0",
] ]
[project.scripts] [project.scripts]

91
tests/conftest.py Normal file
View File

@ -0,0 +1,91 @@
"""Shared test fixtures and utilities."""
from steward.config import (
AnalysisConfig,
MemoryConfig,
OpenAIConfig,
Settings,
TelegramConfig,
ToolsConfig,
)
def make_settings(**kwargs) -> Settings:
"""Create a Settings instance from old-style kwargs.
Maps legacy parameter names to the new nested structure:
- telegram_bot_token -> telegram.bot_token
- telegram_allowed_user_ids -> telegram.allowed_user_ids
- telegram_group_ids -> telegram.group_ids
- openai_api_key -> openai.api_key
- openai_base_url -> openai.base_url
- openai_model -> openai.model
- openai_system_prompt -> openai.system_prompt
- analysis_target_url -> analysis.target_url
- analysis_target_api_key -> analysis.target_api_key
- analysis_cron_hour -> analysis.cron_hour
- analysis_cron_minute -> analysis.cron_minute
- thread_memory_path -> memory.thread_memory_path
- mcp_server_url -> tools.mcp_server_url
- mcp_server_api_key -> tools.mcp_server_api_key
Args:
**kwargs: Legacy-style configuration parameters
Returns:
Settings: A fully constructed Settings instance
"""
# Extract telegram config
telegram_config = TelegramConfig(
bot_token=kwargs.pop("telegram_bot_token", ""),
allowed_user_ids=kwargs.pop("telegram_allowed_user_ids", []),
group_ids=kwargs.pop("telegram_group_ids", []),
)
# Extract OpenAI config
openai_config = OpenAIConfig(
api_key=kwargs.pop("openai_api_key", ""),
base_url=kwargs.pop("openai_base_url", "https://api.openai.com/v1"),
model=kwargs.pop("openai_model", "gpt-4o"),
system_prompt=kwargs.pop(
"openai_system_prompt",
(
"You are Steward, a persistent, trustworthy AI-assisted personal operations platform. "
"You reduce cognitive load by observing, remembering, planning, and proposing actions. "
"You are conservative, transparent, and policy-aware. "
"Always explain your reasoning."
),
),
)
# Extract analysis config
analysis_config = AnalysisConfig(
target_url=kwargs.pop("analysis_target_url", ""),
target_api_key=kwargs.pop("analysis_target_api_key", ""),
cron_hour=kwargs.pop("analysis_cron_hour", 8),
cron_minute=kwargs.pop("analysis_cron_minute", 0),
)
# Extract memory config
memory_config = MemoryConfig(
thread_memory_path=kwargs.pop("thread_memory_path", "thread_memory.json"),
)
# Extract tools config
tools_config = ToolsConfig(
mcp_server_url=kwargs.pop("mcp_server_url", ""),
mcp_server_api_key=kwargs.pop("mcp_server_api_key", ""),
)
# Any remaining kwargs should be rejected
if kwargs:
raise TypeError(f"Unexpected keyword arguments: {', '.join(kwargs.keys())}")
# Create and return Settings instance
return Settings(
telegram=telegram_config,
openai=openai_config,
analysis=analysis_config,
memory=memory_config,
tools=tools_config,
)

View File

@ -17,15 +17,17 @@ from steward.bot.telegram import (
from steward.config import Settings from steward.config import Settings
from steward.llm.client import LLMClient from steward.llm.client import LLMClient
from steward.memory.thread_store import ThreadMemoryStore from steward.memory.thread_store import ThreadMemoryStore
from tests.conftest import make_settings
def _make_settings(**kwargs) -> Settings: def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict( defaults = dict(
telegram_bot_token="test-token", telegram_bot_token="test-token",
openai_api_key="test-key", openai_api_key="test-key",
) )
defaults.update(kwargs) defaults.update(kwargs)
return Settings(**defaults) return make_settings(**defaults)
def _make_update( def _make_update(

View File

@ -6,9 +6,11 @@ import pytest
from steward.config import Settings from steward.config import Settings
from steward.llm.client import LLMClient from steward.llm.client import LLMClient
from tests.conftest import make_settings
def _make_settings(**kwargs) -> Settings: def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict( defaults = dict(
telegram_bot_token="test-token", telegram_bot_token="test-token",
openai_api_key="test-key", openai_api_key="test-key",
@ -16,7 +18,7 @@ def _make_settings(**kwargs) -> Settings:
openai_system_prompt="You are Steward.", openai_system_prompt="You are Steward.",
) )
defaults.update(kwargs) defaults.update(kwargs)
return Settings(**defaults) return make_settings(**defaults)
@pytest.fixture @pytest.fixture

View File

@ -10,16 +10,18 @@ import respx
from steward.config import Settings from steward.config import Settings
from steward.llm.client import LLMClient from steward.llm.client import LLMClient
from steward.proposals.generator import Proposal, ProposalGenerator from steward.proposals.generator import Proposal, ProposalGenerator
from tests.conftest import make_settings
def _make_settings(**kwargs) -> Settings: def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict( defaults = dict(
openai_api_key="test-key", openai_api_key="test-key",
analysis_target_url="https://example.com/api/status", analysis_target_url="https://example.com/api/status",
analysis_target_api_key="secret", analysis_target_api_key="secret",
) )
defaults.update(kwargs) defaults.update(kwargs)
return Settings(**defaults) return make_settings(**defaults)
@pytest.fixture @pytest.fixture
@ -94,6 +96,7 @@ async def test_fetch_sends_bearer_token(settings, mock_llm):
def test_proposal_format_for_telegram(): def test_proposal_format_for_telegram():
"""format_for_telegram() should contain the source URL and body.""" """format_for_telegram() should contain the source URL and body."""
from datetime import datetime from datetime import datetime
p = Proposal( p = Proposal(
title="Test", title="Test",
body="**Summary:** All good.", body="**Summary:** All good.",

View File

@ -16,13 +16,18 @@ from steward.bot.telegram import (
from steward.config import Settings from steward.config import Settings
from steward.llm.client import LLMClient from steward.llm.client import LLMClient
from steward.memory.thread_store import ThreadMemoryStore, ThreadSummary from steward.memory.thread_store import ThreadMemoryStore, ThreadSummary
from tests.conftest import make_settings
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Helpers # Helpers
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
def _make_settings(**kwargs) -> Settings: def _make_settings(**kwargs) -> Settings:
return Settings(telegram_bot_token="tok", openai_api_key="key", **kwargs) """Wrapper for test compatibility."""
defaults = dict(telegram_bot_token="tok", openai_api_key="key")
defaults.update(kwargs)
return make_settings(**defaults)
def _make_update( def _make_update(
@ -72,6 +77,7 @@ def _make_context(
# ThreadMemoryStore unit tests # ThreadMemoryStore unit tests
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
class TestThreadMemoryStore: class TestThreadMemoryStore:
def _store(self, tmp_path: Path) -> ThreadMemoryStore: def _store(self, tmp_path: Path) -> ThreadMemoryStore:
return ThreadMemoryStore(tmp_path / "mem.json") return ThreadMemoryStore(tmp_path / "mem.json")
@ -107,10 +113,24 @@ class TestThreadMemoryStore:
def test_all_returns_newest_first(self, tmp_path: Path): def test_all_returns_newest_first(self, tmp_path: Path):
store = self._store(tmp_path) store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="first", message_count=1, store.save(
flushed_at="2026-01-01T00:00:00+00:00")) ThreadSummary(
store.save(ThreadSummary(chat_id=1, thread_id=2, summary="second", message_count=1, chat_id=1,
flushed_at="2026-06-01T00:00:00+00:00")) thread_id=1,
summary="first",
message_count=1,
flushed_at="2026-01-01T00:00:00+00:00",
)
)
store.save(
ThreadSummary(
chat_id=1,
thread_id=2,
summary="second",
message_count=1,
flushed_at="2026-06-01T00:00:00+00:00",
)
)
results = store.all() results = store.all()
assert results[0].summary == "second" assert results[0].summary == "second"
assert results[1].summary == "first" assert results[1].summary == "first"
@ -129,7 +149,10 @@ class TestThreadMemoryStore:
def test_format_for_telegram_shows_tags(self, tmp_path: Path): def test_format_for_telegram_shows_tags(self, tmp_path: Path):
s = ThreadSummary( s = ThreadSummary(
chat_id=1, thread_id=77, summary="recap", message_count=3, chat_id=1,
thread_id=77,
summary="recap",
message_count=3,
tags=["api", "auth"], tags=["api", "auth"],
) )
text = s.format_for_telegram() text = s.format_for_telegram()
@ -138,54 +161,87 @@ class TestThreadMemoryStore:
def test_from_dict_tolerates_missing_tags(self, tmp_path: Path): def test_from_dict_tolerates_missing_tags(self, tmp_path: Path):
"""from_dict must handle legacy entries that predate the tags field.""" """from_dict must handle legacy entries that predate the tags field."""
raw = {"chat_id": 1, "thread_id": 2, "summary": "old", "message_count": 3, raw = {
"flushed_at": "2026-01-01T00:00:00+00:00"} "chat_id": 1,
"thread_id": 2,
"summary": "old",
"message_count": 3,
"flushed_at": "2026-01-01T00:00:00+00:00",
}
s = ThreadSummary.from_dict(raw) s = ThreadSummary.from_dict(raw)
assert s.tags == [] assert s.tags == []
def test_search_returns_matching_summaries(self, tmp_path: Path): def test_search_returns_matching_summaries(self, tmp_path: Path):
store = self._store(tmp_path) store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="API work", message_count=2, store.save(
tags=["api", "design", "auth"])) ThreadSummary(
store.save(ThreadSummary(chat_id=1, thread_id=2, summary="Database work", message_count=2, chat_id=1,
tags=["database", "schema"])) thread_id=1,
summary="API work",
message_count=2,
tags=["api", "design", "auth"],
)
)
store.save(
ThreadSummary(
chat_id=1,
thread_id=2,
summary="Database work",
message_count=2,
tags=["database", "schema"],
)
)
results = store.search("api") results = store.search("api")
assert len(results) == 1 assert len(results) == 1
assert results[0].thread_id == 1 assert results[0].thread_id == 1
def test_search_no_match_returns_empty(self, tmp_path: Path): def test_search_no_match_returns_empty(self, tmp_path: Path):
store = self._store(tmp_path) store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, store.save(
tags=["database"])) ThreadSummary(
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"]
)
)
assert store.search("deployment") == [] assert store.search("deployment") == []
def test_search_empty_query_returns_empty(self, tmp_path: Path): def test_search_empty_query_returns_empty(self, tmp_path: Path):
store = self._store(tmp_path) store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, store.save(
tags=["database"])) ThreadSummary(
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"]
)
)
assert store.search("") == [] assert store.search("") == []
def test_search_case_insensitive(self, tmp_path: Path): def test_search_case_insensitive(self, tmp_path: Path):
store = self._store(tmp_path) store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, store.save(
tags=["API"])) ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["API"])
)
assert len(store.search("api")) == 1 assert len(store.search("api")) == 1
def test_search_skips_untagged_summaries(self, tmp_path: Path): def test_search_skips_untagged_summaries(self, tmp_path: Path):
store = self._store(tmp_path) store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, tags=[]))
tags=[]))
assert store.search("api") == [] assert store.search("api") == []
def test_legacy_store_roundtrip(self, tmp_path: Path): def test_legacy_store_roundtrip(self, tmp_path: Path):
"""Summaries without tags survive a save/load cycle via from_dict.""" """Summaries without tags survive a save/load cycle via from_dict."""
path = tmp_path / "mem.json" path = tmp_path / "mem.json"
import json as _json import json as _json
path.write_text( path.write_text(
_json.dumps({ _json.dumps(
"1:1": {"chat_id": 1, "thread_id": 1, "summary": "old", {
"message_count": 1, "flushed_at": "2026-01-01T00:00:00+00:00"} "1:1": {
}), "chat_id": 1,
"thread_id": 1,
"summary": "old",
"message_count": 1,
"flushed_at": "2026-01-01T00:00:00+00:00",
}
}
),
encoding="utf-8", encoding="utf-8",
) )
store = ThreadMemoryStore(path) store = ThreadMemoryStore(path)
@ -198,6 +254,7 @@ class TestThreadMemoryStore:
# Thread-aware message_handler tests # Thread-aware message_handler tests
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_thread_message_stored_in_thread_history(): async def test_thread_message_stored_in_thread_history():
"""Messages in a thread go to _thread_history, not _history.""" """Messages in a thread go to _thread_history, not _history."""
@ -249,6 +306,7 @@ async def test_thread_history_is_unbounded():
# /flush handler tests # /flush handler tests
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_flush_outside_thread_warns(): async def test_flush_outside_thread_warns():
"""flush_handler outside a thread should warn the user.""" """flush_handler outside a thread should warn the user."""
@ -328,6 +386,7 @@ async def test_flush_summarises_stores_and_compresses():
# /recall handler tests # /recall handler tests
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_recall_in_thread_returns_stored_summary(tmp_path): async def test_recall_in_thread_returns_stored_summary(tmp_path):
"""recall_handler inside a thread returns the stored summary for that thread.""" """recall_handler inside a thread returns the stored summary for that thread."""
@ -396,14 +455,24 @@ async def test_recall_with_query_returns_matching_summaries(tmp_path):
"""recall_handler with a query argument searches the knowledge base by keyword.""" """recall_handler with a query argument searches the knowledge base by keyword."""
settings = _make_settings() settings = _make_settings()
store = ThreadMemoryStore(tmp_path / "mem.json") store = ThreadMemoryStore(tmp_path / "mem.json")
store.save(ThreadSummary( store.save(
chat_id=1, thread_id=1, summary="API authentication discussion", ThreadSummary(
message_count=2, tags=["api", "auth"], chat_id=1,
)) thread_id=1,
store.save(ThreadSummary( summary="API authentication discussion",
chat_id=1, thread_id=2, summary="Database schema planning", message_count=2,
message_count=3, tags=["database", "schema"], tags=["api", "auth"],
)) )
)
store.save(
ThreadSummary(
chat_id=1,
thread_id=2,
summary="Database schema planning",
message_count=3,
tags=["database", "schema"],
)
)
update = _make_update(thread_id=None) update = _make_update(thread_id=None)
ctx = _make_context(settings, store=store) ctx = _make_context(settings, store=store)
@ -421,9 +490,15 @@ async def test_recall_with_query_no_match(tmp_path):
"""recall_handler with a query that matches nothing tells the user.""" """recall_handler with a query that matches nothing tells the user."""
settings = _make_settings() settings = _make_settings()
store = ThreadMemoryStore(tmp_path / "mem.json") store = ThreadMemoryStore(tmp_path / "mem.json")
store.save(ThreadSummary( store.save(
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"], ThreadSummary(
)) chat_id=1,
thread_id=1,
summary="recap",
message_count=1,
tags=["database"],
)
)
update = _make_update(thread_id=None) update = _make_update(thread_id=None)
ctx = _make_context(settings, store=store) ctx = _make_context(settings, store=store)
@ -443,13 +518,19 @@ async def test_message_handler_injects_kb_context_transiently(tmp_path):
mock_llm.chat = AsyncMock(return_value="reply with context") mock_llm.chat = AsyncMock(return_value="reply with context")
store = ThreadMemoryStore(tmp_path / "mem.json") store = ThreadMemoryStore(tmp_path / "mem.json")
store.save(ThreadSummary( store.save(
chat_id=1, thread_id=1, summary="Previous API discussion", ThreadSummary(
message_count=2, tags=["api"], chat_id=1,
)) thread_id=1,
summary="Previous API discussion",
message_count=2,
tags=["api"],
)
)
user_id = 999 user_id = 999
from steward.bot.telegram import _history from steward.bot.telegram import _history
_history[user_id].clear() _history[user_id].clear()
update = _make_update(user_id=user_id, text="tell me about the api work") update = _make_update(user_id=user_id, text="tell me about the api work")
@ -467,10 +548,7 @@ async def test_message_handler_injects_kb_context_transiently(tmp_path):
assert any("knowledge base" in m.get("content", "").lower() for m in history_arg) assert any("knowledge base" in m.get("content", "").lower() for m in history_arg)
# But the KB message must NOT be stored in _history # But the KB message must NOT be stored in _history
assert all( assert all("knowledge base" not in m.get("content", "").lower() for m in _history[user_id])
"knowledge base" not in m.get("content", "").lower()
for m in _history[user_id]
)
@pytest.mark.asyncio @pytest.mark.asyncio
@ -482,12 +560,19 @@ async def test_message_handler_no_kb_injection_when_no_match(tmp_path):
store = ThreadMemoryStore(tmp_path / "mem.json") store = ThreadMemoryStore(tmp_path / "mem.json")
# Store a summary with unrelated tags # Store a summary with unrelated tags
store.save(ThreadSummary( store.save(
chat_id=1, thread_id=1, summary="Database recap", message_count=1, tags=["database"], ThreadSummary(
)) chat_id=1,
thread_id=1,
summary="Database recap",
message_count=1,
tags=["database"],
)
)
user_id = 888 user_id = 888
from steward.bot.telegram import _history from steward.bot.telegram import _history
_history[user_id].clear() _history[user_id].clear()
update = _make_update(user_id=user_id, text="what is the weather like?") update = _make_update(user_id=user_id, text="what is the weather like?")

View File

@ -19,6 +19,7 @@ from steward.tools.client import (
_schema_to_json_schema, _schema_to_json_schema,
_spec_to_openai_tools, _spec_to_openai_tools,
) )
from tests.conftest import make_settings
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Fixtures # Fixtures
@ -86,6 +87,7 @@ SIMPLE_SPEC: dict[str, Any] = {
def _make_llm_settings(**kwargs: Any) -> Settings: def _make_llm_settings(**kwargs: Any) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict( defaults = dict(
telegram_bot_token="t", telegram_bot_token="t",
openai_api_key="k", openai_api_key="k",
@ -93,7 +95,7 @@ def _make_llm_settings(**kwargs: Any) -> Settings:
openai_system_prompt="You are Steward.", openai_system_prompt="You are Steward.",
) )
defaults.update(kwargs) defaults.update(kwargs)
return Settings(**defaults) return make_settings(**defaults)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@ -281,9 +283,7 @@ async def test_tool_client_call_get_with_query():
respx.get("http://tools.local/openapi.json").mock( respx.get("http://tools.local/openapi.json").mock(
return_value=Response(200, json=SIMPLE_SPEC) return_value=Response(200, json=SIMPLE_SPEC)
) )
respx.get("http://tools.local/items").mock( respx.get("http://tools.local/items").mock(return_value=Response(200, json=[{"id": 1}]))
return_value=Response(200, json=[{"id": 1}])
)
client = ToolClient("http://tools.local") client = ToolClient("http://tools.local")
result = await client.call("list_items", {"limit": 10}) result = await client.call("list_items", {"limit": 10})
assert "id" in result assert "id" in result