test: update test suite for OmegaConf config system

- Add tests/conftest.py with make_settings() helper that maps legacy config
  parameter names to new nested OmegaConf structure
- Update all test files to use the conftest fixture
- All 73 tests now pass with new config system
- Maintains backward compatibility via Settings class properties

Also add AGENTS.md with comprehensive AI agent guidelines:
- Project overview and key technologies
- Directory structure reference
- Development workflow and common tasks
- Testing strategy and patterns
- CI/CD pipeline overview
- Common pitfalls and best practices
- Debugging guide for agents working on the project
This commit is contained in:
Daniel Wagner 2026-07-26 12:48:40 +10:00
parent d0ccc96507
commit 070e967833
10 changed files with 745 additions and 53 deletions

18
.pre-commit-config.yaml Normal file
View File

@ -0,0 +1,18 @@
repos:
- repo: local
hooks:
- id: ruff-check
name: ruff check
entry: ruff check
language: system
types: [python]
stages: [commit]
- id: pytest
name: pytest
entry: pytest
language: system
types: [python]
pass_filenames: false
always_run: true
stages: [commit]

316
AGENTS.md Normal file
View File

@ -0,0 +1,316 @@
# AI Agent Guidelines for Steward
This document provides guidance for AI agents (Copilot, Claude, etc.) working on the Steward project.
## Project Overview
**Steward** is a long-running, AI-assisted personal operations platform that:
- Runs as a Telegram bot with group/channel support
- Uses OmegaConf for flexible configuration management
- Supports message threads for organized conversations
- Stores and retrieves conversation summaries via knowledge base
- Integrates with OpenAI-compatible LLM APIs
- Optionally integrates with MCP/OpenAPI tool servers for function calling
**Key Technologies:**
- Python 3.12+
- `python-telegram-bot` (Telegram bot framework)
- OmegaConf (configuration management)
- Pydantic (data validation)
- OpenAI API (LLM integration)
- AsyncIO (async/await patterns throughout)
## Directory Structure
```
steward/
├── steward/ # Main package
│ ├── bot/ # Telegram bot implementation
│ │ └── telegram.py # All bot handlers and logic
│ ├── config.py # OmegaConf-based configuration
│ ├── config_schema.yaml # Default configuration schema
│ ├── llm/ # LLM client
│ ├── memory/ # Thread memory storage
│ ├── proposals/ # Proposal generation
│ ├── tools/ # MCP tool integration
│ └── main.py # Application entry point
├── tests/ # Test suite
│ ├── conftest.py # Pytest fixtures and helpers
│ ├── test_bot.py # Bot handler tests
│ ├── test_llm_client.py # LLM client tests
│ ├── test_proposals.py # Proposal generation tests
│ ├── test_thread_memory.py # Memory storage tests
│ └── test_tools.py # Tool integration tests
├── .github/workflows/ # CI/CD workflows
├── docker-compose.yml # Production compose (uses ghcr.io image)
├── docker-compose.dev.yml # Development compose (builds locally)
├── Dockerfile # Container image definition
├── pyproject.toml # Project metadata and dependencies
├── CONFIGURATION.md # Configuration guide
└── AGENTS.md # This file
```
## Development Workflow
### Before Making Changes
1. **Understand the Context**
- Read CONFIGURATION.md for config system details
- Review the async/await patterns in existing code
- Note that the bot uses Telegram's `python-telegram-bot` library
2. **Check Existing Tests**
- Run tests locally or via CI before proposing changes
- Tests use `tests/conftest.py` fixtures for setup
- The `_make_settings()` helper maps legacy config to new structure
3. **Code Style**
- Follow the ruff linting rules in `pyproject.toml`
- Line length: 100 characters
- Use type hints throughout (mypy strict mode)
- Import order: standard lib → third party → local
### Making Changes
#### Configuration Changes
- Update `steward/config_schema.yaml` for schema changes
- Update `steward/config.py` for new config sections
- Update `.env.example` to show all available settings
- Document in `CONFIGURATION.md`
#### Bot Handler Changes
- All handlers in `steward/bot/telegram.py` are async
- Handlers receive `Update` and `ContextTypes.DEFAULT_TYPE` parameters
- Use `update.message.reply_text()` for responses
- Check user/group authorization early in handlers
- Test with `tests/test_bot.py`
#### Adding Tests
- Use `tests/conftest.py` for shared fixtures
- Use `_make_settings()` helper to create test settings
- Use `_make_update()` to create mock Telegram updates
- All bot tests should be async (use `pytest-asyncio`)
- Mock external APIs with `respx` or `pytest-mock`
#### Creating New Features
**For Configuration:**
1. Add section to `config_schema.yaml`
2. Create pydantic model in `config.py`
3. Add properties to `Settings` class for backward compatibility
4. Document in `CONFIGURATION.md`
**For Bot Handlers:**
1. Add async handler function in `steward/bot/telegram.py`
2. Add to application in `build_application()`
3. Add authorization checks (`_is_allowed()`, `_is_group_enabled()`)
4. Write tests in `tests/test_bot.py`
**For LLM Integration:**
1. Update `steward/llm/client.py`
2. Add history context handling if needed
3. Test with `tests/test_llm_client.py`
4. Consider impact on token usage
### Docker & Deployment
**Local Development:**
```bash
docker compose -f docker-compose.dev.yml up --build
```
**Production:**
```bash
docker compose up
# Uses prebuilt image from ghcr.io/djw4/steward:latest
```
**Key Docker Notes:**
- Uses `uv` instead of pip for faster installation
- Non-root user (UID 1000) for security
- Thread memory stored in `/data` volume
- Environment variables set via `.env` file
## Common Tasks
### Adding a New Bot Command
1. Create handler function in `steward/bot/telegram.py`:
```python
async def mycommand_handler(update: Update, context: ContextTypes.DEFAULT_TYPE) -> None:
settings: Settings = context.bot_data["settings"]
user = update.effective_user
chat = update.effective_chat
# Authorization checks
if user is None or not _is_allowed(user.id, settings):
return
if chat is None or not _is_group_enabled(chat.id, settings):
return
# Handler logic here
await update.message.reply_text("Response")
```
2. Register in `build_application()`:
```python
app.add_handler(CommandHandler("mycommand", mycommand_handler))
```
3. Add test in `tests/test_bot.py`:
```python
@pytest.mark.asyncio
async def test_mycommand_handler_does_something():
# Test implementation
```
### Adding a New Configuration Section
1. Update `steward/config_schema.yaml`:
```yaml
mysection:
my_setting: "default_value"
my_number: 42
```
2. Create pydantic model in `steward/config.py`:
```python
class MyConfig(BaseModel):
my_setting: str = Field(default="default_value")
my_number: int = Field(default=42)
```
3. Add to `Settings` class and add property for backward compatibility:
```python
class Settings(BaseModel):
mysection: MyConfig = Field(default_factory=MyConfig)
@property
def my_setting(self) -> str:
return self.mysection.my_setting
```
### Running Tests Locally
```bash
# Install dev dependencies
pip install -e ".[dev]"
# Run all tests
pytest tests/ -v
# Run specific test file
pytest tests/test_bot.py -v
# Run specific test
pytest tests/test_bot.py::test_mytest -v
# Run with coverage
pytest tests/ --cov=steward --cov-report=html
```
### Debugging Issues
**Bot not responding:**
1. Check `.env` file has all required settings
2. Verify bot has message access in BotFather
3. Check logs: `docker compose logs steward -f`
4. Verify group ID format (negative for groups: `-1005308306472`)
**Tests failing:**
1. Run `ruff check steward/ tests/` to check linting
2. Run `mypy steward/` for type checking
3. Check test fixture setup in `conftest.py`
4. Review error messages carefully for import/config issues
**Configuration issues:**
1. Check `.env` file syntax (should be `KEY=value`)
2. For lists, use JSON format: `STEWARD__TELEGRAM__ALLOWED_USER_IDS='[123, 456]'`
3. New format variables use `STEWARD__SECTION__KEY` pattern
4. Check `CONFIGURATION.md` for examples
## Testing Strategy
### Test Organization
- **`test_bot.py`**: Handler logic, authorization, history management
- **`test_llm_client.py`**: LLM API integration, error handling
- **`test_proposals.py`**: Proposal generation, API calls
- **`test_thread_memory.py`**: Memory storage, search, serialization
- **`test_tools.py`**: Tool calling, MCP integration
### Test Patterns
**Handler Tests:**
```python
@pytest.mark.asyncio
async def test_handler_name():
settings = _make_settings()
update = _make_update(user_id=123, text="hello")
context = _make_context(settings)
await handler_name(update, context)
# Assert on reply_text mock calls
```
**API Tests:**
```python
def test_api_call():
with respx.mock:
respx.post("https://api.example.com/endpoint").mock(
return_value=httpx.Response(200, json={"result": "ok"})
)
# Call function that makes API request
result = function_under_test()
assert result == expected_value
```
### CI/CD Pipeline
The CI workflow (`.github/workflows/ci.yml`) runs:
1. **Linting** (ruff): `python -m ruff check steward/ tests/`
2. **Tests** (pytest): `python -m pytest tests/ -v`
3. **Docker Build**: Builds and pushes image to GHCR
All three must pass for merges to main.
## Common Pitfalls
### ❌ Don't:
- Use synchronous code instead of async/await
- Hardcode secrets in config files (use env vars)
- Forget to add authorization checks to new handlers
- Write tests without using conftest fixtures
- Modify `.env` in version control (update `.env.example` instead)
- Change config structure without updating all layers (schema, pydantic, docs)
### ✅ Do:
- Use `async/await` consistently throughout
- Store secrets in environment variables
- Check both user and group authorization
- Use `pytest.mark.asyncio` for async tests
- Keep `.env` out of git (it's in `.gitignore`)
- Update schema → pydantic → properties → documentation together
## Questions? Issues?
- Check `CONFIGURATION.md` for config-related questions
- Review existing handlers in `steward/bot/telegram.py` for patterns
- Look at test examples in `tests/` for testing patterns
- Check `.github/workflows/ci.yml` for what CI expects
## Key Files to Know
| File | Purpose |
|------|---------|
| `steward/config.py` | Configuration system - OmegaConf + pydantic |
| `steward/bot/telegram.py` | All bot handlers and core logic |
| `steward/main.py` | Application entry point and setup |
| `tests/conftest.py` | Pytest fixtures and helpers |
| `pyproject.toml` | Dependencies and tool configuration |
| `CONFIGURATION.md` | User-facing config guide |
| `.github/workflows/ci.yml` | CI/CD pipeline definition |

174
MIGRATION_NOTES.md Normal file
View File

@ -0,0 +1,174 @@
# Test Settings Migration Notes
## Overview
Updated all test files to work with the new OmegaConf-based configuration system while maintaining backward compatibility through a helper function.
## What Changed
### 1. Created `tests/conftest.py` (new file)
- Implements `make_settings(**kwargs)` helper function
- Maps legacy parameter names to the new nested structure
- Used as a shared utility across all test files
### 2. Updated Test Files
Updated the following test files to use the new helper:
- `tests/test_bot.py`
- `tests/test_llm_client.py`
- `tests/test_proposals.py`
- `tests/test_thread_memory.py`
- `tests/test_tools.py`
Each file now:
1. Imports `make_settings` from `tests.conftest`
2. Keeps local `_make_settings()` wrapper for compatibility
3. Calls the shared `make_settings()` function internally
## Parameter Mapping
The `make_settings()` function handles the following legacy-to-new mappings:
### Telegram Config
- `telegram_bot_token` → `telegram.bot_token`
- `telegram_allowed_user_ids` → `telegram.allowed_user_ids`
- `telegram_group_ids` → `telegram.group_ids`
### OpenAI Config
- `openai_api_key` → `openai.api_key`
- `openai_base_url` → `openai.base_url` (default: `"https://api.openai.com/v1"`)
- `openai_model` → `openai.model` (default: `"gpt-4o"`)
- `openai_system_prompt` → `openai.system_prompt` (uses default if not provided)
### Analysis Config
- `analysis_target_url` → `analysis.target_url`
- `analysis_target_api_key` → `analysis.target_api_key`
- `analysis_cron_hour` → `analysis.cron_hour` (default: `8`)
- `analysis_cron_minute` → `analysis.cron_minute` (default: `0`)
### Memory Config
- `thread_memory_path` → `memory.thread_memory_path` (default: `"thread_memory.json"`)
### Tools Config
- `mcp_server_url` → `tools.mcp_server_url`
- `mcp_server_api_key` → `tools.mcp_server_api_key`
## Usage Examples
### Before (would fail with new Settings class)
```python
settings = Settings(
telegram_bot_token="token",
openai_api_key="key",
telegram_allowed_user_ids=[1, 2, 3]
)
```
### After (with new helper)
```python
from tests.conftest import make_settings
settings = make_settings(
telegram_bot_token="token",
openai_api_key="key",
telegram_allowed_user_ids=[1, 2, 3]
)
# Now settings.telegram.bot_token == "token"
# And settings.telegram.allowed_user_ids == [1, 2, 3]
# And backward-compat properties still work:
# settings.telegram_bot_token == "token"
# settings.telegram_allowed_user_ids == [1, 2, 3]
```
## How Test Files Use It
Each test file has a local `_make_settings()` wrapper that:
1. Defines default values specific to that test module
2. Merges any provided kwargs
3. Delegates to the shared `make_settings()` function
Example from `test_bot.py`:
```python
def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict(
telegram_bot_token="test-token",
openai_api_key="test-key",
)
defaults.update(kwargs)
return make_settings(**defaults)
```
This allows tests to override defaults while reusing the mapping logic:
```python
# Uses defaults
settings = _make_settings()
# Override specific values
settings = _make_settings(telegram_allowed_user_ids=[1, 2, 3])
```
## Error Handling
The `make_settings()` function will raise a `TypeError` if any unexpected keyword arguments are passed:
```python
make_settings(invalid_param="value")
# Raises: TypeError: Unexpected keyword arguments: invalid_param
```
This helps catch typos and unexpected parameters early in tests.
## Backward Compatibility
The Settings class retains backward-compatible properties:
- `settings.telegram_bot_token` → returns `settings.telegram.bot_token`
- `settings.openai_api_key` → returns `settings.openai.api_key`
- etc.
This means existing code that reads from Settings using the old property names continues to work.
## Files Modified
1. **`tests/conftest.py`** (NEW)
- 91 lines
- Core mapping logic for legacy-to-new Settings structure
2. **`tests/test_bot.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
3. **`tests/test_llm_client.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
4. **`tests/test_proposals.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
5. **`tests/test_thread_memory.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_settings()` to call shared function
6. **`tests/test_tools.py`**
- Added import: `from tests.conftest import make_settings`
- Updated `_make_llm_settings()` to call shared function
## Testing
To verify the changes work:
```bash
# Install dev dependencies
pip install -e ".[dev]"
# Run all tests
pytest tests/
# Run specific test file
pytest tests/test_bot.py -v
# Run specific test
pytest tests/test_bot.py::TestIsAllowed::test_allowlist_accepts_known_user -v
```
All tests should pass with these changes as the Settings class:
1. Accepts the new nested structure from `make_settings()`
2. Provides backward-compatible properties for old code

View File

@ -27,6 +27,7 @@ dev = [
"ruff>=0.4",
"mypy>=1.10",
"respx>=0.21",
"pre-commit>=3.0",
]
[project.scripts]

91
tests/conftest.py Normal file
View File

@ -0,0 +1,91 @@
"""Shared test fixtures and utilities."""
from steward.config import (
AnalysisConfig,
MemoryConfig,
OpenAIConfig,
Settings,
TelegramConfig,
ToolsConfig,
)
def make_settings(**kwargs) -> Settings:
"""Create a Settings instance from old-style kwargs.
Maps legacy parameter names to the new nested structure:
- telegram_bot_token -> telegram.bot_token
- telegram_allowed_user_ids -> telegram.allowed_user_ids
- telegram_group_ids -> telegram.group_ids
- openai_api_key -> openai.api_key
- openai_base_url -> openai.base_url
- openai_model -> openai.model
- openai_system_prompt -> openai.system_prompt
- analysis_target_url -> analysis.target_url
- analysis_target_api_key -> analysis.target_api_key
- analysis_cron_hour -> analysis.cron_hour
- analysis_cron_minute -> analysis.cron_minute
- thread_memory_path -> memory.thread_memory_path
- mcp_server_url -> tools.mcp_server_url
- mcp_server_api_key -> tools.mcp_server_api_key
Args:
**kwargs: Legacy-style configuration parameters
Returns:
Settings: A fully constructed Settings instance
"""
# Extract telegram config
telegram_config = TelegramConfig(
bot_token=kwargs.pop("telegram_bot_token", ""),
allowed_user_ids=kwargs.pop("telegram_allowed_user_ids", []),
group_ids=kwargs.pop("telegram_group_ids", []),
)
# Extract OpenAI config
openai_config = OpenAIConfig(
api_key=kwargs.pop("openai_api_key", ""),
base_url=kwargs.pop("openai_base_url", "https://api.openai.com/v1"),
model=kwargs.pop("openai_model", "gpt-4o"),
system_prompt=kwargs.pop(
"openai_system_prompt",
(
"You are Steward, a persistent, trustworthy AI-assisted personal operations platform. "
"You reduce cognitive load by observing, remembering, planning, and proposing actions. "
"You are conservative, transparent, and policy-aware. "
"Always explain your reasoning."
),
),
)
# Extract analysis config
analysis_config = AnalysisConfig(
target_url=kwargs.pop("analysis_target_url", ""),
target_api_key=kwargs.pop("analysis_target_api_key", ""),
cron_hour=kwargs.pop("analysis_cron_hour", 8),
cron_minute=kwargs.pop("analysis_cron_minute", 0),
)
# Extract memory config
memory_config = MemoryConfig(
thread_memory_path=kwargs.pop("thread_memory_path", "thread_memory.json"),
)
# Extract tools config
tools_config = ToolsConfig(
mcp_server_url=kwargs.pop("mcp_server_url", ""),
mcp_server_api_key=kwargs.pop("mcp_server_api_key", ""),
)
# Any remaining kwargs should be rejected
if kwargs:
raise TypeError(f"Unexpected keyword arguments: {', '.join(kwargs.keys())}")
# Create and return Settings instance
return Settings(
telegram=telegram_config,
openai=openai_config,
analysis=analysis_config,
memory=memory_config,
tools=tools_config,
)

View File

@ -17,15 +17,17 @@ from steward.bot.telegram import (
from steward.config import Settings
from steward.llm.client import LLMClient
from steward.memory.thread_store import ThreadMemoryStore
from tests.conftest import make_settings
def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict(
telegram_bot_token="test-token",
openai_api_key="test-key",
)
defaults.update(kwargs)
return Settings(**defaults)
return make_settings(**defaults)
def _make_update(

View File

@ -6,9 +6,11 @@ import pytest
from steward.config import Settings
from steward.llm.client import LLMClient
from tests.conftest import make_settings
def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict(
telegram_bot_token="test-token",
openai_api_key="test-key",
@ -16,7 +18,7 @@ def _make_settings(**kwargs) -> Settings:
openai_system_prompt="You are Steward.",
)
defaults.update(kwargs)
return Settings(**defaults)
return make_settings(**defaults)
@pytest.fixture

View File

@ -10,16 +10,18 @@ import respx
from steward.config import Settings
from steward.llm.client import LLMClient
from steward.proposals.generator import Proposal, ProposalGenerator
from tests.conftest import make_settings
def _make_settings(**kwargs) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict(
openai_api_key="test-key",
analysis_target_url="https://example.com/api/status",
analysis_target_api_key="secret",
)
defaults.update(kwargs)
return Settings(**defaults)
return make_settings(**defaults)
@pytest.fixture
@ -94,6 +96,7 @@ async def test_fetch_sends_bearer_token(settings, mock_llm):
def test_proposal_format_for_telegram():
"""format_for_telegram() should contain the source URL and body."""
from datetime import datetime
p = Proposal(
title="Test",
body="**Summary:** All good.",

View File

@ -16,13 +16,18 @@ from steward.bot.telegram import (
from steward.config import Settings
from steward.llm.client import LLMClient
from steward.memory.thread_store import ThreadMemoryStore, ThreadSummary
from tests.conftest import make_settings
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _make_settings(**kwargs) -> Settings:
return Settings(telegram_bot_token="tok", openai_api_key="key", **kwargs)
"""Wrapper for test compatibility."""
defaults = dict(telegram_bot_token="tok", openai_api_key="key")
defaults.update(kwargs)
return make_settings(**defaults)
def _make_update(
@ -72,6 +77,7 @@ def _make_context(
# ThreadMemoryStore unit tests
# ---------------------------------------------------------------------------
class TestThreadMemoryStore:
def _store(self, tmp_path: Path) -> ThreadMemoryStore:
return ThreadMemoryStore(tmp_path / "mem.json")
@ -107,10 +113,24 @@ class TestThreadMemoryStore:
def test_all_returns_newest_first(self, tmp_path: Path):
store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="first", message_count=1,
flushed_at="2026-01-01T00:00:00+00:00"))
store.save(ThreadSummary(chat_id=1, thread_id=2, summary="second", message_count=1,
flushed_at="2026-06-01T00:00:00+00:00"))
store.save(
ThreadSummary(
chat_id=1,
thread_id=1,
summary="first",
message_count=1,
flushed_at="2026-01-01T00:00:00+00:00",
)
)
store.save(
ThreadSummary(
chat_id=1,
thread_id=2,
summary="second",
message_count=1,
flushed_at="2026-06-01T00:00:00+00:00",
)
)
results = store.all()
assert results[0].summary == "second"
assert results[1].summary == "first"
@ -129,7 +149,10 @@ class TestThreadMemoryStore:
def test_format_for_telegram_shows_tags(self, tmp_path: Path):
s = ThreadSummary(
chat_id=1, thread_id=77, summary="recap", message_count=3,
chat_id=1,
thread_id=77,
summary="recap",
message_count=3,
tags=["api", "auth"],
)
text = s.format_for_telegram()
@ -138,54 +161,87 @@ class TestThreadMemoryStore:
def test_from_dict_tolerates_missing_tags(self, tmp_path: Path):
"""from_dict must handle legacy entries that predate the tags field."""
raw = {"chat_id": 1, "thread_id": 2, "summary": "old", "message_count": 3,
"flushed_at": "2026-01-01T00:00:00+00:00"}
raw = {
"chat_id": 1,
"thread_id": 2,
"summary": "old",
"message_count": 3,
"flushed_at": "2026-01-01T00:00:00+00:00",
}
s = ThreadSummary.from_dict(raw)
assert s.tags == []
def test_search_returns_matching_summaries(self, tmp_path: Path):
store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="API work", message_count=2,
tags=["api", "design", "auth"]))
store.save(ThreadSummary(chat_id=1, thread_id=2, summary="Database work", message_count=2,
tags=["database", "schema"]))
store.save(
ThreadSummary(
chat_id=1,
thread_id=1,
summary="API work",
message_count=2,
tags=["api", "design", "auth"],
)
)
store.save(
ThreadSummary(
chat_id=1,
thread_id=2,
summary="Database work",
message_count=2,
tags=["database", "schema"],
)
)
results = store.search("api")
assert len(results) == 1
assert results[0].thread_id == 1
def test_search_no_match_returns_empty(self, tmp_path: Path):
store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
tags=["database"]))
store.save(
ThreadSummary(
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"]
)
)
assert store.search("deployment") == []
def test_search_empty_query_returns_empty(self, tmp_path: Path):
store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
tags=["database"]))
store.save(
ThreadSummary(
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"]
)
)
assert store.search("") == []
def test_search_case_insensitive(self, tmp_path: Path):
store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
tags=["API"]))
store.save(
ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["API"])
)
assert len(store.search("api")) == 1
def test_search_skips_untagged_summaries(self, tmp_path: Path):
store = self._store(tmp_path)
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1,
tags=[]))
store.save(ThreadSummary(chat_id=1, thread_id=1, summary="recap", message_count=1, tags=[]))
assert store.search("api") == []
def test_legacy_store_roundtrip(self, tmp_path: Path):
"""Summaries without tags survive a save/load cycle via from_dict."""
path = tmp_path / "mem.json"
import json as _json
path.write_text(
_json.dumps({
"1:1": {"chat_id": 1, "thread_id": 1, "summary": "old",
"message_count": 1, "flushed_at": "2026-01-01T00:00:00+00:00"}
}),
_json.dumps(
{
"1:1": {
"chat_id": 1,
"thread_id": 1,
"summary": "old",
"message_count": 1,
"flushed_at": "2026-01-01T00:00:00+00:00",
}
}
),
encoding="utf-8",
)
store = ThreadMemoryStore(path)
@ -198,6 +254,7 @@ class TestThreadMemoryStore:
# Thread-aware message_handler tests
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_thread_message_stored_in_thread_history():
"""Messages in a thread go to _thread_history, not _history."""
@ -249,6 +306,7 @@ async def test_thread_history_is_unbounded():
# /flush handler tests
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_flush_outside_thread_warns():
"""flush_handler outside a thread should warn the user."""
@ -328,6 +386,7 @@ async def test_flush_summarises_stores_and_compresses():
# /recall handler tests
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_recall_in_thread_returns_stored_summary(tmp_path):
"""recall_handler inside a thread returns the stored summary for that thread."""
@ -396,14 +455,24 @@ async def test_recall_with_query_returns_matching_summaries(tmp_path):
"""recall_handler with a query argument searches the knowledge base by keyword."""
settings = _make_settings()
store = ThreadMemoryStore(tmp_path / "mem.json")
store.save(ThreadSummary(
chat_id=1, thread_id=1, summary="API authentication discussion",
message_count=2, tags=["api", "auth"],
))
store.save(ThreadSummary(
chat_id=1, thread_id=2, summary="Database schema planning",
message_count=3, tags=["database", "schema"],
))
store.save(
ThreadSummary(
chat_id=1,
thread_id=1,
summary="API authentication discussion",
message_count=2,
tags=["api", "auth"],
)
)
store.save(
ThreadSummary(
chat_id=1,
thread_id=2,
summary="Database schema planning",
message_count=3,
tags=["database", "schema"],
)
)
update = _make_update(thread_id=None)
ctx = _make_context(settings, store=store)
@ -421,9 +490,15 @@ async def test_recall_with_query_no_match(tmp_path):
"""recall_handler with a query that matches nothing tells the user."""
settings = _make_settings()
store = ThreadMemoryStore(tmp_path / "mem.json")
store.save(ThreadSummary(
chat_id=1, thread_id=1, summary="recap", message_count=1, tags=["database"],
))
store.save(
ThreadSummary(
chat_id=1,
thread_id=1,
summary="recap",
message_count=1,
tags=["database"],
)
)
update = _make_update(thread_id=None)
ctx = _make_context(settings, store=store)
@ -443,13 +518,19 @@ async def test_message_handler_injects_kb_context_transiently(tmp_path):
mock_llm.chat = AsyncMock(return_value="reply with context")
store = ThreadMemoryStore(tmp_path / "mem.json")
store.save(ThreadSummary(
chat_id=1, thread_id=1, summary="Previous API discussion",
message_count=2, tags=["api"],
))
store.save(
ThreadSummary(
chat_id=1,
thread_id=1,
summary="Previous API discussion",
message_count=2,
tags=["api"],
)
)
user_id = 999
from steward.bot.telegram import _history
_history[user_id].clear()
update = _make_update(user_id=user_id, text="tell me about the api work")
@ -467,10 +548,7 @@ async def test_message_handler_injects_kb_context_transiently(tmp_path):
assert any("knowledge base" in m.get("content", "").lower() for m in history_arg)
# But the KB message must NOT be stored in _history
assert all(
"knowledge base" not in m.get("content", "").lower()
for m in _history[user_id]
)
assert all("knowledge base" not in m.get("content", "").lower() for m in _history[user_id])
@pytest.mark.asyncio
@ -482,12 +560,19 @@ async def test_message_handler_no_kb_injection_when_no_match(tmp_path):
store = ThreadMemoryStore(tmp_path / "mem.json")
# Store a summary with unrelated tags
store.save(ThreadSummary(
chat_id=1, thread_id=1, summary="Database recap", message_count=1, tags=["database"],
))
store.save(
ThreadSummary(
chat_id=1,
thread_id=1,
summary="Database recap",
message_count=1,
tags=["database"],
)
)
user_id = 888
from steward.bot.telegram import _history
_history[user_id].clear()
update = _make_update(user_id=user_id, text="what is the weather like?")

View File

@ -19,6 +19,7 @@ from steward.tools.client import (
_schema_to_json_schema,
_spec_to_openai_tools,
)
from tests.conftest import make_settings
# ---------------------------------------------------------------------------
# Fixtures
@ -86,6 +87,7 @@ SIMPLE_SPEC: dict[str, Any] = {
def _make_llm_settings(**kwargs: Any) -> Settings:
"""Wrapper for test compatibility."""
defaults = dict(
telegram_bot_token="t",
openai_api_key="k",
@ -93,7 +95,7 @@ def _make_llm_settings(**kwargs: Any) -> Settings:
openai_system_prompt="You are Steward.",
)
defaults.update(kwargs)
return Settings(**defaults)
return make_settings(**defaults)
# ---------------------------------------------------------------------------
@ -281,9 +283,7 @@ async def test_tool_client_call_get_with_query():
respx.get("http://tools.local/openapi.json").mock(
return_value=Response(200, json=SIMPLE_SPEC)
)
respx.get("http://tools.local/items").mock(
return_value=Response(200, json=[{"id": 1}])
)
respx.get("http://tools.local/items").mock(return_value=Response(200, json=[{"id": 1}]))
client = ToolClient("http://tools.local")
result = await client.call("list_items", {"limit": 10})
assert "id" in result