Thank you for your interest in contributing! This document provides guidelines and instructions for contributing to the project.
We are committed to providing a welcoming and inclusive environment for all contributors. Please:
- Be respectful and professional
- Welcome diverse perspectives
- Focus on constructive feedback
- Report unacceptable behavior responsibly
git clone https://github.com/richardogundele/rag_knowledge_assistant.git
cd rag_knowledge_assistantgit checkout -b feature/your-feature-namepython -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
pip install pytest black flake8 mypy # Development tools- Keep commits focused and atomic
- Write clear, descriptive commit messages
- Follow the code style guidelines below
pytest tests/
black .
flake8
mypygit push origin feature/your-feature-name- Follow PEP 8
- Use Black for auto-formatting
- Use type hints for functions
- Maximum line length: 100 characters
from typing import List, Dict, Optional
def retrieve_documents(
query: str,
top_k: int = 5,
similarity_threshold: float = 0.3
) -> List[Dict[str, str]]:
"""
Retrieve documents matching the query.
Args:
query: Search query string
top_k: Number of results to return
similarity_threshold: Minimum similarity score
Returns:
List of document chunks with metadata
"""
# Implementation
passUse clear, descriptive commit messages:
feat: Add support for DOCX files
fix: Correct confidence score calculation
docs: Update API documentation
refactor: Simplify retriever logic
test: Add tests for guardrails
style: Format code with black
chore: Update dependencies
- Update Documentation: Update README.md and relevant docs
- Add Tests: Include tests for new features
- Write Description: Clear description of changes and motivation
- Link Issues: Reference related issues with "Closes #123"
- Request Review: Ask maintainers for review
## Description
Brief description of what this PR does.
## Motivation
Why is this change needed?
## Type of Change
- [ ] Bug fix
- [ ] New feature
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Refactoring
## Testing
How was this tested?
## Checklist
- [ ] Code follows style guidelines
- [ ] Tests added/updated
- [ ] Documentation updated
- [ ] No breaking changes (or documented)- Document format support (DOCX, Excel, HTML)
- Performance optimization for large datasets
- Production deployment guides
- Test coverage expansion
- Multi-language support
- Additional LLM integrations
- UI/UX improvements
- Analytics dashboard
- Documentation improvements
- Bug fixes
- Code comments and docstrings
- Test additions
# Run all tests
pytest
# Run specific test file
pytest tests/test_retriever.py
# Run with coverage
pytest --cov=services tests/
# Run with verbose output
pytest -v- Maintainers will review your PR
- Provide feedback and request changes if needed
- Once approved, your PR will be merged
- Your contribution will be acknowledged in release notes
When submitting changes that affect documentation:
- Update relevant README sections
- Update docstrings in code
- Update API documentation if endpoints change
- Add changelog entry in CHANGELOG.md
Submit bugs through GitHub Issues:
Title: Brief description of the bug
Description:
## Bug Description
Clear description of what's not working.
## Steps to Reproduce
1. Step 1
2. Step 2
3. Step 3
## Expected Behavior
What should happen.
## Actual Behavior
What actually happens.
## Environment
- OS: (Windows/Mac/Linux)
- Python: 3.10/3.11/etc
- Ollama: (installed/version)
- Documents: (number and type)
## Additional Context
Any additional information.
Submit feature requests through GitHub Discussions:
Title: Brief feature description
Description:
## Use Case
Why is this feature needed?
## Proposed Solution
How should it work?
## Alternatives Considered
Other approaches considered.
## Additional Context
Any additional information.
# Add debug logging
import logging
logging.basicConfig(level=logging.DEBUG)
logger = logging.getLogger(__name__)
logger.debug(f"Query: {query}")
logger.info("Documents retrieved")
logger.warning("Low confidence score")
logger.error("Failed to process document")# Profile your code
python -m cProfile -s cumulative main.py
# Memory profiling
pip install memory_profiler
python -m memory_profiler main.py# Test FAISS operations in isolation
from services.vector_store import FAISSVectorStore
store = FAISSVectorStore()
store.add_embeddings(embeddings)
results = store.search(query_embedding, top_k=5)def function_name(param1: str, param2: int) -> Dict:
"""
Brief one-line description.
Longer description if needed. Explain what the function does,
why it exists, and how to use it.
Args:
param1: Description of param1
param2: Description of param2
Returns:
Description of return value
Raises:
ValueError: When invalid input provided
TimeoutError: When document processing exceeds timeout
Example:
>>> result = function_name("test", 42)
>>> result['status']
'success'
"""
pass- Update version in
config.py - Update
CHANGELOG.md - Create GitHub release with tag
- Publish to PyPI (if applicable)
- Announce in discussions
- 📖 Read the README
- 💬 Start a Discussion
- 🐛 File an Issue
By contributing, you agree that your contributions will be licensed under the MIT License.
Thank you for contributing to Government Knowledge Assistant! 🎉