Skip to content
This repository was archived by the owner on Nov 10, 2025. It is now read-only.

Adds RAG feature - #406

Merged
lucasgomide merged 18 commits into
mainfrom
lg-custom-rag
Aug 19, 2025
Merged

lucasgomide merged 18 commits into
mainfrom
lg-custom-rag

Conversation

@lucasgomide

Copy link
Copy Markdown
Contributor

No description provided.

Comment thread crewai_tools/rag/data_types.py Dismissed
@lucasgomide
lucasgomide force-pushed the lg-custom-rag branch 2 times, most recently from f8ec3b6 to a773bd7 Compare August 4, 2025 21:59
Comment thread crewai_tools/rag/core.py

@lorenzejay lorenzejay left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are we going

Comment thread crewai_tools/rag/loaders/docx_loader.py Outdated
Comment thread crewai_tools/adapters/rag_adapter.py
- Each chunker uses separators specific to its content type
- Implement document deduplication logic in RAG
  * Check for existing documents by source reference
  * Compare doc IDs to detect content changes
  * Automatically replace outdated content while preventing duplicates

- Centralize common functionality for better maintainability
  * Create SourceContent class to handle URLs, files, and text uniformly
  * Extract shared utilities (compute_sha256) to misc.py
  * Standardize doc ID generation across all loaders

- Improve RAG system architecture
  * All loaders now inherit consistent patterns from centralized BaseLoader
  * Better separation of concerns with dedicated content management classes
  * Standardized LoaderResult structure across all loader implementations
@lucasgomide
lucasgomide merged commit eb770d0 into main Aug 19, 2025
7 checks passed
mplachta pushed a commit to mplachta/crewAI-tools that referenced this pull request Aug 27, 2025
* feat: initialize rag

* refactor: using cosine distance metric for chromadb

* feat: use RecursiveCharacterTextSplitter as chunker strategy

* feat: support chucker and loader per data_type

* feat: adding JSON loader

* feat: adding CSVLoader

* feat: adding loader for DOCX files

* feat: add loader for MDX files

* feat: add loader for XML files

* feat: add loader for parser Webpage

* feat: support to load files from an entire directory

* feat: support to auto-load the loaders for additional DataType

* feat: add chuckers for some specific data type

- Each chunker uses separators specific to its content type

* feat: prevent document duplication and centralize content management

- Implement document deduplication logic in RAG
  * Check for existing documents by source reference
  * Compare doc IDs to detect content changes
  * Automatically replace outdated content while preventing duplicates

- Centralize common functionality for better maintainability
  * Create SourceContent class to handle URLs, files, and text uniformly
  * Extract shared utilities (compute_sha256) to misc.py
  * Standardize doc ID generation across all loaders

- Improve RAG system architecture
  * All loaders now inherit consistent patterns from centralized BaseLoader
  * Better separation of concerns with dedicated content management classes
  * Standardized LoaderResult structure across all loader implementations

* chore: split text loaders file

* test: adding missing tests about RAG loaders

* refactor: QOL

* fix: add missing uv syntax on DOCXLoader
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants