hzywhite
9bc5f1578c
Merge branch 'main' into RAGAnything
2025-09-16 17:52:58 +08:00
yangdx
050a00b693
Update webui assets
2025-09-16 17:33:05 +08:00
yangdx
db524532f1
Bump core version to v.1.4.8.2 and API version to 0223
2025-09-16 17:16:57 +08:00
yangdx
0e8d973d44
Shorten progress prefix in entity extraction error messages
2025-09-16 15:48:37 +08:00
hzywhite
680b7c5b89
merge
2025-09-16 14:32:36 +08:00
yangdx
ecaee43788
Add error handling with chunk ID prefixing in entity extraction
2025-09-16 13:41:49 +08:00
yangdx
5f45ff56be
Merge remote-tracking branch 'origin/main'
2025-09-15 12:34:04 +08:00
yangdx
7b371309dd
Update README
2025-09-15 12:31:39 +08:00
yangdx
02c0066df0
Bump core version to 1.4.8.1
2025-09-15 05:34:34 +08:00
yangdx
37d01e2df8
fix: Ensures complete metadata (source_id, created_at, file_path) is preserved in aquery_data responses
2025-09-15 03:45:09 +08:00
yangdx
e71229698d
refactor: centralize metadata generation in query functions
...
- Remove processing_info generation from _convert_to_user_format function
- Move all metadata generation (keywords, processing_info) to kg_query and naive_query functions
- Simplify _convert_to_user_format to focus only on data format conversion
2025-09-15 03:11:07 +08:00
yangdx
c0d5abba6b
Fix linting
2025-09-15 02:59:21 +08:00
yangdx
b1c8206346
Add aquery_data endpoint for structured retrieval without LLM generation
...
- Add QueryDataResponse model
- Implement /query/data endpoint
- Add aquery_data method to LightRAG
- Return entities, relationships, chunks
2025-09-15 02:15:14 +08:00
yangdx
f69c5dfd9a
Add language control and format clarity to extraction prompts
2025-09-14 18:26:41 +08:00
yangdx
3ae827c255
Bump API version to 0222
2025-09-14 17:52:27 +08:00
yangdx
6e37460964
Improve entity extraction prompt clarity and make sure LLM output content only
2025-09-14 17:50:56 +08:00
yangdx
82a67354d0
Code formatting improvements and style consistency fixes
...
* Remove trailing whitespace
* Fix function signature ellipsis style
2025-09-14 17:49:02 +08:00
yangdx
87bb8a023b
Fix tuple delimiter regex patterns and add debug logging
...
- Add debug logs for malformed records
- Fix regex for consecutive delimiters
- Handle missing closing brackets
2025-09-14 17:29:27 +08:00
yangdx
4de1473875
Improve entity extraction prompts and error message formatting
...
• Fix typo in error log message
• Clarify format requirements in prompts
• Make extraction instructions clearer
• Improve user prompt consistency
2025-09-14 13:45:59 +08:00
yangdx
70fee5bbeb
Fix syntax warning by removin examples from fix_tuple_delimiter_corruption docstring
2025-09-14 12:37:21 +08:00
yangdx
20c5127c7c
Merge branch 'optimize-extraction' into return-data-only
2025-09-14 12:33:37 +08:00
yangdx
619553021e
Fix delimiter processing and optimize case-sensitive handling
...
• Fix completion_delimiter reference bug
• Add case check before lowercase conversion
• Improve delimiter corruption handling
• Optimize redundant processing logic
2025-09-14 12:23:48 +08:00
yangdx
ff705a2323
Fix tuple delimiter corruption when missing closing bracket, Handle <|#: -> <|#|> pattern
2025-09-14 11:44:21 +08:00
yangdx
fd48afdb00
Use "relation" instead of "relationship" in extration prompt, and support both format for safty
2025-09-14 11:43:35 +08:00
yangdx
1dc96f3959
Merge branch 'optimize-extraction' into return-data-only
2025-09-14 05:37:48 +08:00
yangdx
b820d8d588
Fix entity/relationship record parsing in extraction result processing
2025-09-14 05:35:01 +08:00
yangdx
4f5ad76c2c
Add entity vector database upsert for newly added entities by edges upserts
2025-09-14 05:04:45 +08:00
yangdx
7cc2b69bcf
Fix linting
2025-09-14 05:02:02 +08:00
yangdx
cddd81a86c
Fix LLM output format errors in extraction result processing
...
- Handle tuple_delimiter as record separator
- Add format validation and correction
- Add warning for format errors
2025-09-14 04:13:01 +08:00
yangdx
419f4f0268
Update web assets
2025-09-14 02:31:42 +08:00
yangdx
d993464a92
Restructure entity extraction prompt with clearer formatting and examples
...
* Improved instruction clarity
* Added better formatting structure
* Enhanced delimiter usage rules
* Clarified relationship handling
* Better third-person guidelines
2025-09-14 02:30:32 +08:00
yangdx
5311083f43
Rename "Process" entity type to "Method" across all components
2025-09-14 02:30:05 +08:00
yangdx
7060cf17f0
Add Process and Data entity types to LLM extraction system
...
• Add Process and Data to default types
• Update env.example configuration
• Add translations for new entities
• Support 5 languages (en/zh/fr/ar/tw)
2025-09-14 01:14:47 +08:00
yangdx
2686fc526e
Change entity type from CreativeWork to Content and update delimiter
...
• Replace CreativeWork with Content type
• Improve LLM output error messages
• Update prompt for binary relationships
• Fix delimiter corruption examples
2025-09-14 00:55:15 +08:00
yangdx
4a5ab5121d
Change delimiter from <|S|> to <|#|> and clarify formatting rules
2025-09-13 22:58:56 +08:00
yangdx
244122094d
Merge branch 'optimize-extraction' into return-data-only
2025-09-13 15:38:50 +08:00
yangdx
41cdeaeaad
Add Concept and NaturalObject to default entity types
2025-09-13 15:37:11 +08:00
yangdx
0ffb5d5f2d
Replace search API with aquery_data for consistent raw data retrieval, mirroring aquery results
...
• Reuse existing query logic paths and remove kg_search function entirely
• Update kg_query/naive_query to return raw data as needed
2025-09-13 15:30:29 +08:00
yangdx
c2d064b580
Bump API version to 0221
2025-09-13 14:06:20 +08:00
yangdx
4ce5f9014c
Improve error messages in entity and relationship extraction
2025-09-13 11:20:03 +08:00
yangdx
f3b5352019
Refine default entity types
2025-09-13 11:17:06 +08:00
yangdx
bf423a4ce1
Clarify output structure in prompt instructions by adding field count specifications
2025-09-13 09:51:33 +08:00
yangdx
369f799b16
Refine entity extraction prompts for clarity and consistency
...
• Clarify tuple delimiter usage
• Soften proper noun translation rules
• Standardize language requirements
• Improve output format consistency
2025-09-13 08:14:46 +08:00
yangdx
9a2e8be5a7
Fix extraction validation and delimiter comment accuracy
...
• Change < to != for exact length check
• Fix entity validation from 4 to exact 4
• Fix relationship validation to exact 5
• Correct delimiter comment example
2025-09-12 18:13:25 +08:00
yangdx
8088b7e07a
Fix tuple delimiter corruption handling and update documentation
2025-09-12 18:03:37 +08:00
yangdx
8a3e2c03a9
Fix tuple delimiter corruption patterns with pipes and brackets
...
- Handle <||S||> malformed delimiters
- Fix <||> empty pipe sequences
- Repair <|| incomplete patterns
- Process ||S|| missing brackets
- Improve delimiter normalization
2025-09-12 17:45:32 +08:00
yangdx
43f6fcea6c
Fix linting
2025-09-12 17:00:53 +08:00
yangdx
1ee1fe895b
Merge branch 'qdrant1.7' into optimize-extraction
2025-09-12 16:40:53 +08:00
yangdx
69ca447f45
Sort description by timestamp then description length to improves merge consistency
2025-09-12 13:59:26 +08:00
yangdx
668a7c1f16
Bump API vesrion to 0220
2025-09-12 12:32:42 +08:00