Cải thiện truy xuất đồ thị

Retrieval Augmented Generation (RAG) với LangChain

Meri Nova

Machine Learning Engineer

Kỹ thuật

Hạn chế chính: độ tin cậy khi dịch từ người dùng → Cypher

Chiến lược cải thiện hệ thống truy xuất đồ thị:

  • Lọc lược đồ đồ thị
  • Xác thực truy vấn Cypher
  • Few-shot prompting
Retrieval Augmented Generation (RAG) với LangChain

Lọc

from langchain_community.chains.graph_qa.cypher import GraphCypherQAChain

llm = ChatOpenAI(api_key="...", model="gpt-4o-mini", temperature=0)

chain = GraphCypherQAChain.from_llm(
graph=graph, llm=llm, exclude_types=["Concept"], verbose=True
)
print(graph.get_schema)
Thuộc tính nút:
Document {title: STRING, id: STRING, text: STRING, summary: STRING, source: STRING}
Organization {id: STRING}
Retrieval Augmented Generation (RAG) với LangChain

Xác thực truy vấn Cypher

  • Khó xác định hướng của các quan hệ
chain = GraphCypherQAChain.from_llm(
    graph=graph, llm=llm, verbose=True, validate_cypher=True
)
  1. Phát hiện các nút và quan hệ
  2. Xác định hướng của quan hệ
  3. Kiểm tra lược đồ đồ thị
  4. Cập nhật hướng của quan hệ
Retrieval Augmented Generation (RAG) với LangChain

Few-shot prompting

examples = [
    {
        "question": "How many notable large language models are mentioned in the article?",
        "query": "MATCH (m:Concept {id: 'Large Language Model'}) RETURN count(DISTINCT m)",
    },
    {
        "question": "Which companies or organizations have developed the large language models mentioned?",
        "query": "MATCH (o:Organization)-[:DEVELOPS]->(m:Concept {id: 'Large Language Model'}) RETURN DISTINCT o.id",
    },
    {
        "question": "What is the largest model size mentioned in the article, in terms of number of parameters?",
        "query": "MATCH (m:Concept {id: 'Large Language Model'}) RETURN max(m.parameters) AS largest_model",
    },
]
Retrieval Augmented Generation (RAG) với LangChain

Triển khai few-shot prompting

from langchain_core.prompts import FewShotPromptTemplate, PromptTemplate

example_prompt = PromptTemplate.from_template( "User input: {question}\nCypher query: {query}" )
cypher_prompt = FewShotPromptTemplate( examples=examples, example_prompt=example_prompt, prefix="You are a Neo4j expert. Given an input question, create a syntactically correct Cypher query to run.\n\nHere is the schema information\n{schema}.\n\n Below are a number of examples of questions and their corresponding Cypher queries.", suffix="User input: {question}\nCypher query: ", input_variables=["question"], )
Retrieval Augmented Generation (RAG) với LangChain

Hoàn chỉnh prompt

Bạn là chuyên gia Neo4j. Với một câu hỏi đầu vào, hãy tạo truy vấn Cypher đúng cú pháp để chạy.

Dưới đây là một số ví dụ về câu hỏi và truy vấn Cypher tương ứng.

User input: How many notable large language models are mentioned in the article?
Cypher query: MATCH (p:Paper) RETURN count(DISTINCT p)

User input: Which companies or organizations have developed the large language models?
Cypher query: MATCH (o:Organization)-[:DEVELOPS]->(m:Concept {id: 'Large Language Model'}) RETURN DISTINCT o.id

User input: What is the largest model size mentioned in the article, in terms of number of parameters?
Cypher query: MATCH (m:Concept {id: 'Large Language Model'}) RETURN max(m.parameters) AS largest_model

User input: How many papers were published in 2016?
Cypher query:
Retrieval Augmented Generation (RAG) với LangChain

Thêm các ví dụ few-shot

chain = GraphCypherQAChain.from_llm(
    graph=graph, llm=llm, cypher_prompt=cypher_prompt,
    verbose=True, validate_cypher=True
)
Retrieval Augmented Generation (RAG) với LangChain

Cùng luyện tập nào!

Retrieval Augmented Generation (RAG) với LangChain

Preparing Video For Download...