LLM Note 4
Data Mastery Series — Episode 49: Understanding RAG Taxonomy
LLM Note 4
Data Mastery Series — Episode 49: Understanding RAG Taxonomy

Connect with me and follow our journey: Linkedin, Facebook
1) A Taxonomy of Retrieval Augmented Generation
(SOURCE)เนื้อหาใน Blog จะแบ่งหัวข้อออกเป็น 8 ส่วนหลัก เพื่ออธิบายองค์ประกอบต่างๆ ของ RAG อย่างละเอียด:

Retrieval Augmented Generation enhances the reliability and the trustworthiness in LLM responses (Source: A Taxonomy of Retrieval Augmented Generation)
- RAG Basics (พื้นฐานของ RAG): อธิบายข้อจำกัดของ LLM และแนะนำแนวคิด RAG เบื้องต้น
- Core Components (องค์ประกอบหลัก): อธิบายกระบวนการสร้าง index และ generation
- Evaluation (การประเมินผล): อธิบาย metrics และ frameworks ที่ใช้ในการประเมินประสิทธิภาพของ RAG
- Pipeline Design (การออกแบบ pipeline): อธิบายรูปแบบต่างๆ ของ RAG pipelines (เช่น แบบง่าย, แบบซับซ้อน, แบบ modular)
- Operations Stack (Operations Stack): อธิบายเครื่องมือและเทคนิคที่ใช้ในการจัดการและ monitor RAG systems
- Emerging Patterns (รูปแบบใหม่ๆ): อธิบายแนวทาง RAG ที่กำลังพัฒนา เช่น multimodal RAG, agentic RAG และ KG-powered RAG
- Technology Providers (ผู้ให้บริการเทคโนโลยี): รายชื่อโซลูชันและผู้ให้บริการที่เกี่ยวข้องกับ RAG
- Applied RAG (RAG ในการใช้งานจริง): รูปแบบการใช้งาน RAG, ขอบเขตการใช้งาน, และความท้าทายในปัจจุบัน
① RAG Basics (พื้นฐานของ RAG)
1️⃣ LLM Limitations (ข้อจำกัดของ LLM)
- Knowledge Cut-off Date (ข้อมูลอัพเดตถึงเมื่อไร):
LLMs ถูก train บนข้อมูลจำนวนมหาศาล แต่ข้อมูลนั้นไม่ได้เป็นปัจจุบัน - Training Data Limitation (ข้อจำกัดของข้อมูลที่ใช้ train):
LLMs ถูก train บนข้อมูลสาธารณะเท่านั้น - Hallucinations (การสร้างข้อมูลเท็จ):
LLMs อาจให้ข้อมูลที่ไม่ถูกต้อง แต่ฟังดูน่าเชื่อถือ - Context Window (ความยาวของคำตอบ):
คำตอบของ LLMs ถูกจำกัดตามจำนวน tokens (หน่วยของคำ) ถ้า prompt ยาวเกิน context window ส่วนเกินจะถูกตัดทิ้ง
2️⃣ RAG Concepts (แนวคิด RAG)
- Parametric Memory (หน่วยความจำเชิงพารามิเตอร์):
ความรู้ที่ฝังอยู่ในตัว LLM เอง, ได้มาจากการ train (การสอน) เช่น 1.5 trillion parameters ของ Gemini 1.5 Pro ซึ่งเปลี่ยนแปลงไม่ได้ง่ายๆ - Non-parametric Memory (หน่วยความจำที่ไม่ใช่เชิงพารามิเตอร์):
ข้อมูลที่ LLM ไม่มีอยู่ใน parameters แต่สามารถเข้าถึงได้จากภายนอก เช่น search engine หรือ database RAG ซึ่งสามารถเปลี่ยนแปลงและอัพเดทได้ - Knowledge Base (ฐานความรู้):
แหล่งข้อมูลภายนอกที่สร้างขึ้นสำหรับ RAG application ข้อมูลถูกประมวลผลและจัดเก็บใน persistent memory (เช่นเก็บใน Database สำหรับ RAG) - User Query (คำถามของผู้ใช้):
Prompt ที่ผู้ใช้ส่งไปยัง LLM เพื่อขอคำตอบ - Retrieval (การดึงข้อมูล):
กระบวนการค้นหาและดึงข้อมูลที่เกี่ยวข้องกับ user query จาก knowledge base - Augmentation (การเสริมข้อมูล):
กระบวนการเพิ่มข้อมูลที่ถูกดึงมาลงใน user query - Generation (การสร้างข้อความ):
กระบวนการสร้างผลลัพธ์โดย LLM โดยใช้ augmented prompt - Source Citation (การอ้างอิงแหล่งที่มา):
ความสามารถของ RAG system ในการชี้ไปยังข้อมูลจาก knowledge base ที่ใช้ในการสร้างคำตอบ - Unlimited Memory (หน่วยความจำไม่จำกัด):
แนวคิดที่ว่าสามารถเพิ่ม documents จำนวนเท่าใดก็ได้ลงใน RAG knowledge base

How does RAG help? (Source: A Taxonomy of Retrieval Augmented Generation)
② Core Components (องค์ประกอบหลัก)
1️⃣ Indexing (การสร้างดัชนี)
- Indexing Pipeline (ขั้นตอนการสร้างดัชนี):
ชุดของกระบวนการที่ใช้ในการสร้าง knowledge base สำหรับ RAG applications เป็น pipeline ที่ไม่ใช่ real-time และจะ update knowledge base เป็นช่วงๆ - Source Systems (ระบบต้นทาง):
ที่เก็บข้อมูลดั้งเดิมที่ต้องการใช้สำหรับ RAG application เช่น data lakes, file systems, CMSs, SQL & NoSQL databases - Data Loading (การโหลดข้อมูล):
ขั้นตอนแรกของ indexing pipeline ที่เชื่อมต่อกับ source systems เพื่อดึงและ parse files เพื่อนำข้อมูลไปใช้ใน RAG knowledge base - Metadata (ข้อมูลอธิบายข้อมูล):
ข้อมูลที่อธิบายข้อมูลอื่นๆ ให้ข้อมูล เช่น รายละเอียดของข้อมูล, เวลาที่สร้าง, ผู้สร้าง Metadata ช่วยให้ค้นหาข้อมูลได้ง่ายขึ้นใน RAG - Data Masking (การปกปิดข้อมูล):
การปิดบังข้อมูลที่ละเอียดอ่อน เช่น PII (Personally Identifiable Information) หรือข้อมูลที่เป็นความลับ - Chunking (การแบ่งข้อมูลเป็นชิ้น):
กระบวนการแบ่งข้อความยาวๆ ออกเป็นชิ้นเล็กๆ เพื่อให้จัดการได้ง่ายขึ้น Chunking ช่วยให้ค้นหาได้ง่ายขึ้น และแก้ปัญหา context window limits ของ LLMs - Lost in the middle problem (ปัญหาการหลงลืมตรงกลาง):
LLMs อาจอ่านข้อมูลได้ไม่แม่นยำถ้าข้อมูลที่เกี่ยวข้องอยู่ตรงกลางของ prompt การส่งเฉพาะข้อมูลที่เกี่ยวข้องไปยัง LLM สามารถแก้ปัญหานี้ได้ - Fixed Size Chunking (การแบ่งข้อมูลเป็นชิ้นขนาดคงที่):
กำหนดขนาดของ chunk และ overlap (ส่วนที่ซ้อนทับกัน) ล่วงหน้า - Structure-Based Chunking (การแบ่งข้อมูลตามโครงสร้าง):
แบ่งข้อมูลตามโครงสร้าง เช่น HTML, Markdown, JSON แทนที่จะใช้ขนาดคงที่ - Context-Enriched Chunking (การแบ่งข้อมูลเสริมบริบท):
เพิ่ม summary ของ document ที่ใหญ่กว่าลงในแต่ละ chunk เพื่อเสริม context - Agentic Chunking (การแบ่งข้อมูลแบบ Agent):
สร้าง chunks โดยอิงตามเป้าหมายหรือ task เช่น วิเคราะห์ customer reviews โดยแบ่ง reviews ตาม topic หรือ sentiment - Semantic Chunking (การแบ่งข้อมูลเชิงความหมาย):
แบ่งข้อมูลโดยพิจารณา semantic similarity (ความคล้ายคลึงกันในความหมาย) ระหว่าง sentences - Small to big chunking (การแบ่งข้อมูลจากเล็กไปใหญ่):
แบ่งข้อความเป็นหน่วยเล็กๆ ก่อน (เช่น sentences, paragraphs) แล้วรวมหน่วยเล็กๆ เข้าด้วยกันจนได้ขนาด chunk ที่ต้องการ

Small to big chunking (Source: A Taxonomy of Retrieval Augmented Generation)
- Chunk Size (ขนาดของชิ้นข้อมูล):
ขนาดของ chunk มีผลต่อคุณภาพของ RAG system chunks ขนาดใหญ่ให้ context ที่ดีกว่า แต่ก็มี noise มากกว่า chunks ขนาดเล็กมีข้อมูลที่แม่นยำกว่า แต่อาจพลาดข้อมูลสำคัญ - Metadata Filtering (การกรองด้วย metadata):
เพิ่ม metadata เช่น timestamp, author, category เพื่อกรอง chunks ก่อนทำ similarity search - Metadata Enhancement (การเพิ่มประสิทธิภาพ metadata):
ใช้ metadata ที่ inferred (อนุมาน) ได้ เช่น chunk summary, sentiment, category เพื่อเพิ่มประสิทธิภาพการ retrieval - Parent Child Indexing (การสร้างดัชนีแบบ Parent-Child):
จัดระเบียบ documents ในลักษณะ hierarchical parent document มี overarching themes (หัวข้อหลัก) child documents เจาะลึกรายละเอียด ระบบจะ locate child documents ที่เกี่ยวข้องก่อน แล้วอ้างอิง parent documents เพื่อ context เพิ่มเติม - Embeddings (การแปลงข้อมูลเป็นเวกเตอร์):
การแปลงข้อมูลที่ไม่ใช่ตัวเลข เช่น text หรือ image ให้เป็นรูปแบบตัวเลข (vectors) Embeddings ใช้สำหรับ RAG เพราะช่วยในการสร้าง semantic relationship (ความสัมพันธ์เชิงความหมาย) ระหว่างคำ, phrases และ documents - Cosine Similarity (ความคล้ายคลึงโคไซน์):
คำนวณความคล้ายคลึงกันระหว่าง vectors โดยวัด cosine ของมุมระหว่าง vectors Terms ที่เกี่ยวข้องจะมี cosine similarity ใกล้ 1 terms ที่ไม่เกี่ยวข้องจะมีค่าใกล้ 0 - Word2Vec:
โมเดล shallow neural network ที่ใช้เรียนรู้ word embeddings พัฒนาโดย Google - GloVe:
Unsupervised learning technique ที่พัฒนาโดย Stanford University - FastText:
ส่วนขยายของ Word2Vec พัฒนาโดย Facebook AI Research มีประโยชน์สำหรับการจัดการ misspellings และ rare words - ELMo:
Embeddings from Language Models พัฒนาโดย Allen Institute for AI - BERT:
Bidirectional Encoder Representations from Transformers พัฒนาโดย Google เป็นโมเดลที่ใช้ architecture แบบ Transformers ให้ contextualized word embeddings - Pre-trained Embeddings Models:
Embeddings models ที่ถูก train บนข้อมูลจำนวนมาก สามารถ generalize ได้ดีกับ tasks และ domains ต่างๆ - Vector Databases (ฐานข้อมูลเวกเตอร์):
สร้างขึ้นเพื่อจัดการ high dimensional vectors เช่น embeddings ฐานข้อมูลเหล่านี้ specialize ในการ indexing และจัดเก็บ vector embeddings เพื่อให้ค้นหา semantic search และ retrieval ได้อย่างรวดเร็ว - Vector Indices (ดัชนีเวกเตอร์):
Libraries ที่เน้น core features ของ indexing และ search ไม่ support data management, query processing, interfaces
2️⃣ Generation (การสร้างข้อความ)
- Generation Pipeline (ขั้นตอนการสร้างข้อความ):
ชุดของกระบวนการที่ใช้ในการค้นหาและดึงข้อมูลจาก knowledge base เพื่อสร้าง responses ต่อ user queries - Information Retrieval (IR) (การดึงข้อมูล):
ศาสตร์แห่งการค้นหา ไม่ว่าจะเป็นการค้นหาข้อมูลใน document หรือค้นหา documents เอง - Retriever (ตัวดึงข้อมูล):
องค์ประกอบของ generation pipeline ที่ใช้อัลกอริทึมในการค้นหาและดึงข้อมูลที่เกี่ยวข้องจาก knowledge base - Boolean retrieval (การดึงข้อมูลแบบ Boolean):
การค้นหาแบบ keyword ที่ใช้ Boolean logic เพื่อ match documents กับ queries โดยอิงตาม absence หรือ presence ของคำ - TF-IDF (Term Frequency-Inverse Document Frequency):
Statistical measure ที่ใช้ประเมินความสำคัญของคำใน document เมื่อเทียบกับ collection of documents (corpus)

TF-IDF calculation (Source: A Taxonomy of Retrieval Augmented Generation)
- BM25 (Best Match 25):
Advanced probabilistic model ที่ใช้ rank documents โดยอิงตาม query terms ที่ปรากฏในแต่ละ document ปรับแก้ความยาวของ documents เพื่อไม่ให้ documents ที่ยาวกว่าได้คะแนนสูงกว่าอย่างไม่เป็นธรรม - Static Word Embeddings (Embeddings คำแบบคงที่):
Embeddings เช่น Word2Vec และ GloVe ที่ represent คำเป็น dense vectors ใน continuous vector space capturing semantic relationships - Contextual Embeddings (Embeddings คำตามบริบท):
Embeddings ที่ generate โดย models เช่น BERT หรือ OpenAI’s text embeddings ให้ high-dimensional, context-aware representations

Static vs Contextual Embeddings (Source: A Taxonomy of Retrieval Augmented Generation)
- Learned Sparse Retrieval (การดึงข้อมูลแบบ Sparse ที่เรียนรู้ได้): Generate sparse representations โดยใช้ neural networks
- Dense Retrieval (การดึงข้อมูลแบบ Dense):
Encode queries และ documents เป็น dense vectors - Hybrid Retrieval (การดึงข้อมูลแบบ Hybrid):
รวม sparse และ dense methods - Cross-Encoder Retrieval (การดึงข้อมูลแบบ Cross-Encoder):
เปรียบเทียบ query-document pairs โดยใช้ transformer models - Graph-based Retrieval (การดึงข้อมูลแบบ Graph):
ใช้ graph structures เพื่อ model relationships ระหว่าง documents - Quantum-inspired Retrieval (การดึงข้อมูลแบบ Quantum):
ใช้วิธีการของ quantum computing ในการดึงข้อมูล - Neural IR models (โมเดล IR แบบ Neural):
ใช้วิธีการของ neural network ในการดึงข้อมูล - Augmentation (การเสริมข้อมูล):
กระบวนการรวม user query และ documents ที่ดึงมาจาก knowledge base - Prompt Engineering:
เทคนิคในการให้ instructions แก่ LLM เพื่อให้ได้ผลลัพธ์ที่ต้องการ สร้าง prompts เพื่อให้ LLM สร้าง responses ที่ถูกต้องและเกี่ยวข้อง - Contextual Prompting (การใช้พรอมต์ตามบริบท):
เพิ่ม instruction เช่น “Answer only based on the context provided.” เพื่อให้ LLM focus เฉพาะข้อมูลที่ให้มา - Controlled Generation Prompting (การใช้พรอมต์ควบคุมการสร้างข้อความ):
เพิ่ม instruction เช่น “If the question cannot be answered based on the provided context, say I don’t know.” - Few Shot Prompting (การใช้พรอมต์แบบตัวอย่างน้อย):
ให้ examples ใน prompt เพื่อ guide การ generation ในแบบที่ต้องการ - Chain of Thought Prompting (การใช้พรอมต์แบบลูกโซ่ความคิด):
เพิ่ม intermediate “reasoning” steps เพื่อปรับปรุง performance ของ LLMs ใน tasks ที่ต้องการ complex reasoning - Self Consistency (ความสอดคล้องในตัวเอง):
Sample multiple reasoning paths และใช้ generations ของแต่ละ path เพื่อให้ได้คำตอบที่สอดคล้องกันมากที่สุด - Generated Knowledge Prompting (การใช้พรอมต์แบบสร้างความรู้):
สร้าง knowledge chains โดยใช้ latent knowledge ของ models เพื่อเสริมสร้าง reasoning - Tree of Thoughts Prompting (ToT) (การใช้พรอมต์แบบต้นไม้ความคิด): เทคนิคการ prompt ที่ LLM สร้างความคิดที่เป็นไปได้หลายทาง (เหมือนกิ่งก้านของต้นไม้) แล้วประเมินและเลือกความคิดที่ดีที่สุดเพื่อแก้ปัญหา มีจุดเด่นคือ เหมาะกับปัญหาที่ซับซ้อนที่ต้องการการคิดหลายขั้นตอนและการตัดสินใจ
Note: Chain-of-Thought (CoT) สร้างแค่เส้นทางความคิดเดียว แต่ ToT สร้างหลายเส้นทาง - Automatic Reasoning and Tool-use (ART) (การให้เหตุผลอัตโนมัติและการใช้เครื่องมือ):
Framework ที่ LLM สามารถใช้เครื่องมือภายนอก (เช่น search engine, API) เพื่อช่วยในการให้เหตุผลและแก้ปัญหาที่ซับซ้อน มีจุดเด่นคือ ช่วยให้ LLM สามารถเข้าถึงข้อมูลและดำเนินการที่อยู่นอกเหนือความสามารถของตัวเอง
Note: LLM ทั่วไปจำกัดอยู่แค่ข้อมูลที่ตัวเองมี แต่ ART สามารถใช้เครื่องมือช่วยในการหาข้อมูลและตัดสินใจ - Automatic Prompt Engineer (APE):
Framework ที่ LLM ใช้เพื่อสร้างและเลือก prompt ที่ดีที่สุดสำหรับ task นั้นๆ โดยอัตโนมัติ มีจุดเด่นคือ ช่วยลดภาระในการออกแบบ prompt ด้วยมือ และช่วยให้ได้ prompt ที่มีประสิทธิภาพมากยิ่งขึ้น - Active Prompt:
เทคนิคที่ปรับปรุง Chain-of-Thought โดยการปรับ prompt ให้เข้ากับ task นั้นๆ แบบไดนามิก (ปรับเปลี่ยนตามสถานการณ์) - ReAct Prompting:
เทคนิคที่รวม LLM สำหรับการให้เหตุผลและการดำเนินการ (Action) ไปพร้อมๆ กัน มีจุดเด่นคือ LLM สามารถใช้เครื่องมือภายนอกเพื่อหาข้อมูลและดำเนินการได้ และยังสามารถให้เหตุผลเกี่ยวกับข้อมูลที่ได้มา - Recursive Prompting (การใช้พรอมต์แบบเรียกซ้ำ):
เทคนิคการแก้ปัญหาที่ซับซ้อนโดยการแบ่งออกเป็นปัญหาย่อยๆ แล้วใช้ prompt ในการแก้ปัญหาย่อยๆ เหล่านั้นทีละขั้นตอน ผลลัพธ์จากขั้นตอนก่อนหน้าจะถูกนำไปใช้ในขั้นตอนถัดไป มีจุดเด่นคือ เหมาะกับปัญหาที่ต้องใช้การคิดเชิงประกอบ (compositional generalization) เช่น โจทย์คณิตศาสตร์ หรือการตอบคำถามที่ต้องวิเคราะห์หลายขั้นตอน - Foundation Models (แบบจำลองพื้นฐาน):
LLM ขนาดใหญ่ที่ถูก train บนข้อมูลจำนวนมหาศาล มักถูกนำไป fine-tune เพื่อใช้งานเฉพาะทาง ตัวอย่าง: GPT-4o, Gemini 1.5 Pro etc. - Supervised Fine-Tuning (SFT) (การปรับแต่งแบบ Supervised):
การนำ Foundation Model ที่ train ไว้แล้ว มา train เพิ่มเติมบน labeled dataset (ชุดข้อมูลที่มีการระบุคำตอบที่ถูกต้อง) เพื่อให้ model เก่งในงานเฉพาะทาง เช่น การตอบคำถาม หรือการสร้าง chatbot

Supervised Fine-tuning (SFT) of an LLM (Source: A Taxonomy of Retrieval Augmented Generation)
- Small Language Models (SLMs) (แบบจำลองภาษาขนาดเล็ก):
LLM ที่มีขนาดเล็กกว่า (parameters น้อยกว่า) Foundation Models
③ Evaluation (การประเมินผล)
1️⃣ Metrics (ตัวชี้วัด)
- Evaluation Metrics (ตัวชี้วัดการประเมินผล):
Quantitative measures (มาตรวัดเชิงปริมาณ) ที่ใช้ประเมิน performance ของ retrieval & generation และโดยรวมของ RAG system

Precision & Recall (Source: A Taxonomy of Retrieval Augmented Generation)
- Accuracy (ความแม่นยำ):
สัดส่วนของการทำนายที่ถูกต้อง (ทั้งถูกและไม่ถูก) จากทั้งหมด
Note: ในฐานข้อมูลขนาดใหญ่ ส่วนใหญ่ documents ไม่เกี่ยวข้องกับ query ทำให้ accuracy อาจสูงเกินจริงและ misleading - Precision (ความเที่ยงตรง):
สัดส่วนของ documents ที่ retrieved มา ที่เกี่ยวข้องกับ query จริงๆ หรือเป็นคำตอบของคำถามที่ว่า “จาก documents ที่ retrieved มาทั้งหมด มีกี่ documents ที่เกี่ยวข้องจริงๆ?” ซึ่งเป็น Metrics ที่เน้นความถูกต้องของผลลัพธ์ที่ได้ - Precision@k (ความเที่ยงตรงที่ k):
Precision ที่วัดเฉพาะใน documents ที่ retrieved มา k อันดับแรก - Recall (ความครบถ้วน):
สัดส่วนของ documents ที่เกี่ยวข้องทั้งหมดในฐานข้อมูล ที่ถูก retrieved มา หรือเป็นคำตอบของคำถามที่ว่า “จาก documents ที่เกี่ยวข้องทั้งหมด มีกี่ documents ที่ถูก retrieved มาจริงๆ?” ซึ่งเน้นความครบถ้วนของ documents ที่ retrieved มา - F1-score:
ค่าเฉลี่ย harmonic ของ Precision และ Recall ซึ่งเน้นความ balance ทั้ง Precision และ Recall - Mean Reciprocal Rank (MRR) (ค่าเฉลี่ยของส่วนกลับของอันดับ):
ค่าเฉลี่ยของส่วนกลับ (reciprocal) ของอันดับ (rank) ของผลลัพธ์ที่เกี่ยวข้อง (relevant) อันแรก ในแต่ละ query มีจุดเด่นคือ ให้ความสำคัญกับอันดับของผลลัพธ์ที่เกี่ยวข้องอันแรก หากผลลัพธ์ที่เกี่ยวข้องอันแรกอยู่ในอันดับต้นๆ จะได้คะแนนสูง - Mean Average Precision (MAP) (ค่าเฉลี่ยของความเที่ยงตรงเฉลี่ย):
ค่าเฉลี่ยของ Average Precision (AP) ในแต่ละ query โดย AP คือค่าที่รวม Precision และ Recall ที่ cut-off ต่างๆ (k อันดับแรก) มีจุดเด่นคือ พิจารณาทั้งความถูกต้อง (Precision) และความครบถ้วน (Recall) ของผลลัพธ์ที่ได้จากหลายๆ อันดับ - Normalised Discounted Cumulative Gain (nDCG):
วัดคุณภาพการจัดอันดับ โดยให้คะแนนสูงแก่ผลลัพธ์ที่เกี่ยวข้องที่อยู่ในอันดับต้นๆ และให้คะแนนลดหลั่นลงไปตามอันดับที่ต่ำลง (discounted) แล้ว normalize ค่า มีจุดเด่นคือ สามารถจัดการกับสถานการณ์ที่ผลลัพธ์มีความเกี่ยวข้องในระดับต่างๆ กันได้

Calculating nDCG (Source: A Taxonomy of Retrieval Augmented Generation)
- Context relevance (ความเกี่ยวข้องของบริบท):
ประเมินว่า documents ที่ retrieved มาเกี่ยวข้องกับ query เดิมมากน้อยแค่ไหน โดยดูที่ topical alignment (ความสอดคล้องของหัวข้อ), information usefulness (ประโยชน์ของข้อมูล) และ redundancy (ความซ้ำซ้อน) โดยมีหลักการว่า retrieved context ควรมีข้อมูลที่เกี่ยวข้องกับ query เท่านั้น - Answer Faithfulness (ความน่าเชื่อถือของคำตอบ):
วัดว่าคำตอบที่สร้างขึ้นมีความถูกต้องตามข้อเท็จจริง (factually grounded) ใน retrieved context มากน้อยแค่ไหน โดยมีหลักการว่า ข้อเท็จจริงในคำตอบต้องไม่ขัดแย้งกับ context และสามารถ traced back to the source (อ้างอิงแหล่งที่มาได้) - Hallucination Rate (อัตราการสร้างข้อมูลเท็จ):
สัดส่วนของ generated claims (ข้อความที่สร้างขึ้น) ในคำตอบที่ไม่อยู่ใน retrieved context โดยวัดว่า LLM สร้างข้อมูลที่ไม่เกี่ยวข้องกับข้อมูลที่ retrieved มามากน้อยแค่ไหน - Coverage (ความครอบคลุม):
วัดว่าข้อมูลที่เกี่ยวข้องจาก retrieved passages ถูกรวมอยู่ในคำตอบมากน้อยแค่ไหน โดยมีหลักการว่า คำตอบควรครอบคลุมข้อมูลที่สำคัญจาก retrieved passages - Answer Relevance (ความเกี่ยวข้องของคำตอบ):
วัดว่าคำตอบที่สร้างขึ้นมีความเกี่ยวข้องกับ query มากน้อยแค่ไหน โดยดูที่ system’s ability to comprehend the query (ความสามารถในการเข้าใจ query), response being pertinent to the query (คำตอบที่ตรงประเด็น) และ completeness of the response (ความสมบูรณ์ของคำตอบ) - Ground truth (ข้อมูลที่แท้จริง):
ข้อมูลที่รู้ว่าเป็นจริง ใช้เป็น benchmark (เกณฑ์มาตรฐาน) ในการประเมิน - Human Evaluation (การประเมินโดยมนุษย์):
ผู้เชี่ยวชาญ (Subject matter expert) ดู documents และ determine the relevance and accuracy of the outputs (ตัดสินความเกี่ยวข้องและความถูกต้องของผลลัพธ์) outputs - Noise Robustness (ความทนทานต่อสัญญาณรบกวน):
ความสามารถของระบบ RAG ในการแยก noisy documents (documents ที่เกี่ยวข้องกับ query แต่ไม่มีข้อมูลที่เป็นประโยชน์) ออกจาก relevant ones - Negative Rejection (การปฏิเสธเชิงลบ):
ความสามารถของระบบ RAG ในการ “ไม่ให้คำตอบ” เมื่อไม่มีข้อมูลที่เกี่ยวข้องกับ query ใน documents ใน knowledge base - Information Integration (การบูรณาการข้อมูล):
ความสามารถของระบบในการ assimilate information (รวมข้อมูล) จาก multiple documents เพื่อตอบ query ได้อย่างครอบคลุม - Counterfactual Robustness (ความทนทานต่อข้อเท็จจริงที่ขัดแย้งกัน): ความสามารถของระบบ RAG ในการ address (จัดการ) และ reject known inaccuracies (ปฏิเสธความไม่ถูกต้องที่รู้แล้ว) ใน retrieved information
2️⃣ Frameworks (เฟรมเวิร์ก)
- Frameworks (เฟรมเวิร์ก):
เครื่องมือ (Tools) ที่ออกแบบมาเพื่อช่วยในการ evaluation โดยมี automation ของ evaluation process และ data generation - RAGAS (Retrieval Augmented Generation Assessment):
framework ที่พัฒนาโดย Exploding Gradients ที่ assesses retrieval และ generation components ของ RAG systems โดยไม่ต้องใช้ human annotations มากนัก - Synthetic Test Dataset Generation (การสร้างชุดข้อมูลทดสอบสังเคราะห์):
การใช้ models เช่น LLMs ในการ automatically generate ground truth data จาก knowledge base

Synthetic Data Generation in RAGAS (Source: A Taxonomy of Retrieval Augmented Generation)
- LLM as a judge (LLM เป็นผู้ตัดสิน):
การใช้ LLM ในการ evaluate a task (ประเมินงาน) - ARES (Automated RAG evaluation system):
framework ที่พัฒนาโดย researchers ที่ Stanford University และ Databricks โดยใช้หลักการ LLM as a judge
3️⃣ Benchmarks (เกณฑ์มาตรฐาน)
- Benchmarks (เกณฑ์มาตรฐาน):
Standardised datasets และ evaluation metrics ที่ใช้ measure performance ของ RAG systems ซึ่งช่วย provide a common ground สำหรับ comparing different RAG approaches และ ensure consistency across evaluations โดย considering fixed tasks และ evaluation criteria - BEIR (Benchmarking Information Retrieval):
A comprehensive heterogeneous benchmark ที่ based on 9 IR tasks และ 19 Question-Answer datasets

BEIR — 9 tasks and 18 (of 19) datasets (Source: BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models)
- Retrieval Augmented Generation Benchmark (RGB):
A benchmark ที่ focusses on 4 key abilities ของ RAG system — 1) Noise Robustness, 2) Negative Rejection, 3) Information Integration และ 4) Counterfactual Robustness

Four abilities required of RAG systems (Source: Benchmarking Large Language Models in Retrieval-Augmented Generation, Chen et al )
- Multihop RAG:
a benchmark ที่ contains queries ที่ requires reasoning across multiple documents โดยจะ queries involve document metadata, reflecting complex scenarios ที่ commonly found ใน real-world RAG applications - Comprehensive RAG (CRAG):
a benchmark ที่ focusses on factual question answering และ simulates web และ Knowledge Graph (KG) search โดยจะ contains 8 types ของ queries across 5 domains
④ Pipeline Design (การออกแบบไปป์ไลน์)
1️⃣ Naive RAG (RAG แบบดั้งเดิม)
- Naive RAG:
RAG แบบพื้นฐานที่สุด ทำงานเป็นเส้นตรง (linear) ตามลำดับ: indexing -> retrieval -> augmentation -> generation - Retrieve-Read:
A retriever ที่ดึงข้อมูล และ LLM อ่านข้อมูลนั้นเพื่อสร้างคำตอบ - RAG Failure Points (จุดบกพร่องของ RAG):
- retriever ดึงข้อมูลไม่ครบ หรือดึงข้อมูลที่ไม่เกี่ยวข้อง
- LLM ไม่สนใจ context ที่ให้มา หรือเลือกข้อมูลที่ไม่เกี่ยวข้องจาก context - Disjointed Context (บริบทที่ไม่ต่อเนื่อง):
ข้อมูลมาจากหลายแหล่ง ทำให้การเชื่อมต่อระหว่าง chunks ไม่ราบรื่น - Over-reliance on Context (การพึ่งพาบริบทมากเกินไป):
LLM ลืมความรู้เดิมของตัวเอง (parametric memory)
2️⃣ Advanced RAG (RAG ขั้นสูง)
- Advanced RAG:
RAG ที่ปรับปรุงขั้นตอนต่างๆ (pre-retrieval, retrieval, post-retrieval) เพื่อแก้ข้อจำกัดของ Naive RAG

Advanced RAG as Rewrite-Retrieve-Rerank-Read pattern (Source: A Taxonomy of Retrieval Augmented Generation)
- Rewrite-Retrieve-Rerank-Read:
เพิ่มขั้นตอนการ rewrite (เขียนใหม่) query และ rerank (จัดอันดับใหม่) ผลลัพธ์ - Index Optimisation (การปรับปรุงดัชนีให้ดีที่สุด):
จัดเตรียม knowledge base ให้ดีขึ้นสำหรับการ retrieval - Query Optimisation (การปรับปรุงคำถามให้ดีที่สุด):
ปรับ query ให้เหมาะกับการ retrieval - Query Expansion (การขยายคำถาม):
เพิ่มคำใน query เพื่อให้ได้ข้อมูลที่เกี่ยวข้องมากขึ้น
- Multi-query expansion: สร้างหลาย variations ของ query
- Sub-query expansion: แบ่ง query ซับซ้อน เป็น queries ย่อยๆ
- Step back expansion: สร้าง query ที่เป็น conceptual query (query เชิงความคิดรวบยอด) - Query Transformation (การแปลงคำถาม):
ใช้ query ที่ transformed (แปลงแล้ว) แทน query เดิม - Query Rewriting (การเขียนคำถามใหม่):
เขียน query ใหม่จาก input ที่อาจไม่ใช่ query โดยตรง - Hypothetical document embedding, HyDE (การฝังเอกสารสมมติ): LLM สร้าง hypothetical answer (คำตอบสมมติ) แล้วใช้ answer นั้นในการค้นหา
- Query Routing (การจัดเส้นทางคำถาม):
ส่ง query ไปยัง workflow ที่เหมาะสม ตาม criteria ต่างๆ - Hybrid Retrieval (การดึงข้อมูลแบบผสมผสาน):
ใช้หลาย retrieval methods ร่วมกัน (เช่น keyword-based search + semantic similarity)

Hybrid of sparse, dense and graph retrieval (Source: A Taxonomy of Retrieval Augmented Generation)
- Iterative Retrieval (การดึงข้อมูลแบบวนซ้ำ):
ดึงข้อมูลหลายครั้ง โดยใช้ generated response ในการดึงข้อมูลครั้งต่อไป - Recursive Retrieval (การดึงข้อมูลแบบเรียกซ้ำ):
เหมือน Iterative Retrieval แต่มีการ transform query หลังจากการ generate แต่ละครั้ง - Adaptive Retrieval (การดึงข้อมูลแบบปรับตัว):
LLM กำหนดเวลาและ content ที่เหมาะสมสำหรับการ retrieval - Contextual Compression (การบีบอัดบริบท):
ลดความยาวของข้อมูลที่ retrieved โดยเลือกเฉพาะส่วนที่เกี่ยวข้อง - Explanation: Reducing the amount of retrieved information to only the relevant parts.
- Reranking (การจัดอันดับใหม่):
จัดอันดับ retrieved information เพื่อเลือก documents ที่เกี่ยวข้องที่สุด
3️⃣ Modular RAG (RAG แบบแยกส่วน)
- Modular RAG:
แบ่งโครงสร้าง RAG แบบเดิม (monolithic) ออกเป็น components ที่สามารถสลับเปลี่ยนได้ (interchangeable) เพื่อให้ปรับแต่งระบบได้ตาม use cases - Search Module (โมดูลการค้นหา):
ค้นหาข้อมูลจากแหล่งต่างๆ - RAG-Fusion:
ปรับปรุงระบบ search เดิม โดยใช้ multi-query approach - Memory Module (โมดูลหน่วยความจำ):
ใช้ประโยชน์จาก “memory” ของ LLM (ความรู้ที่อยู่ใน parameters) - Routing (การจัดเส้นทาง):
นำทาง query ผ่านแหล่งข้อมูลต่างๆ เพื่อเลือก pathway ที่เหมาะสม - Task Adapter (ตัวปรับงาน):
ทำให้ RAG ปรับตัวเข้ากับ downstream tasks ต่างๆ ได้ เช่น summarisation, translation
⑤ Operations Stack (Operations Stack)
1️⃣ Critical Layers (เลเยอร์ที่สำคัญยิ่ง)
- Critical Layers (เลเยอร์ที่สำคัญยิ่ง):
องค์ประกอบพื้นฐานที่ RAG system ขาดไม่ได้ ถ้าขาดไป ระบบจะล้มเหลว - Data Layer (เลเยอร์ข้อมูล):
สร้างและจัดเก็บ knowledge base โดย collect data จากแหล่งต่างๆ, transform ให้เป็น format ที่ใช้งานได้, และ store เพื่อ efficient retrieval - Model Layer (เลเยอร์โมเดล):
- จัดเก็บ model library, training & fine-tuning components, และ inference optimisation components
- ทำให้ LLM สามารถ generate คำตอบได้อย่างรวดเร็วและคุ้มค่า

Model Layer of the RAGOps stack (Source: A Taxonomy of Retrieval Augmented Generation)
- Fully managed deployment (การ deployment แบบจัดการเต็มรูปแบบ):
ผู้ให้บริการจัดการ infrastructure ทั้งหมด - Self-hosted deployment (การ deployment แบบ self-hosted):
ผู้พัฒนาจัดการ infrastructure เอง บน private clouds หรือ on-premises - Local/edge deployment (การ deployment แบบ local/edge):
- รัน model บน local hardware หรือ edge devices เพื่อ data privacy, reduced latency, และ offline functionality
- Application Orchestration Layer (เลเยอร์ประสานงานแอปพลิเคชัน): จัดการ interactions ระหว่าง layers ต่างๆ, เป็น central coordinator ที่ enable communication ระหว่าง data, retrieval systems, generation models, และ services อื่นๆ
2️⃣ Essential Layers (เลเยอร์ที่จำเป็น)
- Essential Layers (เลเยอร์ที่จำเป็น):
เลเยอร์ที่ focus บน performance, reliability, และ safety ของระบบ ทำให้ระบบมีมาตรฐานและให้ value แก่ผู้ใช้ - Prompt Layer (เลเยอร์พรอมต์):
จัดการ augmentation และ LLM prompts ต่างๆ - Evaluation Layer (เลเยอร์การประเมินผล):
จัดการ regular evaluation ของ retrieval accuracy, context relevance, faithfulness, และ answer relevance - Monitoring Layer (เลเยอร์การตรวจสอบ):
- Continuous monitoring เพื่อ long-term health ของ RAG system
- เข้าใจ system behaviour, identify points of failure, assess relevance & adequacy of information, และ track system metrics ต่างๆ - LLM Security & Privacy Layer (เลเยอร์ความปลอดภัยและความเป็นส่วนตัวของ LLM):
- Ensure data privacy และ protection ด้วย strategies เช่น anonymisation, encryption, differential privacy, query validation & sanitisation, และ output filtering
- Implement guardrails, access controls, monitoring, และ auditing - Caching Layer (เลเยอร์แคช):
ลด cost และ latency โดย store frequently accessed data

RAGOps stack with critical and essential layers (Source: A Taxonomy of Retrieval Augmented Generation)
3️⃣ Enhancement Layers (เลเยอร์เสริม)
- Enhancement Layer (เลเยอร์เสริม):
เลเยอร์ที่ improving efficiency, scalability, และ usability ของระบบ เลือกใช้ตาม end requirements - Human-in-the-loop Layer (เลเยอร์มนุษย์ในวงจร):
ให้ human judgment ใน use-cases ที่ต้องการ higher accuracy หรือ ethical considerations - Cost Optimisation Layer (เลเยอร์การปรับต้นทุนให้เหมาะสม):
Manage resources efficiently สำหรับ large-scale systems - Explainability and Interpretability Layer (เลเยอร์ความสามารถในการอธิบายและตีความ):
Provide transparency สำหรับ system decisions ใน domains ที่ requiring accountability - Collaboration and Experimentation Layer (เลเยอร์การทำงานร่วมกันและการทดลอง):
Useful สำหรับ teams ที่ทำงานบน development และ experimentation
⑥ Emerging Patterns (รูปแบบที่เกิดขึ้นใหม่)
1️⃣ Knowledge Graphs (กราฟความรู้)
- Knowledge Graph powered RAG (RAG ที่ขับเคลื่อนด้วยกราฟความรู้):
ใช้โครงสร้าง knowledge graph เพื่อเพิ่ม contextual understanding, enhanced reasoning capabilities, และ improved explainability - Knowledge Graphs (กราฟความรู้):
จัดระเบียบข้อมูลเป็น entities (วัตถุ, แนวคิด) และ relationships (ความสัมพันธ์) ในรูปแบบ structured manner - GraphRAG:
An open-source framework ที่สร้าง knowledge graphs จาก source documents โดยอัตโนมัติ และใช้ knowledge graph นั้นในการ retrieval - Graph Communities (ชุมชนกราฟ):
แบ่ง entities และ relationships เป็นกลุ่มๆ - Community Summaries (บทสรุปชุมชน):
LLM generate summaries สำหรับ communities เพื่อ insights into topical structure และ semantics - Local Search (การค้นหาในท้องถิ่น):
หา set ของ entities ที่ semantically-related กับ user input จาก knowledge graph - Global Search (การค้นหาระดับโลก):
Similarity based search บน community summaries - Ontology (ออนโทโลยี):
A formal representation of knowledge as a set of concepts ภายใน domain, และ relationships ระหว่าง concepts เหล่านั้น
2️⃣ Multimodal (มัลติโมดอล)
- Multimodal RAG (RAG แบบมัลติโมดอล):
ใช้ modalities อื่นๆ นอกเหนือจาก text (เช่น images, audio, video) ใน both retrieval และ generation - Modality (โมดอลลิตี้):
Specific type ของ input data (เช่น text, image, video, audio) - Multimodal Embeddings (Embeddings แบบมัลติโมดอล):
A unified vector representation ที่ encode multiple data types (เช่น text และ image embeddings combined) - CLIP (Contrastive Language-Image Pre-training):
A model ที่ learns visual concepts จาก natural language supervision, ใช้สำหรับการ cross-modal retrieval และ generation - Contrastive Learning (การเรียนรู้แบบเปรียบเทียบ):
A learning method ที่ align data across different modalities โดย bringing semantically similar data points closer ใน shared embedding space
3️⃣ Agentic (แบบ Agent)
- Agentic RAG (RAG แบบ Agent):
Leverage LLM based agents สำหรับ adapting RAG workflow to query types และ type ของ documents ใน knowledge base - Adaptive Frameworks (เฟรมเวิร์กแบบปรับตัว):
Dynamic systems ที่ adjust retrieval และ generation strategies based บน evolving context และ data - Routing Agents (เอเจนต์การจัดเส้นทาง):
Agents responsible สำหรับ direct user queries to the most appropriate sources หรือ sub-systems - Query Planning Agents (เอเจนต์การวางแผนคำถาม):
Agents ที่ break down complex queries into sub-queries และ manage execution across different retrieval pipelines - Multiple Vectors per Document (หลายเวกเตอร์ต่อเอกสาร):
A technique ที่ multiple vector representations are generated multiple vector representations สำหรับ each document เพื่อ capture different aspects ของ content
⑦ Technology Providers (ผู้ให้บริการเทคโนโลยี)
- Model Access, Training & FineTuning (การเข้าถึงโมเดล, การฝึกฝน, และการปรับแต่ง):
ผู้ให้บริการที่ให้ access เข้าถึง, train, และ fine-tune LLMs รวมถึง cloud providers และบริษัทที่พัฒนา LLMs เอง เช่น OpenAI, HuggingFace ซึ่งเป็นแหล่งของ models ที่ใช้ใน RAG systems - Data Loading (การโหลดข้อมูล):
ผู้ให้บริการ tools และ services สำหรับ loading และ preparing data เพื่อใช้ใน RAG systems เช่น LlamaIndex, LangChain ซึ่งช่วยให้การเตรียมข้อมูลเป็นไปอย่างมีประสิทธิภาพ - Vector DB and Indexing (ฐานข้อมูลเวกเตอร์และการทำดัชนี):
ผู้ให้บริการ vector databases และ indexing solutions ที่สำคัญสำหรับการ efficient retrieval ใน RAG systems เช่น Pinecone, Chroma ซึ่งเป็น infrastructure หลักสำหรับการค้นหาข้อมูลที่เกี่ยวข้อง - Application Framework (เฟรมเวิร์กแอปพลิเคชัน):
Frameworks ที่ provide tools และ components สำหรับ building RAG applications เช่น LangChain, LlamaIndex, CrewAI (Agentic Orchestration), LangGraph (Agentic Orchestration) ซึ่งช่วยลดความซับซ้อนในการพัฒนาระบบ RAG - Prompt Engineering (วิศวกรรมพรอมต์):
ผู้ให้บริการ tools และ platforms สำหรับ prompt engineering ที่ช่วย optimize LLM performance เช่น W&B (Weights & Biases), PromptLayer ซึ่ง prompt ที่ดีช่วยให้ LLM ทำงานได้ดีขึ้น - Deployment Frameworks (เฟรมเวิร์กการ Deployment):
Frameworks ที่ facilitate deployment ของ LLMs และ RAG systems เช่น MLflow ซึ่งช่วยให้การนำ LLMs และ RAG ไปใช้งานจริงง่ายขึ้น - Deployment & Inferencing (การ Deployment และการอนุมาน):
Cloud providers และ services ที่ offer deployment และ inference capabilities สำหรับ LLMs เช่น AWS, GCP, OpenAI API, Azure ซึ่งเป็น infrastructure สำหรับรัน LLMs และ RAG systems - Monitoring (การตรวจสอบ):
ผู้ให้บริการ monitoring tools และ services ที่ช่วย ensure health และ performance ของ RAG systems เช่น HoneyHive, TruEra ซึ่งช่วยให้ระบบทำงานได้อย่างราบรื่นและมีประสิทธิภาพ - Proprietary LLMs/VLMs (LLMs/VLMs ที่เป็นกรรมสิทธิ์):
รายชื่อ LLMs และ Vision Language Models (VLMs) ที่เป็นกรรมสิทธิ์ (closed-source) เช่น GPT series by OpenAI, Gemini series by Google ซึ่งเป็นแหล่งของ models ที่มี license กำกับ - Open Source LLMs (LLMs โอเพนซอร์ส):
รายชื่อ LLMs ที่เป็น open source สามารถนำไปใช้และแก้ไขได้อย่างอิสระ เช่น Llama series by Meta ซึ่งเป็นทางเลือกสำหรับผู้ที่ต้องการความยืดหยุ่นและควบคุม - Small Language Models (แบบจำลองภาษาขนาดเล็ก):
รายชื่อ LLMs ที่มีขนาดเล็กกว่า เหมาะกับการใช้งานที่ต้องการความรวดเร็วและประหยัดทรัพยากร เช่น Gemma series by Google AI ซึ่งเป็นทางเลือกสำหรับ edge devices และสภาพแวดล้อมที่มีทรัพยากรจำกัด - Managed RAG solutions (โซลูชัน RAG ที่มีการจัดการ):
โซลูชัน RAG ที่มีการจัดการทั้งหมด ช่วยให้การ setup และ management ของ RAG systems ง่ายขึ้น เช่น OpenAI File Search, Azure AI File Search ซึ่งเหมาะสำหรับผู้ที่ต้องการความสะดวกและไม่ต้องดูแล infrastructure เอง - Knowledge Graph and Ontology (กราฟความรู้และ Ontology):
ผู้ให้บริการ knowledge graph databases และ ontology management tools เช่น Neo4j ซึ่งเป็น infrastructure สำหรับสร้างและจัดการ knowledge graphs - Security and Privacy (ความปลอดภัยและความเป็นส่วนตัว):
ผู้ให้บริการ security และ privacy solutions สำหรับ AI systems รวมถึง RAG เช่น Hazy, Duality, BigID ซึ่งช่วย protect ข้อมูลและ ensure compliance - Synthetic Data (ข้อมูลสังเคราะห์):
ผู้ให้บริการ synthetic data generation tools ที่สามารถใช้ augment training datasets และ improve model performance เช่น Mostly AI, Tonic.ai, Synthesis AI ซึ่งช่วยเพิ่มปริมาณและคุณภาพของข้อมูลสำหรับ training - Others (อื่นๆ):
หมวดหมู่สำหรับ technologies และ services อื่นๆ ที่เกี่ยวข้องกับ RAG เช่น Cohere reranker, Unstructured.io
⑧ Applied RAG (RAG ที่นำไปใช้)
1️⃣ Other RAG Patterns (รูปแบบ RAG อื่นๆ)
- Corrective RAG (RAG ที่แก้ไขได้):
ดึงข้อมูล real-time เพื่อ check factual accuracy ของ LLM generated answer (ใช้ในการ fact-checking, medical & legal domains) - Contrastive RAG (RAG แบบเปรียบเทียบ):
ใช้ contrastive learning เพื่อ enhance retrieval process โดยแยก relevant และ irrelevant documents - Selective RAG (RAG แบบคัดเลือก):
Optimise retrieval phase โดย determine when it is beneficial to retrieve external information (ใช้ใน context ที่ retrieval อาจไม่มี value) - RAG with Active Learning (RAG กับการเรียนรู้เชิงรุก):
ใช้ user feedback เพื่อ fine-tune หรือ adapt retrieval process over time (ใช้ใน continuous improvement systems เช่น recommendation engines) - Personalised RAG (RAG ส่วนบุคคล):
ใช้ user preferences, behaviour, และ historical interactions เพื่อ personalise retrieval process (ใช้ใน personalization-heavy domains เช่น recommendation engines, customer service) - Self-RAG:
Adaptive retrieval mechanism ที่ selectively decides when to retrieve knowledge based บน query’s context - RAFT (Retrieval-Augmented Fine-Tuning):
Combine retrieval mechanisms with traditional fine-tuning techniques - RAPTOR (Recursive Abstractive Processing for Tree-Organised Retrieval):
สร้าง recursive, tree-like structure จาก documents เพื่อ improve context-aware information retrieval
2️⃣ Application Areas (ขอบเขตการใช้งาน)
- Search Engine (เครื่องมือค้นหา):
ใช้ RAG เพื่อ present coherent text ใน natural language with source citation

Prominent Search Engines using RAG (Source: Google.com, Bing.com, Perplexity.ai, OpenAI.com)
- Personalised Marketing Content Generation (การสร้างเนื้อหาทางการตลาดส่วนบุคคล):
Content can be personalised to readers, incorporate real-time trends, and be contextually appropriate - Personalised Learning Plans (แผนการเรียนรู้ส่วนบุคคล):
สร้าง personalised learning paths based บน past trends และ automated evaluation and feedback - Real-time Event Commentary (ความคิดเห็นเกี่ยวกับเหตุการณ์แบบเรียลไทม์):
Connect to real-time updates/data via APIs และ pass information to LLM to create virtual commentator - Conversational agents (ตัวแทนสนทนา):
Customise LLMs to product/service manuals, domain knowledge, guidelines, etc. และ serve as support agents - Document Question Answering Systems (ระบบตอบคำถามเกี่ยวกับเอกสาร):
Answer all questions about the organisation with access to proprietary documents - Virtual Assistants (ผู้ช่วยเสมือน):
Enhance user’s experience with more context บน user behaviour using RAG
3️⃣ Applied RAG Challenges (ความท้าทายในการใช้ RAG)
- Relevance Mismatch (ความไม่ตรงกันของความเกี่ยวข้อง):
Difficulty retrieving most relevant documents due to suboptimal ranking - Over-Retrieval (การดึงข้อมูลมากเกินไป):
Retrieving too many documents, leading to unnecessary noise - Sparse vs Dense Retrieval Trade-off (การแลกเปลี่ยนระหว่างการดึงข้อมูลแบบเบาบางและหนาแน่น):
Balancing between sparse retrieval และ dense retrieval เพื่อ maximise relevance - Document Question Answering Systems (ระบบตอบคำถามเกี่ยวกับเอกสาร):
Delays due to retrieval จาก large knowledge bases - Latency (ความหน่วง): Retrieval จาก large หรือ distributed knowledge bases can introduce significant delays affecting real-time applications
- Cost of Storage (ค่าใช้จ่ายในการจัดเก็บ):
Maintaining massive vector databases can be expensive - Narrow Retrieval Focus (การโฟกัสการดึงข้อมูลที่แคบ):
Difficulty retrieving diverse perspectives - Bias in Retrieval (อคติในการดึงข้อมูล):
Biases ใน retrieval results based บน structure ของ data - Context Loss in Long Queries (การสูญเสียบริบทในคำถามยาว):
Loss of context when handling long, multi-turn queries - Incoherent Summarisation (การสรุปที่ไม่สอดคล้องกัน):
Generating inconsistent summaries จาก multiple documents - Over-Generation (การสร้างข้อมูลมากเกินไป):
Generating overly verbose responses - Inconsistent Modal Alignment (การจัดแนว Modal ที่ไม่สอดคล้องกัน): Challenges integrating multimodal data
- Data Silos (ไซโลข้อมูล):
Knowledge is fragmented across multiple sources - Processing Large-Scale Data (การประมวลผลข้อมูลขนาดใหญ่): Difficulty maintaining high throughput as data grows
- Multi-Agent Coordination (การประสานงานหลาย Agent):
Complex coordination among multiple agents - Inefficient Query Routing (การจัดเส้นทางคำถามที่ไม่มีประสิทธิภาพ): Routing queries to wrong sources
- Data Poisoning Attacks (การโจมตีด้วยการวางยาข้อมูล):
External sources feed biased data into the generation pipeline - Adversarial Attacks (การโจมตีแบบ Adversarial):
Attackers influence retrieval or generation results - Knowledge Base Updating (การอัปเดตฐานความรู้): Maintaining an up-to-date knowledge base
- Memory Retention (การเก็บรักษาความทรงจำ):
Ensuring system can store and retrieve long-term memory
หวังว่าบทความนี้จะช่วยให้คุณเข้าใจ RAG ได้อย่างลึกซึ้งและนำไปประยุกต์ใช้ได้อย่างมีประสิทธิภาพนะครับ แล้วพบกันใหม่ในบทความต่อไป 🚀
Data Science Explore the world of data science with Donato_Story
Dashboard Discover the power of data visualization with Donato_Story
Donato_Journey Join me on my journey (Thai version)
Course_Review Discover the training courses with Donato_Story (Thai version)
Let’s Connect!
Your thoughts and feedback are invaluable. Feel free to share them in the comments or connect with me on
- Medium: medium.com/donato-story
- Facebook: web.facebook.com/DonatoStory
- Linkedin: linkedin.com/in/nattapong-thanngam
Originally published on Medium
Related
Corrective RAG
Data Mastery Series — Episode 53: RAG ที่ “คิด” ก่อน “ตอบ” และ “แก้ไข” เมื่อผิดพลาด
Hierarchical Multi-Agent Systems
Data Mastery Series — Episode 59: การสร้างระบบ AI ทีมงานด้วย Supervisor Agent กับทีมย่อย
LangGraph Introduction
Data Mastery Series — Episode 50: Next-Level Chat with Document
LLM Note 2
Data Mastery Series — Episode 47: Summarization Techniques and Advanced RAG