-
GenAI – What Is ‘Trie’ Prefix Tree ?
GenAI – What Is ‘Trie’ Prefix Tree ? Table Of Content: What Is Trie ? Key Properties Of Trie. Example Of Trie. What Can We Do With Trie? Why Not Just Use A List ? What Is Prefix Lookup? How ‘Trie’ Helps Text Tokenization ? (1) What Is ‘Trie’ ? (2) Key Properties Of ‘Trie’ . (3) How ‘Trie’ Works Internally ? (4) How “New York City” Will Get Inserted ? ['Ney York City', 'India', 'South Africa'] (5) What Happens If The Word Does Not Exist ? (6) How Trie Helps In Text Tokenization ? (7) How Trie will give
-
GenAI – Audio File Chunking Process
GenAI – Audio File Chunking Process Table Of Content: How To Chunk Audio Files ? (1) Reference Links https://github.com/infiniflow/ragflow/blob/main/rag/app/audio.py (2) How To Chunk Audio Files ? Imported File Links: https://github.com/infiniflow/ragflow/blob/main/api/db/__init__.py https://github.com/infiniflow/ragflow/blob/main/rag/nlp/rag_tokenizer.py https://github.com/infiniflow/ragflow/blob/main/api/db/services/llm_service.py https://github.com/infiniflow/ragflow/blob/main/rag/nlp/__init__.py import re from api.db import LLMType from rag.nlp import rag_tokenizer from api.db.services.llm_service import LLMBundle from rag.nlp import tokenize def chunk(filename, binary, tenant_id, lang, callback=None, **kwargs): doc = { "docnm_kwd": filename, "title_tks": rag_tokenizer.tokenize(re.sub(r".[a-zA-Z]+$", "", filename)) } doc["title_sm_tks"] = rag_tokenizer.fine_grained_tokenize(doc["title_tks"]) # is it English eng = lang.lower() == "english" # is_english(sections) try: callback(0.1, "USE Sequence2Txt LLM to transcription the audio") seq2txt_mdl = LLMBundle(tenant_id, LLMType.SPEECH2TEXT, lang=lang) ans = seq2txt_mdl.transcription(binary) callback(0.8,
-
Python – What Is Enum Classes ?
Python – What Is Enum Class ? Table Of Content: What Is Enum ? Why Do We Use Enum ? Example Of Enum Class ? What Does StrEnum Means ? Where Do Enums Shine ? Benefits Of Enum ? Summary. (1) What Is Enum ? (2) Why Do We Use Enum ? (4) Example Of Enum. (5) What Is StrEnum ? Example-1: Without StrEnum from enum import Enum class LLMType(Enum): CHAT = "chat" EMBEDDING = "embedding" SPEECH2TEXT = "speech2text" task_type = LLMType.SPEECH2TEXT print(task_type) LLMType.SPEECH2TEXT Example-2: With StrEnum from enum import StrEnum class LLMType(StrEnum): CHAT = "chat" EMBEDDING = "embedding" SPEECH2TEXT
-
GenAI – Query Aware Chunking.
-
GenAI – Adaptive Chunk Sizing
-

GenAI – Galileo’s Chunk Attribution & Utilization Metrics.
-

GenAI – Galileo’s Context Adherence Metric(CAM)
-

GenAI – Data Encryption Algorithms
GenAI – Data Encryption Algorithms Table Of Content: Symmetric Encryption Algorithms. Asymmetric Encryption Algorithms. Hashing Algorithms. Hybrid Encryption. Summary Table. (1) Symmetric Encryption Algorithms. (2) Asymmetric Encryption Algorithms (3) Hashing Algorithms (4) Hybrid Encryption (5) Summary Table (6) Best Python Library For Encryption & Security.
-

GenAI – GenAI Security Breach Scenarios.
GenAI – GenAI Security Breach Scenarios Table Of Content: ChatGPT users reported seeing other users’ chat histories due to a bug. In some cases, payment information was also exposed. A former AWS employee exploited a misconfigured WAF and gained access to over 100 million customer records stored in Amazon S3. Researchers showed they could extract sensitive training data from fine-tuned GPT-style models by crafting adversarial prompts. GitHub Copilot sometimes generated insecure code patterns or replicated licensed code snippets from public repos. Source code for Toyota’s T-Connect app was publicly exposed on GitHub for over 5 years, revealing credentials to the
-

GenAI – Content Moderation Tools
GenAI – Content Moderation Libraries Table Of Content: Content Moderation Libraries.
