'Building a Knowledge-Sharing Chatbot: Turn Expertise Into an AI That Anyone [post] deterministic
How to build a chatbot that captures your knowledge, answers questions
Building a Knowledge-Sharing Chatbot: Turn Expertise Into an AI That Anyone Can Query
We all carry knowledge that others need. Whether you're a seasoned manager with institutional history, a technician with troubleshooting tricks learned over decades, or a founder with lessons from a hundred decisions—the problem is the same: **your knowledge is trapped in your head, and it scales poorly.**
You could write documentation, but documentation is static. It doesn't answer follow-up questions. It doesn't adapt to what someone actually needs in the moment. And most people don't read it anyway—they ask you.
What if you could clone the part of yourself that answers questions? Not a generic AI, but one trained on *your* knowledge, *your* processes, *your* edge cases?
That's what I built: a system that captures expertise through guided interviews, transforms it into structured documentation, and delivers a chatbot that anyone can query. The key constraint? Everything runs locally—no cloud APIs, no monthly fees, complete privacy.
The Real Problem: Knowledge Bottlenecks
Every expert becomes a bottleneck. Here's how it manifests:
**For individuals:** - You answer the same questions repeatedly - Your time gets consumed by knowledge transfer instead of high-value work - When you're unavailable, decisions wait or go wrong
**For organizations:** - Key person dependency creates risk - Onboarding takes months instead of weeks - Hard-won lessons get lost when people leave
**For communities:** - Expertise remains siloed with a few individuals - Newcomers struggle to get up to speed - Knowledge fragments across chat logs, emails, and documents
Traditional solutions don't work well. Wikis go stale. Training videos are passive. Documentation requires people to know what to look for. What people actually want is **conversation**—the ability to ask questions and get answers tailored to their context.
The Solution: A Knowledge-Capture-to-Chatbot Pipeline
The system I built follows a simple but powerful pipeline:
Expert Interview → LLM Structuring → Vector Embeddings → Queryable Chatbot
↓
Unknown Questions → Expert Review
↓
New Knowledge Integrated ←
This creates a **learning loop**: the chatbot answers what it knows, flags what it doesn't, and gets smarter over time.
Why This Approach Works
1. **Interview-based capture**: Experts don't have to write documentation—they just answer questions they already know 2. **LLM structuring**: Raw responses get transformed into organized, readable documentation automatically 3. **Semantic search**: Users ask questions naturally, not with exact keywords 4. **Dynamic learning**: The system improves without manual updates
The Architecture in Practice
Let me show you how each component works, using real code from the implementation.
Phase 1: Capturing Expert Knowledge
The first challenge is getting knowledge out of people's heads. Most experts are too busy to write comprehensive documentation, but they'll answer focused questions.
The interview module uses a structured approach:
```python # Structured interview questions for staff INTERVIEW_QUESTIONS = [ "What are the main tasks you do daily?", "What mistakes do new hires often make?", "Which documents or forms are essential for your role?", "Are there any edge cases you frequently encounter?", "What advice would you give to someone just starting in this role?" ]
Keywords that indicate potential edge cases EDGE_CASE_KEYWORDS = [ "sometimes", "rarely", "depends", "if", "occasionally", "usually", "typically", "in rare cases", "edge case" ]
def detect_edge_cases(response_text: str) -> list: """ Detect potential edge cases based on keywords in the response. """ edge_cases = [] sentences = response_text.split('. ') for sentence in sentences: sentence_lower = sentence.lower() for keyword in EDGE_CASE_KEYWORDS: if keyword in sentence_lower: edge_cases.append(sentence.strip()) break return edge_cases ```
The interview process is deliberately conversational:
python
def run_staff_interview(staff_name: str, role: str) -> StaffResponse:
"""
Run an interactive staff interview via console input.
"""
print(f"\n{'='*50}")
print(f"Expert Interview: {staff_name} - {role}")
print(f"{'='*50}\n")
responses = []
all_edge_cases = []
for question in INTERVIEW_QUESTIONS:
print(f"Question: {question}")
answer = input("Answer: ").strip()
if not answer:
print(" (Skipped - no answer provided)")
continue
responses.append(f"Q: {question}\nA: {answer}")
# Check for edge cases in the answer
detected = detect_edge_cases(answer)
all_edge_cases.extend(detected)
print(f" ✓ Recorded ({len(detected)} potential edge cases detected)\n")
# Combine all responses into single text
response_text = "\n\n".join(responses)
# Create and save staff response
session = get_session()
staff_response = StaffResponse(
staff_name=staff_name,
role=role,
response_text=response_text,
edge_cases=all_edge_cases
)
session.add(staff_response)
session.commit()
print(f"\n{'='*50}")
print(f"Interview complete! {len(all_edge_cases)} edge cases detected.")
print(f"Responses saved for {staff_name} ({role})")
print(f"{'='*50}\n")
session.close()
return staff_response
**Key insight**: The questions are designed to surface not just what to do, but *what goes wrong*. Questions about mistakes and edge cases capture the tacit knowledge that never makes it into formal documentation.
Phase 2: Structuring Knowledge with LLMs
Raw interview responses are valuable but unstructured. The LLM transforms them into coherent documentation:
```python from db import StaffResponse, SOPDraft, get_session from llm import generate_text, SOP_GENERATION_PROMPT
def generate_sop(staff_response_id: int) -> SOPDraft: """ Generate an SOP draft from staff responses. """ session = get_session() # Get staff response sr = session.query(StaffResponse).filter_by(id=staff_response_id).first() if not sr: print(f"Staff response with ID {staff_response_id} not found") session.close() return None print(f"Generating SOP for {sr.staff_name} ({sr.role})...") # Build prompt for SOP generation prompt = f"""Create a detailed Standard Operating Procedure (SOP) based on the following staff responses:
{sr.response_text}
Please organize this into a clear, professional SOP with: 1. Role Overview 2. Daily Tasks (step-by-step) 3. Common Mistakes to Avoid 4. Essential Forms/Documents 5. Edge Cases / Special Circumstances """ # Generate SOP using Ollama sop_text = generate_text( prompt=prompt, system_prompt=SOP_GENERATION_PROMPT, temperature=0.3, max_tokens=1500 ) # Create SOP draft sop = SOPDraft( role=sr.role, sop_text=sop_text, metadata={ "staff_name": sr.staff_name, "staff_response_id": sr.id } ) session.add(sop) session.commit() print(f"SOP generated and saved for role: {sr.role}") session.close() return sop ```
The system prompt guides the LLM to create well-structured output:
```python SOP_GENERATION_PROMPT = """You are an expert process engineer and technical writer. Your task is to create clear, structured Standard Operating Procedures (SOPs) from staff responses. Create well-organized documents with: - Clear step-by-step instructions - Checklists where ap