qualora / data /documentation.md
Pratham Amritkar
final
d833ce9
|
Raw
History Blame Contribute Delete
6.03 kB

Qualora Documentation

Platform Version: 1.0 (Early Access)

Architecture: Multi-Tier LLM Cascade & RAG Retrieval

Welcome to the Qualora Enterprise Quality Auditing Platform documentation. This technical guide outlines the core functional domains, inference architecture, and security protocols governing the Qualora ecosystem.


1. Executive Overview

Qualora is a GenAI-native quality auditing copilot engineered to automate the evaluation of customer support interactions at scale. By synthesizing advanced Retrieval-Augmented Generation (RAG) with a fault-tolerant LLM cascade, Qualora provides objective, deterministic scoring across critical dimensions including agent empathy, regulatory compliance, and operational efficiency.

2. Core Audit Engine

The Audit Engine acts as the primary ingestion and inference layer, natively supporting deterministic pipelines for both voice and text modalities.

2.1 Voice Audits

  • Ingestion: The platform accepts raw audio payloads (MP3, WAV, M4A, OGG) up to 50MB, or direct browser-based streams via the integrated MediaRecorder API.
  • Transcription Cascade (Tier 1): HuggingFace (HF) Space nodes executing localized WhisperX models, yielding high-fidelity transcriptions with speaker diarization and prosody metadata.
  • Transcription Cascade (Tier 2): High-availability fallback routing through API chains, leveraging ElevenLabs Scribe, Deepgram Nova-2, or Groq Whisper-large-v3.
  • Evaluation: Normalized transcripts are routed asynchronously to the LLM cascade for multi-dimensional scoring against the tenant's knowledge base.

2.2 Text and Chat Audits

  • Ingestion: Supports unstructured raw text insertion or structured document parsing (PDF, CSV, JSON, TXT).
  • Sanitization: Payloads undergo pre-processing and sanitization before being audited against organization-specific compliance schemas.

2.3 LLM Inference Cascade

To ensure strict zero-downtime availability, Qualora routes all evaluation queries through a resilient fallback hierarchy:

  1. Primary Tier: Google Gemini 2.0 Flash (Optimized for high-context, low-latency reasoning).
  2. Secondary Tier: Llama 3.3 via OpenRouter (High-performance secondary logic).
  3. Tertiary Tier: Llama 3.3 via Groq (Ultra-fast inference recovery layer).
  4. Resort Tier: HuggingFace Inference Endpoints (Emergency failover).

3. RAG Knowledge Base (KB)

The Knowledge Base is the organizational "Source of Truth," enabling supervisors to map arbitrary policy documents directly into the AI Judge's scoring logic.

  • Atomic Vectorization (Primary Store): MongoDB Atlas Vector Search for secure, cloud-native semantic retrieval.
  • Atomic Vectorization (Caching Store): ChromaDB for localized, high-throughput persistent caching.
  • Retrieval Pipeline: Upon audit execution, the RAG engine retrieves the highest-relevance policy clauses, appending them to the system prompt to constrain and calibrate the model's evaluation.

4. Human-In-The-Loop (HITL) Workflow

Qualora strictly enforces a human-centric auditing philosophy, ensuring automated outputs do not execute autonomously in sensitive HR environments.

  • Supervisor Overrides: Administrators and Supervisors possess unilateral rights to review AI-generated reports, marking them as Approved, Flagged, or Rejected.
  • State Reversibility: Destructive or status-modifying commits carry a 10-second grace window, mitigating accidental state changes.
  • Contextual Annotations: Supervisors may append manual coaching notes directly to the immutable audit trail to supplement the AI's behavioral nudges.

5. Agent Telemetry and Performance

Agent evaluation is quantified through continuous, multi-dimensional tracking:

  • Quality Matrix: Interactions are scored against five core dimensions: Language Proficiency, Cognitive Empathy, Efficiency, Bias Reduction, and Active Listening.
  • F1 Trend Analysis: The dashboard visualizes a rolling time-series graph of an agent's recent 50 audits, highlighting standard deviations in performance.
  • Risk Stratification: The platform algorithms isolate agents demonstrating recurring critical (Red) or cautionary (Amber) compliance flags, optimizing targeted coaching prioritization.

6. Real-Time Compliance Alerts

The system implements a continuous monitoring daemon to trigger localized alerts upon detecting:

  1. Critical Breach: Immediate detection of Red-Risk compliance violations during an interaction.
  2. Performance Degradation: An agent's rolling F1 score dropping below the 0.6 threshold.
  3. Systematic Failure: Compounding failures across generalized matrix criteria.
  • Integration: Enterprise deployment allows dynamic dispatching of these alerts via Webhooks to Slack or Discord channels.

7. Security and Tenant Isolation

Security represents a foundational pillar of the Qualora architecture:

  • Authentication Vectors: Sessions are exclusively governed by HTTP-Only JWT cookies, nullifying exposure to Cross-Site Scripting (XSS) payload extraction.
  • CSRF Mitigation: All state-mutating requests (POST/PUT/DELETE) strictly require header-injected X-CSRF-Token validation.
  • Data Sovereignty: Robust database-level isolation guarantees that tenants can solely query and mutate data tethered strictly to their internal Organizational ID.

8. Troubleshooting Logistics

  • System Diagnostics: Utilize the "System Status" indicator in the application footer for real-time ping tracking of API targets, database connectivity, and average response latencies.
  • Inference Fast Track: Should standard inference pipelines experience degradation, users may manually invoke the "Fast Track" directive, forcefully shifting the request to the highest-speed node available.
  • Session Lifecycles: Strict TTL (Time-To-Live) constraints govern active sessions; expired tokens will execute a secure, automatic redirect to the authentication gateway.