Skip to main content

Overview

The switchAILocal embedding SDK provides ONNX-based local embedding generation using the MiniLM model. This enables advanced features like semantic caching, intelligent routing, and skill matching without requiring external API calls.

Key Features

  • Local Processing: All embeddings computed on-device using ONNX Runtime
  • 384-Dimensional Vectors: Standard MiniLM-L6-v2 model output
  • Fast Inference: Optimized for real-time semantic matching (<20ms)
  • No External Dependencies: Fully offline after model download
  • Thread-Safe: Concurrent embedding generation support

Architecture

When to Use Embeddings

Semantic Tier (Phase 2)

Match user queries to intents using embedding similarity instead of LLM classification:
Benefits:
  • 10-20x faster than LLM classification
  • Deterministic results
  • No API costs

Semantic Caching

Cache responses based on semantic similarity:
Use Cases:
  • Deduplicating similar queries
  • Reducing API costs
  • Faster response times for common questions

Skill Matching

Match queries to domain-specific skills:
Example Skills:
  • Language experts (Go, Python, TypeScript)
  • Infrastructure (Docker, Kubernetes)
  • Security, Testing, Debugging

Model Details

all-MiniLM-L6-v2

Download the Model

This downloads:
  • model.onnx - The ONNX model file
  • vocab.txt - The tokenizer vocabulary
Files are stored in ~/.switchailocal/models/.

Quick Start

1

Download Model

2

Enable in Config

3

Start Server

4

Verify

Check logs for:

Configuration Options

Performance Characteristics

Latency

Memory Usage

  • Model Loading: ~50 MB
  • Per Request: ~1-2 MB (temporary)
  • Cached Embeddings: 384 floats × 4 bytes = 1.5 KB per vector

Accuracy

  • Semantic Similarity: 0.0 (unrelated) to 1.0 (identical)
  • Typical Intent Match: >0.85 for correct matches
  • Typical Skill Match: >0.80 for relevant skills

Comparison with Alternatives

Next Steps

Usage Guide

Learn how to use the embedding SDK

Custom Providers

Integrate custom embedding models

Semantic Tier

Configure semantic intent matching

Semantic Cache

Enable semantic caching