Semantic relevance, computed locally
Your question, every caption, and every transcript are embedded as vectors by a MiniLM sentence-transformer and compared by cosine similarity, so “best budget espresso setup” matches videos that never use those exact words. If the model can't load, ranking falls back to lexical BM25.