Qwen3 Embedding
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models Embedding 使用[EOS]对应的hidden state作为embedding 使用改进的InfoNCE Loss进行训练 分子S = query 和 正例的余弦相似度 分母Z = S + query和难负样本的余弦相似度+batch内正例与其它query的余弦相似度+batch内query与其它文档的余弦相似度 温度系数 防止假负例:相似度>S+0.1 大规模合成数据的弱监督训练 任务类型:检索(Retrieval)、双语对齐(Bitext mining)、语意相似度(STS, semantic textual similarity)、分类(classification) 查询重写:角色库top5的角色放入提示词中,提示词涵盖查询类型、查询长度、难度、语言等 1.5亿对进行预训练,1200万高质量进行SFT 重写Prompt Given a **Character**, **Passage**, and **Requirement**, generate a query from the **Character**’s perspective that satisfies the **Requirement** and can be used to retrieve the **Passage**. Please return the result in JSON format. Here is an example: <example> Now, generate the **output** based on the **Character**, **Passage** and language, the **Character** and **Requirement** will be in English. **Requirement** from user, the **Passage** will be in...