咦什么

文本匹配论文阅读 A Simple but Tough-to-Beat Baseline for Sentence Embeddings阅读笔记

文本匹配论文阅读

DIIN(1) heap(1) STL(1) 排序(1) 文本分类(2) 文本分类论文阅读(10) 未归档(29) 机器学习(3) 算法面试(1) 面试(1)

/ 注册

A Simple but Tough-to-Beat Baseline for Sentence Embeddings阅读笔记

1350 浏览 0 回复 2019-05-08

咦什么

+关注

文章目录

概述
算法
实验
- 1. Textual Similarity Tasks
- 2. Supervised Tasks

概述

一篇17年的论文, 采用无监督的方法.
主要思想可以概括为两步:

利用词嵌入方法，通过词向量的线性的加权组合对一个句子进行编码
利用奇异向量求出最终的句向量。

算法

实验

1. Textual Similarity Tasks

数据集

all the datasets from SemEval semantic textual similarity (STS) tasks (2012-2015)
the SemEval 2015 Twitter task
the SemEval 2014 Semantic Relatedness task

实验设置
词向量分别采用了无监督的GloVe和弱监督的PSL.
$α$ 固定为 $1 0^{- 3}$ , 词频利用commoncrawl dataset进行统计.

实验结果

2. Supervised Tasks

the SICK similarity task
the SICK entailment task
the Stanford Sentiment Treebank (SST) binary classification task

实验结果

举报

收藏

赞

评论加载中...