html2article-golang基于文本密度的 html2article 實(shí)現(xiàn)
html2article — 基于文本密度的html2article實(shí)現(xiàn)[golang]
Install
go get -u -v github.com/sundy-li/html2article
Performance
avg 3.2ms per article, accuracy >= 98% (對(duì)比其他開(kāi)源實(shí)現(xiàn),可能是目前最快的html2article實(shí)現(xiàn),我們測(cè)試的數(shù)據(jù)集約3kw來(lái)自于微信公眾號(hào),各大類中文科技媒體歷史文章,目前能達(dá)到98%以上準(zhǔn)確率)
Examples
參考examples from_url.go
package main
import (
"github.com/sundy-li/html2article"
)
func main() {
article, err := html2article.FromUrl("https://www.leiphone.com/news/201602/DsiQtR6c1jCu7iwA.html")
if err != nil {
panic(err)
}
println("article title is =>", article.Title)
println("article publishtime is =>", article.Publishtime)
println("article content is =>", article.Content)
}
Algorithm
評(píng)論
圖片
表情
