北大天网搜索引擎TSE分析及完全注释[5]倒排索引的建立及文件介绍

2023-12-08 01:56:57

不好意思让大家久等了，前一阵一直在忙考试，终于结束了。呵呵！废话不多说了下面我们开始吧！

TSE用的是将抓取回来的网页文档全部装入一个大文档，让后对这一个大文档内的数据整体统一的建索引，其中包含了几个步骤。

view plain copy to clipboard print ?

1. The document index (Doc.idx) keeps information about each document.
It is a fixed width ISAM (Index sequential access mode) index, orderd by docID.
The information stored in each entry includes a pointer into the repository,
a document length, a document checksum.
//Doc.idx 文档编号文档长度 checksum hash码
0 0 bc9ce846d7987c4534f53d423380ba70
1 76760 4f47a3cad91f7d35f4bb6b2a638420e5
2 141624 d019433008538f65329ae8e39b86026c
3 142350 5705b8f58110f9ad61b1321c52605795
//Doc.idx end
The url index (url.idx) is used to convert URLs into docIDs.
//url.idx
5c36868a9c5117eadbda747cbdb0725f 0
3272e136dd90263ee306a835c6c70d77 1
6b8601bb3bb9ab80f868d549b5c5a5f3 2
3f9eba99fa788954b5ff7f35a5db6e1f 3
//url.idx end
It is a list of URL checksums with their corresponding docIDs and is sorted by
checksum. In order to find the docID of a particular URL, the URL's checksum
is computed and a binary search is performed on the checksums file to find its
docID.
./DocIndex
got Doc.idx, Url.idx, DocId2Url.idx //Data文件夹中的Doc.idx DocId2Url.idx和Doc.idx中
//DocId2Url.idx
0 http://*.*.edu.cn/index.aspx
1 http://*.*.edu.cn/showcontent1.jsp?NewsID=118
2 http://*.*.edu.cn/0102.html
3 http://*.*.edu.cn/0103.html
//DocId2Url.idx end
2. sort Url.idx|uniq > Url.idx.sort_uniq //Data文件夹中的Url.idx.sort_uniq
//Url.idx.sort_uniq
//对hash值进行排序
000bfdfd8b2dedd926b58ba00d40986b 1111
000c7e34b653b5135a2361c6818e48dc 1831
0019d12f438eec910a06a606f570fde8 366
0033f7c005ec776f67f496cd8bc4ae0d 2103
3. Segment document to terms, (with finding document according to the url)
./DocSegment Tianwang.raw.2559638448 //Tianwang.raw.2559638448为爬回来的文件，每个页面包含http头
got Tianwang.raw.2559638448.seg
//Tianwang.raw.2559638448 爬取的原始网页文件在文档内部每一个文档之间应该是通过version，和回车做标志位分割的
version: 1.0
url: http://***.105.138.175/Default2.asp?lang=gb
origin: http://***.105.138.175/
date: Fri, 23 May 2008 20:01:36 GMT
ip: 162.105.138.175
length: 38413
HTTP/1.1 200 OK
Server: Microsoft-IIS/5.0
Date: Fri, 23 May 2008 11:17:49 GMT
Connection: keep-alive
Connection: Keep-Alive
Content-Length: 38088
Content-Type: text/html; Charset=gb2312
Expires: Fri, 23 May 2008 11:17:49 GMT
Set-Cookie: ASPSESSIONIDSSTRDCAB=IMEOMBIAIPDFCKPAEDJFHOIH; path=/
Cache-control: private
"-//W3C//DTD HTML 4.01 Transitional//EN"
"http://www.w3.org/TR/html4/loose.dtd">
Apabi数字资源平台
"Content-Type" content="text/html; charset=gb2312">
"ROBOTS" CONTENT="INDEX,NOFOLLOW">
"DESCRIPTION" CONTENT="数字图书馆方正数字图书馆电子图书电子书 ebook e书 Apabi 数字资源平台">
"stylesheet" type="text/css" href="css/common.css">
"text/css">
"vbscript">
...
"javascript">
...
"0" topmargin="0">
//Tianwang.raw.2559638448 end
//Tianwang.raw.2559638448.seg 将每个页面分成一行如下(注意中间没有回车作为分隔)
1
...
...
...
2
...
...
...
//Tianwang.raw.2559638448.seg end
//下是 Tiny search 非必须因素
4. Create forward index (docic-->termid) //建立正向索引
./CrtForwardIdx Tianwang.raw.2559638448.seg > moon.fidx
//Tianwang.raw.2559638448.seg 将每个页面分成一行如下
//分词   DocID
1
三星/  s/  手机/  论坛/  ,/  手机/  铃声/  下载/  ,/  手机/  图片/  下载/  ,/  手机/
2
...
...
...

1.  The document index (Doc.idx) keeps information about each document.It is a fixed width ISAM (Index sequential access mode) index, orderd by docID.The information stored in each entry includes a pointer into the repository,a document length, a document checksum.//Doc.idx  文档编号	文档长度	checksum hash码0	0	bc9ce846d7987c4534f53d423380ba701	76760	4f47a3cad91f7d35f4bb6b2a638420e52	141624	d019433008538f65329ae8e39b86026c3	142350	5705b8f58110f9ad61b1321c52605795//Doc.idx	endThe url index (url.idx) is used to convert URLs into docIDs.//url.idx5c36868a9c5117eadbda747cbdb0725f	03272e136dd90263ee306a835c6c70d77	16b8601bb3bb9ab80f868d549b5c5a5f3	23f9eba99fa788954b5ff7f35a5db6e1f	3//url.idx	endIt is a list of URL checksums with their corresponding docIDs and is sorted bychecksum. In order to find the docID of a particular URL, the URL's checksumis computed and a binary search is performed on the checksums file to find itsdocID../DocIndexgot Doc.idx, Url.idx, DocId2Url.idx	//Data文件夹中的Doc.idx DocId2Url.idx和Doc.idx中//DocId2Url.idx0	http://*.*.edu.cn/index.aspx1	http://*.*.edu.cn/showcontent1.jsp?NewsID=1182	http://*.*.edu.cn/0102.html3	http://*.*.edu.cn/0103.html//DocId2Url.idx	end2.  sort Url.idx|uniq > Url.idx.sort_uniq	//Data文件夹中的Url.idx.sort_uniq//Url.idx.sort_uniq//对hash值进行排序000bfdfd8b2dedd926b58ba00d40986b	1111000c7e34b653b5135a2361c6818e48dc	18310019d12f438eec910a06a606f570fde8	3660033f7c005ec776f67f496cd8bc4ae0d	21033. Segment document to terms, (with finding document according to the url)./DocSegment Tianwang.raw.2559638448		//Tianwang.raw.2559638448为爬回来的文件 ，每个页面包含http头got Tianwang.raw.2559638448.seg		//Tianwang.raw.2559638448	爬取的原始网页文件在文档内部每一个文档之间应该是通过version，和回车做标志位分割的version: 1.0url: http://***.105.138.175/Default2.asp?lang=gborigin: http://***.105.138.175/date: Fri, 23 May 2008 20:01:36 GMTip: 162.105.138.175length: 38413HTTP/1.1 200 OKServer: Microsoft-IIS/5.0Date: Fri, 23 May 2008 11:17:49 GMTConnection: keep-aliveConnection: Keep-AliveContent-Length: 38088Content-Type: text/html; Charset=gb2312Expires: Fri, 23 May 2008 11:17:49 GMTSet-Cookie: ASPSESSIONIDSSTRDCAB=IMEOMBIAIPDFCKPAEDJFHOIH; path=/Cache-control: privateApabi数字资源平台//Tianwang.raw.2559638448	end//Tianwang.raw.2559638448.seg	将每个页面分成一行如下(注意中间没有回车作为分隔)1.........2.........//Tianwang.raw.2559638448.seg	end//下是 Tiny search 非必须因素4. Create forward index (docic-->termid)		//建立正向索引./CrtForwardIdx Tianwang.raw.2559638448.seg > moon.fidx//Tianwang.raw.2559638448.seg 将每个页面分成一行如下
//分词   DocID
1
三星/  s/  手机/  论坛/  ,/  手机/  铃声/  下载/  ,/  手机/  图片/  下载/  ,/  手机/
2
...
...
...

view plain copy to clipboard print ?

//Tianwang.raw.2559638448.seg end
//moon.fidx
//每篇文档号对应文档内分出来的分词 DocID
都会 2391
使 2391
那些 2391
拥有 2391
它 2391
的 2391
人 2391
的 2391
视野 2391
变 2391
窄 2391
在 2180
研究生部 2180
主页 2180
培养 2180
管理 2180
栏目 2180
下载 2180
） 2180
、 2180
关于 2180
做好 2180
年 2180
国家 2180
公派 2180
研究生 2180
项目 2180
//moon.fidx end
5.# set | grep "LANG"
LANG=en; export LANG;
sort moon.fidx > moon.fidx.sort
6. Create inverted index (termid-->docid) //建立倒排索引
./CrtInvertedIdx moon.fidx.sort > sun.iidx
//sun.iidx //文件规模大概减少1/2
花工 236
花海 2103
花卉 1018 1061 1061 1061 1730 1730 1730 1730 1730 1852 949 949
花蕾 447 447
花木 1061
花呢 1430
花期 447 447 447 447 447 525
花钱 174 236
花色 1730 1730
花色品种 1660
花生 450 526
花式 1428 1430 1430 1430
花纹 1430 1430
花序 447 447 447 447 447 450
花絮 136 137
花芽 450 450
//sun.iidx end
TSESearch CGI program for query
Snapshot CGI program for page snapshot

//Tianwang.raw.2559638448.seg end//moon.fidx//每篇文档号对应文档内分出来的	分词	DocID都会	2391使	2391那些	2391拥有	2391它	2391的	2391人	2391的	2391视野	2391变	2391窄	2391在	2180研究生部	2180主页	2180培养	2180管理	2180栏目	2180下载	2180）	2180、	2180关于	2180做好	2180年	2180国家	2180公派	2180研究生	2180项目	2180//moon.fidx	end5.# set | grep "LANG"LANG=en; export LANG;sort moon.fidx > moon.fidx.sort6. Create inverted index (termid-->docid)	//建立倒排索引./CrtInvertedIdx moon.fidx.sort > sun.iidx//sun.iidx	//文件规模大概减少1/2花工	 236花海	 2103花卉	 1018 1061 1061 1061 1730 1730 1730 1730 1730 1852 949 949花蕾	 447 447花木	 1061花呢	 1430花期	 447 447 447 447 447 525花钱	 174 236花色	 1730 1730花色品种	 1660花生	 450 526花式	 1428 1430 1430 1430花纹	 1430 1430花序	 447 447 447 447 447 450花絮	 136 137花芽	 450 450//sun.iidx	endTSESearch	CGI program for querySnapshot	CGI program for page snapshot
author:http://hi.baidu.com/jrckkyyauthor:http://blog.csdn.net/jrckkyy

本文来自互联网用户投稿，文章观点仅代表作者本人，不代表本站立场，不承担相关法律责任。如若转载，请注明出处。 如若内容造成侵权/违法违规/事实不符，请点击【内容举报】进行投诉反馈！

标签：技术

上一篇 > 解决不能打开1433端口
下一篇 > 开源项目（天网千帆）感受by csdn zihui

Duilib中list控件支持ctrl和shif多行选中的实现

[ICML2015]Batch Normalization:Accelerating Deep Network Training by Reducing Internal Covariate Shif

win10系统微软输入法于eclipse ctrl+shif+f冲突间接处理办法

Codeforces Round #259 (Div. 2) B. Little Pony and Sort by Shif

读LDD3，内存映射与DMA--PAGE_SHIF…

VMware虚拟机安装XP【要先分区，再设置BOOT 启动CD，shif+上移】

更换iBus五笔的左与右Shif

sublime ctrl+shif+f 没用解决办法

idea 对 ctrl + z 的撤销是 ctrl + shif + z

计算机最早的设计师应用于,计算机应用基础选择题doc.doc

win10自带截图神器：Win+Shift+S

Python基础之文件目录操作

python简述目录_Python基础之文件目录操作(示例代码)

tp5 如何做数据采集

任务2-7(服务器字体+阿里巴巴矢量库)

html标签（1)：h1~h6,p,br,pre,hr

TI 电量计介绍与芯片选型指南

几款TI电源芯片简介

TI DSP芯片C2000系列读取FLASH数据

德州仪器(Ti)平台嵌入式开发基础

TI三相电机智能栅极驱动芯片特点分类

省选模拟（12.08） T3 圈圈圈圈圈圈圈圈

Hadoop生态圈技术栈（上）

大数据开发基础入门与项目实战（三）Hadoop核心及生态圈技术栈之6.Impala交互式查询

小猿圈之Linux下Mysql 操作命令

大数据Hadoop生态圈常用面试题

大数据开发基础入门与项目实战（三）Hadoop核心及生态圈技术栈之4.Hive DDL、DQL和数据操作

备战Noip2018模拟赛11（B组）T3 Monogatari 物语

【智能优化算法-圆圈搜索算法】基于圆圈搜索算法Circle Search Algorithm求解单目标优化问题附matlab代码

NYOJ 78 圈水池

递归问题跑道汽车绕圈问题 Python实现

Hadoop生态圈（三）：MapReduce

北大天网搜索引擎TSE分析及完全注释[5]倒排索引的建立及文件介绍

相关文章