<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://blackwiki.org/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=90.91.11.127&amp;*</id>
	<title>blackwiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://blackwiki.org/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=90.91.11.127&amp;*"/>
	<link rel="alternate" type="text/html" href="https://blackwiki.org/index.php?title=Special:Contributions/90.91.11.127"/>
	<updated>2026-09-30T12:03:54Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.34.2</generator>
	<entry>
		<id>https://blackwiki.org/index.php?title=Fuzzy_retrieval&amp;diff=211518</id>
		<title>Fuzzy retrieval</title>
		<link rel="alternate" type="text/html" href="https://blackwiki.org/index.php?title=Fuzzy_retrieval&amp;diff=211518"/>
		<updated>2020-09-18T14:33:25Z</updated>

		<summary type="html">&lt;p&gt;90.91.11.127: test 1&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;'''Fuzzy retrieval''' techniques are based on the [[Extended Boolean model]] and the [[Fuzzy set]] theory. There are two classical fuzzy retrieval models: Mixed Min and Max (MMM) and the Paice model. Both models do not provide a way of evaluating query weights, however this is considered by the [[Extended Boolean model|P-norms]] algorithm.&lt;br /&gt;
&lt;br /&gt;
==Mixed Min and Max model (MMM)==&lt;br /&gt;
&lt;br /&gt;
In fuzzy-set theory, an element has a varying degree of membership, say ''d&amp;lt;sub&amp;gt;A&amp;lt;/sub&amp;gt;'', to a given set ''A'' instead of the traditional membership choice (is an element/is not an element).&amp;lt;br /&amp;gt;&lt;br /&gt;
In MMM&amp;lt;ref&amp;gt;{{citation | last1=Fox | first1=E. A. | author2=S. Sharat | year=1986 | title=A Comparison of Two Methods for Soft Boolean Interpretation in Information Retrieval | publisher=Technical Report TR-86-1, Virginia Tech, Department of Computer Science}}&amp;lt;/ref&amp;gt; each index term has a fuzzy set associated with it. A document's weight with respect to an index term ''A'' is considered to be the degree of membership of the document in the fuzzy set associated with ''A''. The degree of membership for union and intersection are defined as follows in Fuzzy set theory:&amp;lt;br/&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;d_{A\cap B}= min(d_A, d_B)&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;d_{A\cup B}= max(d_A,d_B)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
According to this, documents that should be retrieved for a query of the form ''A or B'', should be in the fuzzy set associated with the union of the two sets ''A'' and ''B''. Similarly, the documents that should be retrieved for a query of the form ''A and B'', should be in the fuzzy set associated with the intersection of the two sets. Hence, it is possible to define the similarity of a document to the ''or'' query to be ''max(d&amp;lt;sub&amp;gt;A&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;B&amp;lt;/sub&amp;gt;)'' and the similarity of the document to the ''and'' query to be ''min(d&amp;lt;sub&amp;gt;A&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;B&amp;lt;/sub&amp;gt;)''. The MMM model tries to soften the Boolean operators by considering the query-document similarity to be a linear combination of the ''min'' and ''max'' document weights.&lt;br /&gt;
&lt;br /&gt;
Given a document ''D'' with index-term weights ''d&amp;lt;sub&amp;gt;A1&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;A2&amp;lt;/sub&amp;gt;, ..., d&amp;lt;sub&amp;gt;An&amp;lt;/sub&amp;gt;'' for terms ''A&amp;lt;sub&amp;gt;1&amp;lt;/sub&amp;gt;, A&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt;, ..., A&amp;lt;sub&amp;gt;n&amp;lt;/sub&amp;gt;'', and the queries:&lt;br /&gt;
&lt;br /&gt;
''Q&amp;lt;sub&amp;gt;or&amp;lt;/sub&amp;gt; = (A&amp;lt;sub&amp;gt;1&amp;lt;/sub&amp;gt; or A&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt; or ... or A&amp;lt;sub&amp;gt;n&amp;lt;/sub&amp;gt;)''&amp;lt;br /&amp;gt;&lt;br /&gt;
''Q&amp;lt;sub&amp;gt;and&amp;lt;/sub&amp;gt; = (A&amp;lt;sub&amp;gt;1&amp;lt;/sub&amp;gt; and A&amp;lt;sub&amp;gt;2&amp;lt;/sub&amp;gt; and ... and A&amp;lt;sub&amp;gt;n&amp;lt;/sub&amp;gt;)''&lt;br /&gt;
&lt;br /&gt;
the query-document similarity in the MMM model is computed as follows:&lt;br /&gt;
&lt;br /&gt;
''SlM(Q&amp;lt;sub&amp;gt;or&amp;lt;/sub&amp;gt;, D) = C&amp;lt;sub&amp;gt;or1&amp;lt;/sub&amp;gt; * max(d&amp;lt;sub&amp;gt;A1&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;A2&amp;lt;/sub&amp;gt;, ..., d&amp;lt;sub&amp;gt;An&amp;lt;/sub&amp;gt;) + C&amp;lt;sub&amp;gt;or2&amp;lt;/sub&amp;gt; * min(d&amp;lt;sub&amp;gt;A1&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;A2&amp;lt;/sub&amp;gt;, ..., d&amp;lt;sub&amp;gt;An&amp;lt;/sub&amp;gt;)''&amp;lt;br /&amp;gt;&lt;br /&gt;
''SlM(Q&amp;lt;sub&amp;gt;and&amp;lt;/sub&amp;gt;, D) = C&amp;lt;sub&amp;gt;and1&amp;lt;/sub&amp;gt; * min(d&amp;lt;sub&amp;gt;A1&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;A2&amp;lt;/sub&amp;gt;, ..., d&amp;lt;sub&amp;gt;An&amp;lt;/sub&amp;gt;) + C&amp;lt;sub&amp;gt;and2&amp;lt;/sub&amp;gt; * max(d&amp;lt;sub&amp;gt;A1&amp;lt;/sub&amp;gt;, d&amp;lt;sub&amp;gt;A2&amp;lt;/sub&amp;gt; ..., d&amp;lt;sub&amp;gt;An&amp;lt;/sub&amp;gt;)''&lt;br /&gt;
&lt;br /&gt;
where ''C&amp;lt;sub&amp;gt;or1&amp;lt;/sub&amp;gt;, C&amp;lt;sub&amp;gt;or2&amp;lt;/sub&amp;gt;'' are &amp;quot;softness&amp;quot; coefficients for the ''or'' operator, and ''C&amp;lt;sub&amp;gt;and1&amp;lt;/sub&amp;gt;, C&amp;lt;sub&amp;gt;and2&amp;lt;/sub&amp;gt;'' are softness coefficients for the ''and'' operator. Since we would like to give the maximum of the document weights more importance while considering an ''or'' query and the minimum more importance while considering an ''and'' query, generally we have ''C&amp;lt;sub&amp;gt;or1&amp;lt;/sub&amp;gt; &amp;gt; C&amp;lt;sub&amp;gt;or2&amp;lt;/sub&amp;gt; and C&amp;lt;sub&amp;gt;and1&amp;lt;/sub&amp;gt; &amp;gt; C&amp;lt;sub&amp;gt;and2&amp;lt;/sub&amp;gt;''. For simplicity it is generally assumed that ''C&amp;lt;sub&amp;gt;or1&amp;lt;/sub&amp;gt; = 1 - C&amp;lt;sub&amp;gt;or2&amp;lt;/sub&amp;gt;'' and ''C&amp;lt;sub&amp;gt;and1&amp;lt;/sub&amp;gt; = 1 - C&amp;lt;sub&amp;gt;and2&amp;lt;/sub&amp;gt;''.&lt;br /&gt;
&lt;br /&gt;
Lee and Fox&amp;lt;ref name=&amp;quot;leefox&amp;quot;&amp;gt;{{citation | last1=Lee | first1=W. C. | author2=E. A. Fox | year=1988 | title=Experimental Comparison of Schemes for Interpreting Boolean Queries}}&amp;lt;/ref&amp;gt; experiments indicate that the best performance usually occurs with ''C&amp;lt;sub&amp;gt;and1&amp;lt;/sub&amp;gt;'' in the range [0.5, 0.8] and with ''C&amp;lt;sub&amp;gt;or1&amp;lt;/sub&amp;gt;'' &amp;gt; 0.2. In general, the computational cost of MMM is low, and retrieval effectiveness is much better than with the [[Standard Boolean model]].&lt;br /&gt;
&lt;br /&gt;
==Paice model==&lt;br /&gt;
&lt;br /&gt;
The [[Christopher D. Paice|Paice]] model&amp;lt;ref&amp;gt;{{citation | last=Paice | first=C. D. | year=1984 | title=Soft Evaluation of Boolean Search Queries in Information Retrieval Systems | publisher=Information Technology, Res. Dev. Applications, 3(1), 33-42 }}&amp;lt;/ref&amp;gt; is a general extension to the MMM model. In comparison to the MMM model that considers only the minimum and maximum weights for the index terms, the Paice model incorporates all of the term weights when calculating the similarity:&lt;br /&gt;
&lt;br /&gt;
:&amp;lt;math&amp;gt;S(D,Q) = \sum_{i=1}^n\frac{r^{i-1}*w_{di}}{\sum_{j=1}^n r^{j-1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
where ''r'' is a constant coefficient and ''w&amp;lt;sub&amp;gt;di&amp;lt;/sub&amp;gt;'' is arranged in ascending order for ''and'' queries and descending order for ''or'' queries. When n = 2 the Paice model shows the same behavior as the MMM model.&lt;br /&gt;
&lt;br /&gt;
The experiments of Lee and Fox&amp;lt;ref name=&amp;quot;leefox&amp;quot;/&amp;gt; have shown that setting the ''r'' to 1.0 for ''and'' queries and 0.7 for ''or'' queries gives good retrieval effectiveness. The computational cost for this model is higher than that for the MMM model. This is because the MMM model only requires the determination of ''min'' or ''max'' of a set of term weights each time an ''and'' or ''or'' clause is considered, which can be done in ''O(n)''. The Paice model requires the term weights to be sorted in ascending or descending order, depending on whether an ''and'' clause or an ''or'' clause is being considered. This requires at least an ''0(n log n)'' sorting algorithm. A good deal of floating point calculation is needed too.&lt;br /&gt;
&lt;br /&gt;
==Improvements over the Standard Boolean model==&lt;br /&gt;
Lee and Fox&amp;lt;ref name=&amp;quot;leefox&amp;quot;/&amp;gt; compared the Standard Boolean model with MMM and Paice models with three test collections, CISI, CACM and INSPEC. These are the reported results for average mean precision improvement:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
!&lt;br /&gt;
! CISI&lt;br /&gt;
! CACM&lt;br /&gt;
! INSPEC&lt;br /&gt;
|-&lt;br /&gt;
! MMM&lt;br /&gt;
| 68%&lt;br /&gt;
| 109%&lt;br /&gt;
| 195%&lt;br /&gt;
|-&lt;br /&gt;
! Paice&lt;br /&gt;
| 77%&lt;br /&gt;
| 104%&lt;br /&gt;
| 206%&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
These are very good improvements over the Standard model. MMM is very close to Paice and P-norm results which indicates that it can be a very good technique, and is the most efficient of the three.&lt;br /&gt;
&lt;br /&gt;
==Recent work==&lt;br /&gt;
&lt;br /&gt;
Recently '''Kang ''et al.'''.&amp;lt;ref&amp;gt;{{citation | chapter=Fuzzy Information Retrieval Indexed by Concept Identification | last1=Kang | first1=Bo-Yeong | title=Text, Speech and Dialogue | volume=3658 | pages=179–186 | author2=Dae-Won Kim |author3=Hae-Jung Kim | publisher=Springer Berlin / Heidelberg | year=2005| doi=10.1007/11551874_23 | series=Lecture Notes in Computer Science | isbn=978-3-540-28789-6 }}&amp;lt;/ref&amp;gt; have devised a fuzzy retrieval system indexed by concept identification.&lt;br /&gt;
&lt;br /&gt;
If we look at documents on a pure [[Tf-idf]] approach, even eliminating stop words, there will be words more relevant to the topic of the document than others and they will have the same weight because they have the same term frequency. If we take into account the user intent on a query we can better weight the terms of a document. Each term can be identified as a concept in a certain lexical chain that translates the importance of that concept for that document.&amp;lt;br /&amp;gt;&lt;br /&gt;
They report improvements over Paice and P-norm on the average precision and recall for the Top-5 retrieved documents.&lt;br /&gt;
&lt;br /&gt;
Zadrozny&amp;lt;ref&amp;gt;{{citation | title=Fuzzy information retrieval model revisited | journal=Fuzzy Sets and Systems | volume=160 | issue=15 | pages=2173–2191 | doi=10.1016/j.fss.2009.02.012 | first1=Sławomir | last1=Zadrozny | last2=Nowacka | first2=Katarzyna | year=2009 | publisher=Elsevier North-Holland, Inc.}}&amp;lt;/ref&amp;gt; revisited the fuzzy information retrieval model. He further extends the fuzzy extended Boolean model by:&lt;br /&gt;
* assuming linguistic terms as importance weights of keywords also in documents&lt;br /&gt;
* taking into account the uncertainty concerning the representation of documents and queries&lt;br /&gt;
* interpreting the linguistic terms in the representation of documents and queries as well as their matching in terms of the Zadeh's fuzzy logic (calculus of linguistic statements)&lt;br /&gt;
* addressing some pragmatic aspects of the proposed model, notably the techniques of indexing documents and queries&lt;br /&gt;
&lt;br /&gt;
The proposed model makes it possible to grasp both imprecision and uncertainty concerning the textual information representation and retrieval.&lt;br /&gt;
&lt;br /&gt;
==See also==&lt;br /&gt;
*[[Information retrieval]]&lt;br /&gt;
&lt;br /&gt;
==Further reading==&lt;br /&gt;
* {{citation | title=Information Retrieval: Algorithms and Data structures; Extended Boolean model | last1=Fox | first1=E. | author2=S. Betrabet | author3=M. Koushik | author4=W. Lee | year=1992 | publisher=Prentice-Hall, Inc. | url=https://www.scribd.com/doc/13742235/Information-Retrieval-Data-Structures-Algorithms-William-B-Frakes}}&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
{{DEFAULTSORT:Fuzzy Retrieval}}&lt;br /&gt;
[[Category:Information retrieval techniques]]&lt;/div&gt;</summary>
		<author><name>90.91.11.127</name></author>
		
	</entry>
</feed>