<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://tcs.nju.edu.cn/wiki/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Etone</id>
	<title>TCS Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://tcs.nju.edu.cn/wiki/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Etone"/>
	<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Special:Contributions/Etone"/>
	<updated>2026-09-18T16:47:56Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.46.0</generator>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13959</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13959"/>
		<updated>2026-09-16T15:33:28Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([http://tcs.nju.edu.cn/slides/aa2026/Hashing.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-zh.pdf AI生成讲义]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
# [[高级算法 (Fall 2026)/Concentration of measure|Concentration of measure]] ([http://tcs.nju.edu.cn/slides/aa2026/Concentration.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-concentration-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-concentration-zh.pdf AI生成讲义]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Conditional expectations|Conditional expectations]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13958</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13958"/>
		<updated>2026-09-16T15:32:30Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([http://tcs.nju.edu.cn/slides/aa2026/Hashing.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-zh.pdf AI生成讲义]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
# [[高级算法 (Fall 2026)/Concentration of measure|Concentration of measure]] ([http://tcs.nju.edu.cn/slides/aa2026/Concentration.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-concentration-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-concentration.pdf-zh.pdf AI生成讲义]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Conditional expectations|Conditional expectations]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Conditional_expectations&amp;diff=13946</id>
		<title>高级算法 (Fall 2026)/Conditional expectations</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Conditional_expectations&amp;diff=13946"/>
		<updated>2026-09-15T12:00:16Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;= Conditional Expectations = The &amp;#039;&amp;#039;&amp;#039;conditional expectation&amp;#039;&amp;#039;&amp;#039; of a random variable &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; with respect to an event &amp;lt;math&amp;gt;\mathcal{E}&amp;lt;/math&amp;gt; is defined by :&amp;lt;math&amp;gt; \mathbf{E}[Y\mid \mathcal{E}]=\sum_{y}y\Pr[Y=y\mid\mathcal{E}]. &amp;lt;/math&amp;gt; In particular, if the event &amp;lt;math&amp;gt;\mathcal{E}&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;X=a&amp;lt;/math&amp;gt;, the conditional expectation :&amp;lt;math&amp;gt; \mathbf{E}[Y\mid X=a] &amp;lt;/math&amp;gt; defines a function :&amp;lt;math&amp;gt; f(a)=\mathbf{E}[Y\mid X=a]. &amp;lt;/math&amp;gt; Thus, &amp;lt;math&amp;gt;\mathbf{E}[Y\mid...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Conditional Expectations =&lt;br /&gt;
The &#039;&#039;&#039;conditional expectation&#039;&#039;&#039; of a random variable &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; with respect to an event &amp;lt;math&amp;gt;\mathcal{E}&amp;lt;/math&amp;gt; is defined by&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y\mid \mathcal{E}]=\sum_{y}y\Pr[Y=y\mid\mathcal{E}].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
In particular, if the event &amp;lt;math&amp;gt;\mathcal{E}&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;X=a&amp;lt;/math&amp;gt;, the conditional expectation&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y\mid X=a]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
defines a function&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
f(a)=\mathbf{E}[Y\mid X=a].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &amp;lt;math&amp;gt;\mathbf{E}[Y\mid X]&amp;lt;/math&amp;gt; can be regarded as a random variable &amp;lt;math&amp;gt;f(X)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
;Example&lt;br /&gt;
:Suppose that we uniformly sample a human from all human beings. Let &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; be his/her height, and let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be the country where he/she is from. For any country &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathbf{E}[Y\mid X=a]&amp;lt;/math&amp;gt; gives the average height of that country. And &amp;lt;math&amp;gt;\mathbf{E}[Y\mid X]&amp;lt;/math&amp;gt; is the random variable which can be defined in either ways:&lt;br /&gt;
:* We choose a human uniformly at random from all human beings, and &amp;lt;math&amp;gt;\mathbf{E}[Y\mid X]&amp;lt;/math&amp;gt; is the average height of the country where he/she comes from.&lt;br /&gt;
:* We choose a country at random with a probability &#039;&#039;proportional to its population&#039;&#039;, and &amp;lt;math&amp;gt;\mathbf{E}[Y\mid X]&amp;lt;/math&amp;gt; is the average height of the chosen country.&lt;br /&gt;
&lt;br /&gt;
The following proposition states some fundamental facts about conditional expectation.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Proposition (fundamental facts about conditional expectation)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;X,Y&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Z&amp;lt;/math&amp;gt; be arbitrary random variables. Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; be arbitrary functions. Then&lt;br /&gt;
:# &amp;lt;math&amp;gt;\mathbf{E}[X]=\mathbf{E}[\mathbf{E}[X\mid Y]]&amp;lt;/math&amp;gt;.&lt;br /&gt;
:# &amp;lt;math&amp;gt;\mathbf{E}[X\mid Z]=\mathbf{E}[\mathbf{E}[X\mid Y,Z]\mid Z]&amp;lt;/math&amp;gt;.&lt;br /&gt;
:# &amp;lt;math&amp;gt;\mathbf{E}[\mathbf{E}[f(X)g(X,Y)\mid X]]=\mathbf{E}[f(X)\cdot \mathbf{E}[g(X,Y)\mid X]]&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The proposition can be formally verified by computing these expectations. Although these equations look formal, the intuitive interpretations to them are very clear.&lt;br /&gt;
&lt;br /&gt;
The first equation:&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X]=\mathbf{E}[\mathbf{E}[X\mid Y]]&amp;lt;/math&amp;gt;&lt;br /&gt;
says that there are two ways to compute an average. Suppose again that &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is the height of a uniform random human and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is the country where he/she is from. There are two ways to compute the average human height: one is to directly average over the heights of all humans; the other is that first compute the average height for each country, and then average over these heights weighted by the populations of the countries.&lt;br /&gt;
&lt;br /&gt;
The second equation:&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X\mid Z]=\mathbf{E}[\mathbf{E}[X\mid Y,Z]\mid Z]&amp;lt;/math&amp;gt;&lt;br /&gt;
is the same as the first one, restricted to a particular subspace. As the previous example, inaddition to the height &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and the country &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;Z&amp;lt;/math&amp;gt; be the gender of the individual. Thus, &amp;lt;math&amp;gt;\mathbf{E}[X\mid Z]&amp;lt;/math&amp;gt; is the average height of a human being of a given sex. Again, this can be computed either directly or on a country-by-country basis.&lt;br /&gt;
&lt;br /&gt;
The third equation:&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[\mathbf{E}[f(X)g(X,Y)\mid X]]=\mathbf{E}[f(X)\cdot \mathbf{E}[g(X,Y)\mid X]]&amp;lt;/math&amp;gt;.&lt;br /&gt;
looks obscure at the first glance, especially when considering that &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; are not necessarily independent. Nevertheless, the equation follows the simple fact that conditioning on any &amp;lt;math&amp;gt;X=a&amp;lt;/math&amp;gt;, the function value &amp;lt;math&amp;gt;f(X)=f(a)&amp;lt;/math&amp;gt; becomes a constant, thus can be safely taken outside the expectation due to the linearity of expectation. For any value &amp;lt;math&amp;gt;X=a&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[f(X)g(X,Y)\mid X=a]=\mathbf{E}[f(a)g(X,Y)\mid X=a]=f(a)\cdot \mathbf{E}[g(X,Y)\mid X=a].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The proposition holds in more general cases when &amp;lt;math&amp;gt;X, Y&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Z&amp;lt;/math&amp;gt; are a sequence of random variables.&lt;br /&gt;
&lt;br /&gt;
= Martingales =&lt;br /&gt;
A &#039;&#039;&#039;martingale&#039;&#039;&#039; is a random sequence &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; satisfying the following so-called &#039;&#039;martingale property&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (martingale)|&lt;br /&gt;
:A sequence of random variables &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; is a &#039;&#039;&#039;martingale&#039;&#039;&#039; if for all &amp;lt;math&amp;gt;i&amp;gt; 0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:: &amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}[X_{i}\mid X_0,\ldots,X_{i-1}]=X_{i-1}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
==Examples ==&lt;br /&gt;
;coin flips&lt;br /&gt;
:A fair coin is flipped for a number of times. Let &amp;lt;math&amp;gt;Z_j\in\{-1,1\}&amp;lt;/math&amp;gt; denote the outcome of the &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;th flip. Let &lt;br /&gt;
::&amp;lt;math&amp;gt;X_0=0\quad \mbox{ and } \quad X_i=\sum_{j\le i}Z_j&amp;lt;/math&amp;gt;. &lt;br /&gt;
:The random variables &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; defines a martingale.&lt;br /&gt;
{{Proof| We first observe that &amp;lt;math&amp;gt;\mathbf{E}[X_i\mid X_0,\ldots,X_{i-1}] = \mathbf{E}[X_i\mid X_{i-1}]&amp;lt;/math&amp;gt;, which intuitively says that the next number of HEADs depends only on the current number of HEADs. This property is also called the &#039;&#039;&#039;Markov property&#039;&#039;&#039; in statistic processes.&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}[X_i\mid X_0,\ldots,X_{i-1}] &lt;br /&gt;
&amp;amp;= \mathbf{E}[X_i\mid X_{i-1}]\\&lt;br /&gt;
&amp;amp;= \mathbf{E}[X_{i-1}+Z_{i}\mid X_{i-1}]\\&lt;br /&gt;
&amp;amp;= \mathbf{E}[X_{i-1}\mid X_{i-1}]+\mathbf{E}[Z_{i}\mid X_{i-1}]\\&lt;br /&gt;
&amp;amp;= X_{i-1}+\mathbf{E}[Z_{i}\mid X_{i-1}]\\&lt;br /&gt;
&amp;amp;= X_{i-1}+\mathbf{E}[Z_{i}] &amp;amp;\quad (\mbox{independence of coin flips})\\&lt;br /&gt;
&amp;amp;= X_{i-1}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;edge exposure in a random graph&lt;br /&gt;
:Consider a &#039;&#039;&#039;random graph&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; generated as follows. Let &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt; be the set of vertices, and let &amp;lt;math&amp;gt;[m]={[n]\choose 2}&amp;lt;/math&amp;gt; be the set of all possible edges. For convenience, we enumerate these potential edges by &amp;lt;math&amp;gt;e_1,\ldots, e_m&amp;lt;/math&amp;gt;. For each potential edge &amp;lt;math&amp;gt;e_j&amp;lt;/math&amp;gt;, we independently flip a fair coin to decide whether the edge &amp;lt;math&amp;gt;e_j&amp;lt;/math&amp;gt; appears in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;I_j&amp;lt;/math&amp;gt; be the random variable that indicates whether &amp;lt;math&amp;gt;e_j\in G&amp;lt;/math&amp;gt;. We are interested in some graph-theoretical parameter, say [http://mathworld.wolfram.com/ChromaticNumber.html chromatic number], of the random graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;\chi(G)&amp;lt;/math&amp;gt; be the chromatic number of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X_0=\mathbf{E}[\chi(G)]&amp;lt;/math&amp;gt;, and for each &amp;lt;math&amp;gt;i\ge 1&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_i=\mathbf{E}[\chi(G)\mid I_1,\ldots,I_{i}]&amp;lt;/math&amp;gt;, namely, the expected chromatic number of the random graph after fixing the first &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; edges. This process is called edges exposure of a random graph, as we &amp;quot;exposing&amp;quot; the edges one by one in a random graph.&lt;br /&gt;
&lt;br /&gt;
It is nontrivial to formally verify that the edge exposure sequence for a random graph is a martingale. However, we will later see that this construction can be put into a more general context.&lt;br /&gt;
&lt;br /&gt;
==Generalization ==&lt;br /&gt;
&lt;br /&gt;
The martingale can be generalized to be with respect to another sequence of random variables.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (martingale, general version)| &lt;br /&gt;
:A sequence of random variables &amp;lt;math&amp;gt;Y_0,Y_1,\ldots&amp;lt;/math&amp;gt; is a martingale with respect to the sequence &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; if, for all &amp;lt;math&amp;gt;i\ge 0&amp;lt;/math&amp;gt;, the following conditions hold:&lt;br /&gt;
:* &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is a function of &amp;lt;math&amp;gt;X_0,X_1,\ldots,X_i&amp;lt;/math&amp;gt;;&lt;br /&gt;
:* &amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}[Y_{i+1}\mid X_0,\ldots,X_{i}]=Y_{i}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
Therefore, a sequence &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; is a martingale if it is a martingale with respect to itself.&lt;br /&gt;
&lt;br /&gt;
The purpose of this generalization is that we are usually more interested in a function of a sequence of random variables, rather than the sequence itself.&lt;br /&gt;
&lt;br /&gt;
The following definition describes a very general approach for constructing an important type of martingales.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (The Doob sequence)|&lt;br /&gt;
: The Doob sequence of a function &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; with respect to a sequence of random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt; is defined by&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i}], \quad 0\le i\le n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:In particular, &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(X_1,\ldots,X_n)]&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y_n=f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The Doob sequence of a function defines a martingale. That is&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y_i\mid X_1,\ldots,X_{i-1}]=Y_{i-1}, &lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
for any &amp;lt;math&amp;gt;0\le i\le n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
To prove this claim, we recall the definition that &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i}]&amp;lt;/math&amp;gt;, thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}[Y_i\mid X_1,\ldots,X_{i-1}]&lt;br /&gt;
&amp;amp;=\mathbf{E}[\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i}]\mid X_1,\ldots,X_{i-1}]\\&lt;br /&gt;
&amp;amp;=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i-1}]\\&lt;br /&gt;
&amp;amp;=Y_{i-1},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the second equation is due to the fundamental fact about conditional expectation introduced in the first section.&lt;br /&gt;
&lt;br /&gt;
The Doob martingale describes a very natural procedure to determine a function value of a sequence of random variables. Suppose that we want to predict the value of a function &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; of random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;. The Doob sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; represents a sequence of refined estimates of the value of &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;, gradually using more information on the values of the random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;. The first element &amp;lt;math&amp;gt;Y_0&amp;lt;/math&amp;gt; is just the expectation of &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;. Element &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is the expected value of &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; when the values of &amp;lt;math&amp;gt;X_1,\ldots,X_{i}&amp;lt;/math&amp;gt; are known, and &amp;lt;math&amp;gt;Y_n=f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; is fully determined by &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The following two Doob martingales arise in evaluating the parameters of random graphs. &lt;br /&gt;
&lt;br /&gt;
===edge exposure martingale===&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; be a random graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a real-valued function of graphs, such as, chromatic number, number of triangles, the size of the largest clique or independent set, etc. Denote that &amp;lt;math&amp;gt;m={n\choose 2}&amp;lt;/math&amp;gt;. Fix an arbitrary numbering of potential edges between the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, and denote the edges as &amp;lt;math&amp;gt;e_1,\ldots,e_m&amp;lt;/math&amp;gt;. Let &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
X_i=\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{if }e_i\in G,\\&lt;br /&gt;
0&amp;amp; \mbox{otherwise}.&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Let &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(G)]&amp;lt;/math&amp;gt; and for &amp;lt;math&amp;gt;i=1,\ldots,m&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(G)\mid X_1,\ldots,X_i]&amp;lt;/math&amp;gt;.&lt;br /&gt;
:The sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; gives a Doob martingale that is commonly called the &#039;&#039;&#039;edge exposure martingale&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
===vertex exposure martingale===&lt;br /&gt;
: Instead of revealing edges one at a time, we could reveal the set of edges connected to a given vertex, one vertex at a time. Suppose that the vertex set is &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by the vertex set &amp;lt;math&amp;gt;[i]&amp;lt;/math&amp;gt;, i.e. the first &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:Let &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(G)]&amp;lt;/math&amp;gt; and for &amp;lt;math&amp;gt;i=1,\ldots,n&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(G)\mid X_1,\ldots,X_i]&amp;lt;/math&amp;gt;.&lt;br /&gt;
:The sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; gives a Doob martingale that is commonly called the &#039;&#039;&#039;vertex exposure martingale&#039;&#039;&#039;.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Concentration_of_measure&amp;diff=13945</id>
		<title>高级算法 (Fall 2026)/Concentration of measure</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Concentration_of_measure&amp;diff=13945"/>
		<updated>2026-09-15T11:59:57Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;=Chernoff Bound=  Suppose that we have a fair coin. If we toss it once, then the outcome is completely unpredictable. But if we toss it, say for 1000 times, then the number of HEADs is very likely to be around 500. This phenomenon, as illustrated in the following figure, is called the &amp;#039;&amp;#039;&amp;#039;concentration&amp;#039;&amp;#039;&amp;#039; of measure. The Chernoff bound is an inequality that characterizes the concentration phenomenon for the sum of independent trials.  File:Coinflip.png|border|450px|cent...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Chernoff Bound=&lt;br /&gt;
&lt;br /&gt;
Suppose that we have a fair coin. If we toss it once, then the outcome is completely unpredictable. But if we toss it, say for 1000 times, then the number of HEADs is very likely to be around 500. This phenomenon, as illustrated in the following figure, is called the &#039;&#039;&#039;concentration&#039;&#039;&#039; of measure. The Chernoff bound is an inequality that characterizes the concentration phenomenon for the sum of independent trials.&lt;br /&gt;
&lt;br /&gt;
[[File:Coinflip.png|border|450px|center]]&lt;br /&gt;
&lt;br /&gt;
Before formally stating the Chernoff bound, let&#039;s introduce the &#039;&#039;&#039;moment generating function&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== Moment generating functions ==&lt;br /&gt;
The more we know about the moments of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, the more information we would have about &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. There is a so-called &#039;&#039;&#039;moment generating function&#039;&#039;&#039;, which &amp;quot;packs&amp;quot; all the information about the moments of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; into one function.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition|&lt;br /&gt;
:The moment generating function of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is defined as &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt; is the parameter of the function.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
By Taylor&#039;s expansion and the linearity of expectations,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\sum_{k=0}^\infty\frac{\lambda^k}{k!}X^k\right]\\&lt;br /&gt;
&amp;amp;=\sum_{k=0}^\infty\frac{\lambda^k}{k!}\mathbf{E}\left[X^k\right]&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
&lt;br /&gt;
The moment generating function &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; is a function of &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== The Chernoff bound ==&lt;br /&gt;
The Chernoff bounds are exponentially sharp tail inequalities for the sum of independent trials.&lt;br /&gt;
The bounds are obtained by applying Markov&#039;s inequality to the moment generating function of the sum of independent trials, with some  appropriate choice of the parameter &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Chernoff bound (the upper tail)|&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| For any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;X\ge (1+\delta)\mu&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
where the last step follows by Markov&#039;s inequality.&lt;br /&gt;
&lt;br /&gt;
Computing the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[e^{\lambda \sum_{i=1}^n X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\prod_{i=1}^n e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right].&lt;br /&gt;
&amp;amp; (\mbox{for independent random variables})&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p_i=\Pr[X_i=1]&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;i=1,2,\ldots,n&amp;lt;/math&amp;gt;. Then,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[X]=\mathbf{E}\left[\sum_{i=1}^n X_i\right]=\sum_{i=1}^n\mathbf{E}[X_i]=\sum_{i=1}^n p_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the moment generating function for each individual &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X_i}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
p_i\cdot e^{\lambda\cdot 1}+(1-p_i)\cdot e^{\lambda\cdot 0}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1+p_i(e^\lambda -1)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
e^{p_i(e^\lambda-1)},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
where in the last step we apply the Taylor&#039;s expansion so that &amp;lt;math&amp;gt;e^y\ge 1+y&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;y=p_i(e^\lambda-1)\ge 0&amp;lt;/math&amp;gt;. (By doing this, we can transform the product to the sum of &amp;lt;math&amp;gt;p_i&amp;lt;/math&amp;gt;, which is &amp;lt;math&amp;gt;\mu&amp;lt;/math&amp;gt;.) &lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\prod_{i=1}^n e^{p_i(e^\lambda-1)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(\sum_{i=1}^n p_i(e^{\lambda}-1)\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
e^{(e^\lambda-1)\mu}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, we have shown that for any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{e^{(e^\lambda-1)\mu}}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}\right)^\mu&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;gt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The idea of the proof is actually quite clear: we apply Markov&#039;s inequality to &amp;lt;math&amp;gt;e^{\lambda X}&amp;lt;/math&amp;gt; and for the rest, we just estimate the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;. To make the bound as tight as possible, we minimized the &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt; by setting &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;lt;/math&amp;gt;, which can be justified by taking derivatives of &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
We then proceed to the lower tail, the probability that the random variable deviates below the mean value:&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Chernoff bound (the lower tail)|&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| For any &amp;lt;math&amp;gt;\lambda&amp;lt;0&amp;lt;/math&amp;gt;, by the same analysis as in the upper tail version,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\le (1-\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1-\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1-\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1-\delta)}}\right)^\mu.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
For any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1-\delta)&amp;lt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Useful forms of the Chernoff bounds==&lt;br /&gt;
Some useful special forms of the bounds can be derived directly from the above general forms of the bounds. We now know better why we say that the bounds are exponentially sharp.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Useful forms of the Chernoff bound|&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:1. for &amp;lt;math&amp;gt;0&amp;lt;\delta\le 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{3}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{2}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
:2. for &amp;lt;math&amp;gt;t\ge 2e\mu&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge t]\le 2^{-t}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| To obtain the bounds in (1), we need to show that for &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt; 1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\le e^{-\delta^2/3}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\le e^{-\delta^2/2}&amp;lt;/math&amp;gt;. We can verify both inequalities by standard analysis techniques.&lt;br /&gt;
&lt;br /&gt;
To obtain the bound in (2), let &amp;lt;math&amp;gt;t=(1+\delta)\mu&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\delta=t/\mu-1\ge 2e-1&amp;lt;/math&amp;gt;. Hence,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge(1+\delta)\mu]&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^\delta}{(1+\delta)^{(1+\delta)}}\right)^\mu\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{1+\delta}\right)^{(1+\delta)\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{2e}\right)^t\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
2^{-t}&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Applications to balls-into-bins ==&lt;br /&gt;
Throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls uniformly and independently to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins, what is the maximum load of all bins with high probability? In the last class, we gave an analysis of this problem by using a counting argument.&lt;br /&gt;
&lt;br /&gt;
Now we give a more &amp;quot;advanced&amp;quot; analysis by using Chernoff bounds.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;j\in[m]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_{ij}&amp;lt;/math&amp;gt; be the indicator variable for the event that ball &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; is thrown to bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. Obviously&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_{ij}]=\Pr[\mbox{ball }j\mbox{ is thrown to bin }i]=\frac{1}{n}&amp;lt;/math&amp;gt;&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y_i=\sum_{j\in[m]}X_{ij}&amp;lt;/math&amp;gt; be the load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Then the expected load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;(*)\qquad  \mu=\mathbf{E}[Y_i]=\mathbf{E}\left[\sum_{j\in[m]}X_{ij}\right]=\sum_{j\in[m]}\mathbf{E}[X_{ij}]=m/n.  &amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For the case &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;\mu=1&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is a sum of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; mutually independent indicator variable. Applying Chernoff bound, for any particular bin &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[Y_i&amp;gt;(1+\delta)\mu] \le \left(\frac{e^{\delta}}{(1+\delta)^{1+\delta}}\right)^\mu.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== The &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt; case ===&lt;br /&gt;
&lt;br /&gt;
When &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mu=1&amp;lt;/math&amp;gt;. Write &amp;lt;math&amp;gt;c=1+\delta&amp;lt;/math&amp;gt;. The above bound can be written as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[Y_i&amp;gt;c] \le \frac{e^{c-1}}{c^c}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;c=\frac{e\ln n}{\ln\ln n}&amp;lt;/math&amp;gt;, we evaluate &amp;lt;math&amp;gt;\frac{e^{c-1}}{c^c}&amp;lt;/math&amp;gt; by taking logarithm to its reciprocal.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\ln\left(\frac{c^c}{e^{c-1}}\right)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
c\ln c-c+1\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
c(\ln c-1)+1\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{e\ln n}{\ln\ln n}\left(\ln\ln n-\ln\ln\ln n\right)+1\\&lt;br /&gt;
&amp;amp;\ge&lt;br /&gt;
\frac{e\ln n}{\ln\ln n}\cdot\frac{2}{e}\ln\ln n+1\\&lt;br /&gt;
&amp;amp;\ge&lt;br /&gt;
2\ln n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[Y_i&amp;gt;\frac{e\ln n}{\ln\ln n}\right] \le \frac{1}{n^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Applying the union bound, the probability that there exists a bin with load &amp;lt;math&amp;gt;&amp;gt;12\ln n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;n\cdot \Pr\left[Y_1&amp;gt;\frac{e\ln n}{\ln\ln n}\right] \le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, for &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt;, with high probability, the maximum load is &amp;lt;math&amp;gt;O\left(\frac{e\ln n}{\ln\ln n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== The &amp;lt;math&amp;gt;m&amp;gt; \ln n&amp;lt;/math&amp;gt; case===&lt;br /&gt;
When &amp;lt;math&amp;gt;m\ge n\ln n&amp;lt;/math&amp;gt;, then according to &amp;lt;math&amp;gt;(*)&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mu=\frac{m}{n}\ge \ln n&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We can apply an easier form of the Chernoff bounds,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[Y_i\ge 2e\mu]\le 2^{-2e\mu}\le 2^{-2e\ln n}&amp;lt;\frac{1}{n^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By the union bound, the probability that there exists a bin with load &amp;lt;math&amp;gt;\ge 2e\frac{m}{n}&amp;lt;/math&amp;gt; is,&lt;br /&gt;
:&amp;lt;math&amp;gt;n\cdot \Pr\left[Y_1&amp;gt;2e\frac{m}{n}\right] = n\cdot \Pr\left[Y_1&amp;gt;2e\mu\right]\le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, for &amp;lt;math&amp;gt;m\ge n\ln n&amp;lt;/math&amp;gt;, with high probability, the maximum load is &amp;lt;math&amp;gt;O\left(\frac{m}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Martingales =&lt;br /&gt;
&amp;quot;Martingale&amp;quot; originally refers to a betting strategy in which the gambler doubles his bet after every loss. Assuming unlimited wealth, this strategy is guaranteed to eventually have a positive net profit. For example, starting from an initial stake 1, after &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; losses, if the &amp;lt;math&amp;gt;(n+1)&amp;lt;/math&amp;gt;th bet wins, then it gives a net profit of&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
2^n-\sum_{i=1}^{n}2^{i-1}=1,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is a positive number.&lt;br /&gt;
&lt;br /&gt;
However, the assumption of unlimited wealth is unrealistic. For limited wealth, with geometrically increasing bet, it is very likely to end up bankrupt. You should never try this strategy in real life.&lt;br /&gt;
&lt;br /&gt;
Suppose that the gambler is allowed to use any strategy. His stake on the next beting is decided based on the results of all the bettings so far. This gives us a highly dependent sequence of random variables &amp;lt;math&amp;gt;X_0,X_1,\ldots,&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_0&amp;lt;/math&amp;gt; is his initial capital, and &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; represents his capital after the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th betting. Up to different betting strategies, &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; can be arbitrarily dependent on &amp;lt;math&amp;gt;X_0,\ldots,X_{i-1}&amp;lt;/math&amp;gt;. However, as long as the game is fair, namely, winning and losing with equal chances, conditioning on the past variables &amp;lt;math&amp;gt;X_0,\ldots,X_{i-1}&amp;lt;/math&amp;gt;, we will expect no change in the value of the present variable &amp;lt;math&amp;gt;X_{i}&amp;lt;/math&amp;gt; on average. Random variables satisfying this property is called a &#039;&#039;&#039;martingale&#039;&#039;&#039; sequence.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (martingale)|&lt;br /&gt;
:A sequence of random variables &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; is a &#039;&#039;&#039;martingale&#039;&#039;&#039; if for all &amp;lt;math&amp;gt;i&amp;gt; 0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:: &amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}[X_{i}\mid X_0,\ldots,X_{i-1}]=X_{i-1}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The martingale can be generalized to be with respect to another sequence of random variables.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (martingale, general version)| &lt;br /&gt;
:A sequence of random variables &amp;lt;math&amp;gt;Y_0,Y_1,\ldots&amp;lt;/math&amp;gt; is a martingale with respect to the sequence &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; if, for all &amp;lt;math&amp;gt;i\ge 0&amp;lt;/math&amp;gt;, the following conditions hold:&lt;br /&gt;
:* &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is a function of &amp;lt;math&amp;gt;X_0,X_1,\ldots,X_i&amp;lt;/math&amp;gt;;&lt;br /&gt;
:* &amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}[Y_{i+1}\mid X_0,\ldots,X_{i}]=Y_{i}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
Therefore, a sequence &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; is a martingale if it is a martingale with respect to itself.&lt;br /&gt;
&lt;br /&gt;
The purpose of this generalization is that we are usually more interested in a function of a sequence of random variables, rather than the sequence itself.&lt;br /&gt;
&lt;br /&gt;
==Azuma&#039;s Inequality==&lt;br /&gt;
&lt;br /&gt;
The Azuma&#039;s inequality is a martingale tail inequality.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Azuma&#039;s Inequality|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; be a martingale such that, for all &amp;lt;math&amp;gt;k\ge 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
|X_{k}-X_{k-1}|\le c_k,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|X_n-X_0|\ge t\right]\le 2\exp\left(-\frac{t^2}{2\sum_{k=1}^nc_k^2}\right).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
Unlike the Chernoff bounds, there is no assumption of independence, which makes the martingale inequalities more useful.&lt;br /&gt;
&lt;br /&gt;
The following &#039;&#039;&#039;bounded difference condition&#039;&#039;&#039; &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
|X_{k}-X_{k-1}|\le c_k&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
says that the martingale &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; as a process evolving over time, never makes big change in a single step. &lt;br /&gt;
&lt;br /&gt;
The Azuma&#039;s inequality says that for any martingale satisfying the bounded difference condition, it is unlikely that process wanders far from its starting point.&lt;br /&gt;
&lt;br /&gt;
A special case is when the differences are bounded by a constant.  The following corollary is directly implied by the Azuma&#039;s inequality. &lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Corollary|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; be a martingale such that, for all &amp;lt;math&amp;gt;k\ge 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
|X_{k}-X_{k-1}|\le c,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|X_n-X_0|\ge ct\sqrt{n}\right]\le 2 e^{-t^2/2}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This corollary states that for any martingale sequence whose diferences are bounded by a constant, the probability that it deviates &amp;lt;math&amp;gt;\omega(\sqrt{n})&amp;lt;/math&amp;gt; far away from the starting point after &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; steps is bounded by &amp;lt;math&amp;gt;o(1)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Generalization ===&lt;br /&gt;
&lt;br /&gt;
Azuma&#039;s inequality can be generalized to a martingale with respect another sequence.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Azuma&#039;s Inequality (general version)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;Y_0,Y_1,\ldots&amp;lt;/math&amp;gt; be a martingale with respect to the sequence &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; such that, for all &amp;lt;math&amp;gt;k\ge 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
|Y_{k}-Y_{k-1}|\le c_k,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|Y_n-Y_0|\ge t\right]\le 2\exp\left(-\frac{t^2}{2\sum_{k=1}^nc_k^2}\right).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== The Proof of Azuma&#039;s Inueqality===&lt;br /&gt;
We will only give the formal proof of the non-generalized version. The proof of the general version is almost identical, with the only difference that we work on random sequence &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; conditioning on sequence &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The proof of Azuma&#039;s Inequality uses several ideas which are used in the proof of the Chernoff bounds. We first observe that the total deviation of the martingale sequence can be represented as the sum of deferences in every steps. Thus, as the Chernoff bounds, we are looking for a bound of the deviation of the sum of random variables. The strategy of the proof is almost the same as the proof of Chernoff bounds: we first apply Markov&#039;s inequality to the moment generating function, then we bound the moment generating function, and at last we optimize the parameter of the moment generating function. However, unlike the Chernoff bounds, the martingale differences are not independent any more. So we replace the use of the independence in the Chernoff bound by the martingale property. The proof is detailed as follows.&lt;br /&gt;
&lt;br /&gt;
In order to bound the probability of &amp;lt;math&amp;gt;|X_n-X_0|\ge t&amp;lt;/math&amp;gt;, we first bound the upper tail &amp;lt;math&amp;gt;\Pr[X_n-X_0\ge t]&amp;lt;/math&amp;gt;. The bound of the lower tail can be symmetrically proved with the &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; replaced by &amp;lt;math&amp;gt;-X_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==== Represent the deviation as the sum of differences ====&lt;br /&gt;
We define the &#039;&#039;&#039;martingale difference sequence&#039;&#039;&#039;: for &amp;lt;math&amp;gt;i\ge 1&amp;lt;/math&amp;gt;, let&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=X_i-X_{i-1}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
It holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}[Y_i\mid X_0,\ldots,X_{i-1}]&lt;br /&gt;
&amp;amp;=\mathbf{E}[X_i-X_{i-1}\mid X_0,\ldots,X_{i-1}]\\&lt;br /&gt;
&amp;amp;=\mathbf{E}[X_i\mid X_0,\ldots,X_{i-1}]-\mathbf{E}[X_{i-1}\mid X_0,\ldots,X_{i-1}]\\&lt;br /&gt;
&amp;amp;=X_{i-1}-X_{i-1}\\&lt;br /&gt;
&amp;amp;=0.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The second to the last equation is due to the fact that &amp;lt;math&amp;gt;X_0,X_1,\ldots&amp;lt;/math&amp;gt; is a martingale and the definition of conditional expectation.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;Z_n&amp;lt;/math&amp;gt; be the accumulated differences&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Z_n=\sum_{i=1}^n Y_i.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The deviation &amp;lt;math&amp;gt;(X_n-X_0)&amp;lt;/math&amp;gt; can be computed by the accumulated differences:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
X_n-X_0&lt;br /&gt;
&amp;amp;=(X_1-X_{0})+(X_2-X_1)+\cdots+(X_n-X_{n-1})\\&lt;br /&gt;
&amp;amp;=\sum_{i=1}^n Y_i\\&lt;br /&gt;
&amp;amp;=Z_n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We then only need to upper bound the probability of the event &amp;lt;math&amp;gt;Z_n\ge t&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==== Apply Markov&#039;s inequality to the moment generating function ====&lt;br /&gt;
The event &amp;lt;math&amp;gt;Z_n\ge t&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;e^{\lambda Z_n}\ge e^{\lambda t}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;. Apply Markov&#039;s inequality, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[Z_n\ge t\right]&lt;br /&gt;
&amp;amp;=\Pr\left[e^{\lambda Z_n}\ge e^{\lambda t}\right]\\&lt;br /&gt;
&amp;amp;\le \frac{\mathbf{E}\left[e^{\lambda Z_n}\right]}{e^{\lambda t}}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This is exactly the same as what we did to prove the Chernoff bound. Next, we need to bound the moment generating function &amp;lt;math&amp;gt;\mathbf{E}\left[e^{\lambda Z_n}\right]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==== Bound the moment generating functions ====&lt;br /&gt;
The moment generating function&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda Z_n}\right]&lt;br /&gt;
&amp;amp;=\mathbf{E}\left[\mathbf{E}\left[e^{\lambda Z_n}\mid X_0,\ldots,X_{n-1}\right]\right]\\&lt;br /&gt;
&amp;amp;=\mathbf{E}\left[\mathbf{E}\left[e^{\lambda (Z_{n-1}+Y_n)}\mid X_0,\ldots,X_{n-1}\right]\right]\\&lt;br /&gt;
&amp;amp;=\mathbf{E}\left[\mathbf{E}\left[e^{\lambda Z_{n-1}}\cdot e^{\lambda Y_n}\mid X_0,\ldots,X_{n-1}\right]\right]\\&lt;br /&gt;
&amp;amp;=\mathbf{E}\left[e^{\lambda Z_{n-1}}\cdot\mathbf{E}\left[e^{\lambda Y_n}\mid X_0,\ldots,X_{n-1}\right]\right]&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The first and the last equations are due to the fundamental facts about conditional expectation which are proved by us in the first section.&lt;br /&gt;
&lt;br /&gt;
We then upper bound the &amp;lt;math&amp;gt;\mathbf{E}\left[e^{\lambda Y_n}\mid X_0,\ldots,X_{n-1}\right]&amp;lt;/math&amp;gt; by a constant. To do so, we need the following technical lemma which is proved by the convexity of  &amp;lt;math&amp;gt;e^{\lambda Y_n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lemma|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be a random variable such that &amp;lt;math&amp;gt;\mathbf{E}[X]=0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|X|\le c&amp;lt;/math&amp;gt;. Then for &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[e^{\lambda X}]\le e^{\lambda^2c^2/2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Observe that for &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;, the function &amp;lt;math&amp;gt;e^{\lambda X}&amp;lt;/math&amp;gt; of the variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is convex in the interval &amp;lt;math&amp;gt;[-c,c]&amp;lt;/math&amp;gt;. We draw a line between the two endpoints points &amp;lt;math&amp;gt;(-c, e^{-\lambda c})&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(c, e^{\lambda c})&amp;lt;/math&amp;gt;. The curve of &amp;lt;math&amp;gt;e^{\lambda X}&amp;lt;/math&amp;gt; lies entirely below this line. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
e^{\lambda X}&lt;br /&gt;
&amp;amp;\le \frac{c-X}{2c}e^{-\lambda c}+\frac{c+X}{2c}e^{\lambda c}\\&lt;br /&gt;
&amp;amp;=\frac{e^{\lambda c}+e^{-\lambda c}}{2}+\frac{X}{2c}(e^{\lambda c}-e^{-\lambda c}).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Since &amp;lt;math&amp;gt;\mathbf{E}[X]=0&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}[e^{\lambda X}]&lt;br /&gt;
&amp;amp;\le \mathbf{E}[\frac{e^{\lambda c}+e^{-\lambda c}}{2}+\frac{X}{2c}(e^{\lambda c}-e^{-\lambda c})]\\&lt;br /&gt;
&amp;amp;=\frac{e^{\lambda c}+e^{-\lambda c}}{2}+\frac{e^{\lambda c}-e^{-\lambda c}}{2c}\mathbf{E}[X]\\&lt;br /&gt;
&amp;amp;=\frac{e^{\lambda c}+e^{-\lambda c}}{2}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By expanding both sides as Taylor&#039;s series, it can be verified that &amp;lt;math&amp;gt;\frac{e^{\lambda c}+e^{-\lambda c}}{2}\le e^{\lambda^2c^2/2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Apply the above lemma to the random variable&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(Y_n \mid X_0,\ldots,X_{n-1})&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We have already shown that its expectation &lt;br /&gt;
&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[(Y_n \mid X_0,\ldots,X_{n-1})]=0,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and by the bounded difference condition of Azuma&#039;s inequality, we have&lt;br /&gt;
&amp;lt;math&amp;gt;&lt;br /&gt;
|Y_n|=|(X_n-X_{n-1})|\le c_n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, due to the above lemma, it holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[e^{\lambda Y_n}\mid X_0,\ldots,X_{n-1}]\le e^{\lambda^2c_n^2/2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Back to our analysis of the expectation &amp;lt;math&amp;gt;\mathbf{E}\left[e^{\lambda Z_n}\right]&amp;lt;/math&amp;gt;, we have &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda Z_n}\right]&lt;br /&gt;
&amp;amp;=\mathbf{E}\left[e^{\lambda Z_{n-1}}\cdot\mathbf{E}\left[e^{\lambda Y_n}\mid X_0,\ldots,X_{n-1}\right]\right]\\&lt;br /&gt;
&amp;amp;\le \mathbf{E}\left[e^{\lambda Z_{n-1}}\cdot e^{\lambda^2c_n^2/2}\right]\\&lt;br /&gt;
&amp;amp;= e^{\lambda^2c_n^2/2}\cdot\mathbf{E}\left[e^{\lambda Z_{n-1}}\right] .&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Apply the same analysis to &amp;lt;math&amp;gt;\mathbf{E}\left[e^{\lambda Z_{n-1}}\right]&amp;lt;/math&amp;gt;, we can solve the above recursion by&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda Z_n}\right]&lt;br /&gt;
&amp;amp;\le \prod_{k=1}^n e^{\lambda^2c_k^2/2}\\&lt;br /&gt;
&amp;amp;= \exp\left(\lambda^2\sum_{k=1}^n c_k^2/2\right).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Go back to the Markov&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[Z_n\ge t\right]&lt;br /&gt;
&amp;amp;\le \frac{\mathbf{E}\left[e^{\lambda Z_n}\right]}{e^{\lambda t}}\\&lt;br /&gt;
&amp;amp;\le \exp\left(\lambda^2\sum_{k=1}^n c_k^2/2-\lambda t\right).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We then only need to choose a proper &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==== Optimization ====&lt;br /&gt;
By choosing &amp;lt;math&amp;gt;\lambda=\frac{t}{\sum_{k=1}^n c_k^2}&amp;lt;/math&amp;gt;, we have that &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\exp\left(\lambda^2\sum_{k=1}^n c_k^2/2-\lambda t\right)=\exp\left(-\frac{t^2}{2\sum_{k=1}^n c_k^2}\right).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, the probability&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[X_n-X_0\ge t\right]&lt;br /&gt;
&amp;amp;=\Pr\left[Z_n\ge t\right]\\&lt;br /&gt;
&amp;amp;\le \exp\left(\lambda^2\sum_{k=1}^n c_k^2/2-\lambda t\right)\\&lt;br /&gt;
&amp;amp;= \exp\left(-\frac{t^2}{2\sum_{k=1}^n c_k^2}\right).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The upper tail of Azuma&#039;s inequality is proved. By replacing &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;-X_i&amp;lt;/math&amp;gt;, the lower tail can be treated just as the upper tail. Applying the union bound, Azuma&#039;s inequality is proved.&lt;br /&gt;
&lt;br /&gt;
==The Doob martingales ==&lt;br /&gt;
The following definition describes a very general approach for constructing an important type of martingales.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (The Doob sequence)|&lt;br /&gt;
: The Doob sequence of a function &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; with respect to a sequence of random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt; is defined by&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i}], \quad 0\le i\le n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:In particular, &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(X_1,\ldots,X_n)]&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y_n=f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The Doob sequence of a function defines a martingale. That is&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y_i\mid X_1,\ldots,X_{i-1}]=Y_{i-1}, &lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
for any &amp;lt;math&amp;gt;0\le i\le n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
To prove this claim, we recall the definition that &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i}]&amp;lt;/math&amp;gt;, thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}[Y_i\mid X_1,\ldots,X_{i-1}]&lt;br /&gt;
&amp;amp;=\mathbf{E}[\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i}]\mid X_1,\ldots,X_{i-1}]\\&lt;br /&gt;
&amp;amp;=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_{i-1}]\\&lt;br /&gt;
&amp;amp;=Y_{i-1},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the second equation is due to the fundamental fact about conditional expectation introduced in the first section.&lt;br /&gt;
&lt;br /&gt;
The Doob martingale describes a very natural procedure to determine a function value of a sequence of random variables. Suppose that we want to predict the value of a function &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; of random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;. The Doob sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; represents a sequence of refined estimates of the value of &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;, gradually using more information on the values of the random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;. The first element &amp;lt;math&amp;gt;Y_0&amp;lt;/math&amp;gt; is just the expectation of &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;. Element &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is the expected value of &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; when the values of &amp;lt;math&amp;gt;X_1,\ldots,X_{i}&amp;lt;/math&amp;gt; are known, and &amp;lt;math&amp;gt;Y_n=f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;f(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; is fully determined by &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The following two Doob martingales arise in evaluating the parameters of random graphs. &lt;br /&gt;
&lt;br /&gt;
;edge exposure martingale&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; be a random graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a real-valued function of graphs, such as, chromatic number, number of triangles, the size of the largest clique or independent set, etc. Denote that &amp;lt;math&amp;gt;m={n\choose 2}&amp;lt;/math&amp;gt;. Fix an arbitrary numbering of potential edges between the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, and denote the edges as &amp;lt;math&amp;gt;e_1,\ldots,e_m&amp;lt;/math&amp;gt;. Let &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
X_i=\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{if }e_i\in G,\\&lt;br /&gt;
0&amp;amp; \mbox{otherwise}.&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Let &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(G)]&amp;lt;/math&amp;gt; and for &amp;lt;math&amp;gt;i=1,\ldots,m&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(G)\mid X_1,\ldots,X_i]&amp;lt;/math&amp;gt;.&lt;br /&gt;
:The sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; gives a Doob martingale that is commonly called the &#039;&#039;&#039;edge exposure martingale&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
;vertex exposure martingale&lt;br /&gt;
: Instead of revealing edges one at a time, we could reveal the set of edges connected to a given vertex, one vertex at a time. Suppose that the vertex set is &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by the vertex set &amp;lt;math&amp;gt;[i]&amp;lt;/math&amp;gt;, i.e. the first &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:Let &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(G)]&amp;lt;/math&amp;gt; and for &amp;lt;math&amp;gt;i=1,\ldots,n&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(G)\mid X_1,\ldots,X_i]&amp;lt;/math&amp;gt;.&lt;br /&gt;
:The sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; gives a Doob martingale that is commonly called the &#039;&#039;&#039;vertex exposure martingale&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
===Chromatic number===&lt;br /&gt;
The random graph &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt; is the graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;, obtained by selecting each pair of vertices to be an edge, randomly and independently, with probability &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. We denote &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is generated in this way.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem [Shamir and Spencer (1987)]|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;\chi(G)&amp;lt;/math&amp;gt; be the chromatic number of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|\chi(G)-\mathbf{E}[\chi(G)]|\ge t\sqrt{n}\right]\le 2e^{-t^2/2}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Consider the vertex exposure martingale &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=\mathbf{E}[\chi(G)\mid X_1,\ldots,X_i]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where each &amp;lt;math&amp;gt;X_k&amp;lt;/math&amp;gt; exposes the induced subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; on vertex set &amp;lt;math&amp;gt;[k]&amp;lt;/math&amp;gt;. A single vertex can always be given a new color so that the graph is properly colored, thus the bounded difference condition &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
|Y_i-Y_{i-1}|\le 1&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
is satisfied. Now apply the Azuma&#039;s inequality for the martingale &amp;lt;math&amp;gt;Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; with respect to &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;t=\omega(1)&amp;lt;/math&amp;gt;, the theorem states that the chromatic number of a random graph is tightly concentrated around its mean. The proof gives no clue as to where the mean is. This actually shows how powerful the martingale inequalities are: we can prove that a distribution is concentrated to its expectation without actually knowing the expectation. &lt;br /&gt;
&lt;br /&gt;
=== Hoeffding&#039;s Inequality===&lt;br /&gt;
The following theorem states the so-called Hoeffding&#039;s inequality. It is a generalized version of the Chernoff bounds. Recall that the Chernoff bounds hold for the sum of independent &#039;&#039;trials&#039;&#039;. When the random variables are not trials, the Hoeffding&#039;s inequality is useful, since it holds for the sum of any independent random variables whose ranges are bounded.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Hoeffding&#039;s inequality|&lt;br /&gt;
: Let &amp;lt;math&amp;gt;X=\sum_{i=1}^nX_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt; are independent random variables with &amp;lt;math&amp;gt;a_i\le X_i\le b_i&amp;lt;/math&amp;gt; for each &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[|X-\mu|\ge t]\le 2\exp\left(-\frac{t^2}{2\sum_{i=1}^n(b_i-a_i)^2}\right).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Define the Doob martingale sequence &amp;lt;math&amp;gt;Y_i=\mathbf{E}\left[\sum_{j=1}^n X_j\,\Big|\, X_1,\ldots,X_{i}\right]&amp;lt;/math&amp;gt;. Obviously &amp;lt;math&amp;gt;Y_0=\mu&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y_n=X&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
|Y_i-Y_{i-1}|&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\left|\mathbf{E}\left[\sum_{j=1}^n X_j\,\Big|\, X_0,\ldots,X_{i}\right]-\mathbf{E}\left[\sum_{j=1}^n X_j\,\Big|\, X_0,\ldots,X_{i-1}\right]\right|\\&lt;br /&gt;
&amp;amp;=\left|\sum_{j=1}^i X_i+\sum_{j=i+1}^n\mathbf{E}[X_j]-\sum_{j=1}^{i-1} X_i-\sum_{j=i}^n\mathbf{E}[X_j]\right|\\&lt;br /&gt;
&amp;amp;=\left|X_i-\mathbf{E}[X_{i}]\right|\\&lt;br /&gt;
&amp;amp;\le b_i-a_i&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Apply Azuma&#039;s inequality for the martingale &amp;lt;math&amp;gt;Y_0,\ldots,Y_n&amp;lt;/math&amp;gt; with respect to &amp;lt;math&amp;gt;X_1,\ldots, X_n&amp;lt;/math&amp;gt;,  the Hoeffding&#039;s inequality is proved.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
==The Bounded Difference Method==&lt;br /&gt;
Combining Azuma&#039;s inequality with the construction of Doob martingales, we have the powerful &#039;&#039;Bounded Difference Method&#039;&#039; for concentration of measures.&lt;br /&gt;
&lt;br /&gt;
=== For arbitrary random variables ===&lt;br /&gt;
Given a sequence of random variables &amp;lt;math&amp;gt;X_1,\ldots,X_n&amp;lt;/math&amp;gt; and a function &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;.  The Doob sequence constructs a martingale from them. Combining this construction with Azuma&#039;s inequality, we can get a very powerful theorem called &amp;quot;the method of averaged bounded differences&amp;quot; which bounds the concentration for arbitrary function on arbitrary random variables (not necessarily a martingale).&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem (Method of averaged bounded differences)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\boldsymbol{X}=(X_1,\ldots, X_n)&amp;lt;/math&amp;gt; be arbitrary random variables and let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a function of &amp;lt;math&amp;gt;X_0,X_1,\ldots, X_n&amp;lt;/math&amp;gt; satisfying that, for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
|\mathbf{E}[f(\boldsymbol{X})\mid X_1,\ldots,X_i]-\mathbf{E}[f(\boldsymbol{X})\mid X_1,\ldots,X_{i-1}]|\le c_i,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|f(\boldsymbol{X})-\mathbf{E}[f(\boldsymbol{X})]|\ge t\right]\le 2\exp\left(-\frac{t^2}{2\sum_{i=1}^nc_i^2}\right).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Define the Doob Martingale sequence &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt; by setting &amp;lt;math&amp;gt;Y_0=\mathbf{E}[f(X_1,\ldots,X_n)]&amp;lt;/math&amp;gt; and, for &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;Y_i=\mathbf{E}[f(X_1,\ldots,X_n)\mid X_1,\ldots,X_i]&amp;lt;/math&amp;gt;. Then the above theorem is a restatement of the Azuma&#039;s inequality holding for &amp;lt;math&amp;gt;Y_0,Y_1,\ldots,Y_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== For independent random variables ===&lt;br /&gt;
The condition of bounded averaged differences is usually hard to check. This severely limits the usefulness of the method. To overcome this, we introduce a property which is much easier to check, called the Lipschitz condition.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (Lipschitz condition)|&lt;br /&gt;
:A function &amp;lt;math&amp;gt;f(x_1,\ldots,x_n)&amp;lt;/math&amp;gt; satisfies the Lipschitz condition, if for any &amp;lt;math&amp;gt;x_1,\ldots,x_n&amp;lt;/math&amp;gt; and any &amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
|f(x_1,\ldots,x_{i-1},x_i,x_{i+1},\ldots,x_n)-f(x_1,\ldots,x_{i-1},y_i,x_{i+1},\ldots,x_n)|\le 1.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
In other words, the function satisfies the Lipschitz condition if an arbitrary change in the value of any one argument does not change the value of the function by more than 1. &lt;br /&gt;
&lt;br /&gt;
The diference of 1 can be replaced by arbitrary constants, which gives a generalized version of Lipschitz condition.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (Lipschitz condition, general version)|&lt;br /&gt;
:A function &amp;lt;math&amp;gt;f(x_1,\ldots,x_n)&amp;lt;/math&amp;gt; satisfies the Lipschitz condition with constants &amp;lt;math&amp;gt;c_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, if for any &amp;lt;math&amp;gt;x_1,\ldots,x_n&amp;lt;/math&amp;gt; and any &amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
|f(x_1,\ldots,x_{i-1},x_i,x_{i+1},\ldots,x_n)-f(x_1,\ldots,x_{i-1},y_i,x_{i+1},\ldots,x_n)|\le c_i.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The following &amp;quot;method of bounded differences&amp;quot; can be developed for functions satisfying the Lipschitz condition. Unfortunately, in order to imply the condition of averaged bounded differences from the Lipschitz condition, we have to restrict the method to independent random variables.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Corollary (Method of bounded differences)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\boldsymbol{X}=(X_1,\ldots, X_n)&amp;lt;/math&amp;gt; be &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;&#039;independent&#039;&#039;&#039; random variables and let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a function satisfying the Lipschitz condition with constants &amp;lt;math&amp;gt;c_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|f(\boldsymbol{X})-\mathbf{E}[f(\boldsymbol{X})]|\ge t\right]\le 2\exp\left(-\frac{t^2}{2\sum_{i=1}^nc_i^2}\right).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Proof| For convenience, we denote that &amp;lt;math&amp;gt;\boldsymbol{X}_{[i,j]}=(X_i,X_{i+1},\ldots, X_j)&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;1\le i\le j\le n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We first show that the Lipschitz condition with constants &amp;lt;math&amp;gt;c_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, implies another condition called the averaged Lipschitz condition (ALC): for any &amp;lt;math&amp;gt;a_i,b_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left|\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a_i\right]-\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=b_i\right]\right|\le c_i.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
And this condition implies the averaged bounded difference condition: for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\left|\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]-\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]}\right]\right|\le c_i.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Then by applying the method of averaged bounded differences, the corollary can be proved.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;, by the law of total expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\quad\, \mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a\right]\\&lt;br /&gt;
&amp;amp;=\sum_{a_{i+1},\ldots,a_n}\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a, \boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right]\cdot\Pr\left[\boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\mid \boldsymbol{X}_{[1,i-1]},X_i=a\right]\\&lt;br /&gt;
&amp;amp;=\sum_{a_{i+1},\ldots,a_n}\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a, \boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right]\cdot\Pr\left[\boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right] \qquad (\mbox{independence})\\&lt;br /&gt;
&amp;amp;= \sum_{a_{i+1},\ldots,a_n} f(\boldsymbol{X}_{[1,i-1]},a,\boldsymbol{a}_{[i+1,n]})\cdot\Pr\left[\boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;a=a_i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;, and take the diference. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\quad\, \left|\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a_i\right]-\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=b_i\right]\right|\\&lt;br /&gt;
&amp;amp;=\left|\sum_{a_{i+1},\ldots,a_n}\left(f(\boldsymbol{X}_{[1,i-1]},a_i,\boldsymbol{a}_{[i+1,n]})-f(\boldsymbol{X}_{[1,i-1]},b_i,\boldsymbol{a}_{[i+1,n]})\right)\Pr\left[\boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right]\right|\\&lt;br /&gt;
&amp;amp;\le \sum_{a_{i+1},\ldots,a_n}\left|f(\boldsymbol{X}_{[1,i-1]},a_i,\boldsymbol{a}_{[i+1,n]})-f(\boldsymbol{X}_{[1,i-1]},b_i,\boldsymbol{a}_{[i+1,n]})\right|\Pr\left[\boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right]\\&lt;br /&gt;
&amp;amp;\le \sum_{a_{i+1},\ldots,a_n}c_i\Pr\left[\boldsymbol{X}_{[i+1,n]}=\boldsymbol{a}_{[i+1,n]}\right] \qquad (\mbox{Lipschitz condition})\\&lt;br /&gt;
&amp;amp;=c_i.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Thus, the Lipschitz condition is transformed to the ALC. We then deduce the averaged bounded difference condition from ALC.&lt;br /&gt;
&lt;br /&gt;
By the law of total expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]}\right]=\sum_{a}\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a\right]\cdot\Pr[X_i=a\mid \boldsymbol{X}_{[1,i-1]}].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We can trivially write &amp;lt;math&amp;gt;\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]&amp;lt;/math&amp;gt; as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]=\sum_{a}\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]\cdot\Pr\left[X_i=a\mid \boldsymbol{X}_{[1,i-1]}\right].&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Hence, the difference is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\quad \left|\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]-\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]}\right]\right|\\&lt;br /&gt;
&amp;amp;=\left|\sum_{a}\left(\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]-\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a\right]\right)\cdot\Pr\left[X_i=a\mid \boldsymbol{X}_{[1,i-1]}\right]\right| \\&lt;br /&gt;
&amp;amp;\le \sum_{a}\left|\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i]}\right]-\mathbf{E}\left[f(\boldsymbol{X})\mid \boldsymbol{X}_{[1,i-1]},X_i=a\right]\right|\cdot\Pr\left[X_i=a\mid \boldsymbol{X}_{[1,i-1]}\right] \\&lt;br /&gt;
&amp;amp;\le \sum_a c_i\Pr\left[X_i=a\mid \boldsymbol{X}_{[1,i-1]}\right] \qquad (\mbox{due to ALC})\\&lt;br /&gt;
&amp;amp;=c_i.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The averaged bounded diference condition is implied. Applying the method of averaged bounded diferences, the corollary follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Applications ===&lt;br /&gt;
&lt;br /&gt;
==== Occupancy problem ====&lt;br /&gt;
Throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls uniformly and independently at random to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins, we ask for the occupancies of bins by the balls. In particular, we are interested in the number of empty bins.&lt;br /&gt;
&lt;br /&gt;
This problem can be described equivalently as follows. Let &amp;lt;math&amp;gt;f:[m]\rightarrow[n]&amp;lt;/math&amp;gt; be a uniform random function from &amp;lt;math&amp;gt;[m]\rightarrow[n]&amp;lt;/math&amp;gt;. We ask for the number of &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;f^{-1}(i)&amp;lt;/math&amp;gt; is empty.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; indicate the emptiness of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X=\sum_{i=1}^nX_i&amp;lt;/math&amp;gt; be the number of empty bins.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[X_i]=\Pr[\mbox{bin }i\mbox{ is empty}]=\left(1-\frac{1}{n}\right)^m.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By the linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[X]=\sum_{i=1}^n\mathbf{E}[X_i]=n\left(1-\frac{1}{n}\right)^m.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We want to know how &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; deviates from this expectation. The complication here is that &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; are not independent. So we alternatively look at a sequence of independent random variables &amp;lt;math&amp;gt;Y_1,\ldots, Y_m&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;Y_j\in[n]&amp;lt;/math&amp;gt; represents the bin into which the &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;th ball falls. Clearly &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is function of &amp;lt;math&amp;gt;Y_1,\ldots, Y_m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We than observe that changing the value of any &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; can change the value of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; by at most 1, because one ball can affect the emptiness of at most one bin. &lt;br /&gt;
Thus as a function of independent random variables &amp;lt;math&amp;gt;Y_1,\ldots, Y_m&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; satisfies the Lipschitz condition. Apply the method of bounded differences, it holds that &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\left|X-n\left(1-\frac{1}{n}\right)^m\right|\ge t\sqrt{m}\right]=\Pr[|X-\mathbf{E}[X]|\ge t\sqrt{m}]\le 2e^{-t^2/2}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Thus, for sufficiently large &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;, the number of empty bins is tightly concentrated around &amp;lt;math&amp;gt;n\left(1-\frac{1}{n}\right)^m\approx \frac{n}{e^{m/n}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Pattern Matching ====&lt;br /&gt;
Let &amp;lt;math&amp;gt;\boldsymbol{X}=(X_1,\ldots,X_n)&amp;lt;/math&amp;gt; be a sequence of characters chosen independently and uniformly at random from an alphabet &amp;lt;math&amp;gt;\Sigma&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;m=|\Sigma|&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;\pi\in\Sigma^k&amp;lt;/math&amp;gt; be an arbitrarily fixed string of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; characters from &amp;lt;math&amp;gt;\Sigma&amp;lt;/math&amp;gt;, called a &#039;&#039;pattern&#039;&#039;. Let &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; be the number of occurrences of the pattern &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; as a substring of the random string &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By the linearity of expectation, it is obvious that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y]=(n-k+1)\left(\frac{1}{m}\right)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We now look at the concentration of &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;. The complication again lies in the dependencies between the matches. Yet we will see that &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is well tightly concentrated around its expectation if &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; is relatively small compared to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For a fixed pattern &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, the random variable &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is a function of the independent random variables &amp;lt;math&amp;gt;(X_1,\ldots,X_n)&amp;lt;/math&amp;gt;. Any character &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; participates in no more than &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; matches, thus changing the value of any &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; can affect the value of &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; by at most &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; satisfies the Lipschitz condition with constant &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. Apply the method of bounded differences,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\left|Y-\frac{n-k+1}{m^k}\right|\ge tk\sqrt{n}\right]=\Pr\left[\left|Y-\mathbf{E}[Y]\right|\ge  tk\sqrt{n}\right]\le 2e^{-t^2/2}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Combining unit vectors ====&lt;br /&gt;
Let &amp;lt;math&amp;gt;u_1,\ldots,u_n&amp;lt;/math&amp;gt; be &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; unit vectors from some normed space. That is, &amp;lt;math&amp;gt;\|u_i\|=1&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;\|\cdot\|&amp;lt;/math&amp;gt; denote the vector norm (e.g. &amp;lt;math&amp;gt;\ell_1,\ell_2,\ell_\infty&amp;lt;/math&amp;gt;) of the space.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;\epsilon_1,\ldots,\epsilon_n\in\{-1,+1\}&amp;lt;/math&amp;gt; be independently chosen and &amp;lt;math&amp;gt;\Pr[\epsilon_i=-1]=\Pr[\epsilon_i=1]=1/2&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Let &lt;br /&gt;
:&amp;lt;math&amp;gt;v=\epsilon_1u_1+\cdots+\epsilon_nu_n,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X=\|v\|.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This kind of construction is very useful in combinatorial proofs of metric problems. We will show that by this construction, the random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is well concentrated around its mean.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is a function of independent random variables &amp;lt;math&amp;gt;\epsilon_1,\ldots,\epsilon_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the triangle inequality for norms, it is easy to verify that changing the sign of a unit vector &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt; can only change the value of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; for at most 2, thus &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; satisfies the Lipschitz condition with constant 2. The concentration result follows by applying the method of bounded differences:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[|X-\mathbf{E}[X]|\ge 2t\sqrt{n}]\le 2e^{-t^2/2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13944</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13944"/>
		<updated>2026-09-15T11:59:31Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([http://tcs.nju.edu.cn/slides/aa2026/Hashing.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-zh.pdf AI生成讲义]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
# [[高级算法 (Fall 2026)/Concentration of measure|Concentration of measure]] &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Conditional expectations|Conditional expectations]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13931</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13931"/>
		<updated>2026-09-09T08:30:10Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([http://tcs.nju.edu.cn/slides/aa2026/Hashing.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-hashing-zh.pdf AI生成讲义]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13930</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13930"/>
		<updated>2026-09-09T08:29:22Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([[Media:Hashing-26.pdf|slides]]) ([[Media:AA26-note-hashing-en.pdf|AI-generated lecture note]], [[Media:AA26-note-hashing-zh.pdf|AI生成讲义]]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13917</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13917"/>
		<updated>2026-09-09T04:44:07Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([[Media:Hashing-26.pdf|slides]]) ([[Media:AA26-note-hashing-en.pdf|AI-generated lecture note]], [[Media:AA26-note-hashing-zh.pdf|AI生成讲义]])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13916</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13916"/>
		<updated>2026-09-09T04:43:29Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([[Meida:Hashing-26.pdf|slides]]) ([[Media:AA26-note-hashing-en.pdf|AI-generated lecture note]], [[Media:AA26-note-hashing-zh.pdf|AI生成讲义]])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=File:Hashing-26.pdf&amp;diff=13915</id>
		<title>File:Hashing-26.pdf</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=File:Hashing-26.pdf&amp;diff=13915"/>
		<updated>2026-09-09T04:43:16Z</updated>

		<summary type="html">&lt;p&gt;Etone: Etone uploaded a new version of File:Hashing-26.pdf&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13914</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13914"/>
		<updated>2026-09-09T04:42:41Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([[file:Hashing-26.pdf|slides]]) ([[Media:AA26-note-hashing-en.pdf|AI-generated lecture note]], [[Media:AA26-note-hashing-zh.pdf|AI生成讲义]])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13913</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13913"/>
		<updated>2026-09-08T13:54:23Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([[Media:Hashing-26.pdf|slides]]) ([[Media:AA26-note-hashing-en.pdf|AI-generated lecture note]], [[Media:AA26-note-hashing-zh.pdf|AI生成讲义]])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=File:AA26-note-hashing-zh.pdf&amp;diff=13912</id>
		<title>File:AA26-note-hashing-zh.pdf</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=File:AA26-note-hashing-zh.pdf&amp;diff=13912"/>
		<updated>2026-09-08T13:53:21Z</updated>

		<summary type="html">&lt;p&gt;Etone: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=File:AA26-note-hashing-en.pdf&amp;diff=13911</id>
		<title>File:AA26-note-hashing-en.pdf</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=File:AA26-note-hashing-en.pdf&amp;diff=13911"/>
		<updated>2026-09-08T13:53:06Z</updated>

		<summary type="html">&lt;p&gt;Etone: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13910</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13910"/>
		<updated>2026-09-08T13:52:09Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]] ([[Media:Hashing-26.pdf|slides]])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=File:Hashing-26.pdf&amp;diff=13909</id>
		<title>File:Hashing-26.pdf</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=File:Hashing-26.pdf&amp;diff=13909"/>
		<updated>2026-09-08T13:47:20Z</updated>

		<summary type="html">&lt;p&gt;Etone: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=File:Hashing.pdf&amp;diff=13908</id>
		<title>File:Hashing.pdf</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=File:Hashing.pdf&amp;diff=13908"/>
		<updated>2026-09-08T13:46:04Z</updated>

		<summary type="html">&lt;p&gt;Etone: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Basic_deviation_inequalities&amp;diff=13906</id>
		<title>高级算法 (Fall 2026)/Basic deviation inequalities</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Basic_deviation_inequalities&amp;diff=13906"/>
		<updated>2026-09-07T10:13:47Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;=Markov&amp;#039;s Inequality=  One of the most natural information about a random variable is its expectation, which is the first moment of the random variable. Markov&amp;#039;s inequality draws a tail bound for a random variable from its expectation. {{Theorem |Theorem (Markov&amp;#039;s Inequality)| :Let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be a random variable assuming only nonnegative values. Then, for all &amp;lt;math&amp;gt;t&amp;gt;0&amp;lt;/math&amp;gt;, ::&amp;lt;math&amp;gt;\begin{align} \Pr[X\ge t]\le \frac{\mathbf{E}[X]}{t}. \end{align}&amp;lt;/math&amp;gt; }} {{Proo...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Markov&#039;s Inequality=&lt;br /&gt;
&lt;br /&gt;
One of the most natural information about a random variable is its expectation, which is the first moment of the random variable. Markov&#039;s inequality draws a tail bound for a random variable from its expectation.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem (Markov&#039;s Inequality)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be a random variable assuming only nonnegative values. Then, for all &amp;lt;math&amp;gt;t&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge t]\le \frac{\mathbf{E}[X]}{t}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Let &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; be the indicator such that &lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
Y &amp;amp;=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1 &amp;amp; \mbox{if }X\ge t,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
It holds that &amp;lt;math&amp;gt;Y\le\frac{X}{t}&amp;lt;/math&amp;gt;. Since &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is 0-1 valued, &amp;lt;math&amp;gt;\mathbf{E}[Y]=\Pr[Y=1]=\Pr[X\ge t]&amp;lt;/math&amp;gt;. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[X\ge t]&lt;br /&gt;
=&lt;br /&gt;
\mathbf{E}[Y]&lt;br /&gt;
\le&lt;br /&gt;
\mathbf{E}\left[\frac{X}{t}\right]&lt;br /&gt;
=\frac{\mathbf{E}[X]}{t}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Generalization ==&lt;br /&gt;
For any random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, for an arbitrary non-negative real function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt;, the &amp;lt;math&amp;gt;h(X)&amp;lt;/math&amp;gt; is a non-negative random variable. Applying Markov&#039;s inequality, we directly have that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(X)\ge t]\le\frac{\mathbf{E}[h(X)]}{t}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This trivial application of Markov&#039;s inequality gives us a powerful tool for proving tail inequalities. With the function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; which extracts more information about the random variable, we can prove sharper tail inequalities.&lt;br /&gt;
&lt;br /&gt;
=Chebyshev&#039;s inequality=&lt;br /&gt;
&lt;br /&gt;
== Variance ==&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (variance)|&lt;br /&gt;
:The &#039;&#039;&#039;variance&#039;&#039;&#039; of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is defined as&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Var}[X]=\mathbf{E}\left[(X-\mathbf{E}[X])^2\right]=\mathbf{E}\left[X^2\right]-(\mathbf{E}[X])^2.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
:The &#039;&#039;&#039;standard deviation&#039;&#039;&#039; of random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\delta[X]=\sqrt{\mathbf{Var}[X]}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The variance is the diagonal case for &#039;&#039;&#039;covariance&#039;&#039;&#039;.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (covariance)|&lt;br /&gt;
:The &#039;&#039;&#039;covariance&#039;&#039;&#039; of two random variables &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Cov}(X,Y)=\mathbf{E}\left[(X-\mathbf{E}[X])(Y-\mathbf{E}[Y])\right].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We have the following theorem for the variance of sum.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:For any two random variables &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Var}[X+Y]=\mathbf{Var}[X]+\mathbf{Var}[Y]+2\mathbf{Cov}(X,Y).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
:Generally, for any random variables &amp;lt;math&amp;gt;X_1,X_2,\ldots,X_n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Var}\left[\sum_{i=1}^n X_i\right]=\sum_{i=1}^n\mathbf{Var}[X_i]+\sum_{i\neq j}\mathbf{Cov}(X_i,X_j).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| The equation for two variables is directly due to the definition of variance and covariance. The equation for &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; variables can be deduced from the equation for two variables.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
For independent random variables, the expectation of a product equals the product of expectations. &lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:For any two independent random variables &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}[X\cdot Y]=\mathbf{E}[X]\cdot\mathbf{E}[Y].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{E}[X\cdot Y]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{x,y}xy\Pr[X=x\wedge Y=y]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{x,y}xy\Pr[X=x]\Pr[Y=y]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{x}x\Pr[X=x]\sum_{y}y\Pr[Y=y]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}[X]\cdot\mathbf{E}[Y].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Consequently, covariance of independent random variables is always zero.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:For any two independent random variables &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Cov}(X,Y)=0.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| &lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Cov}(X,Y) &lt;br /&gt;
&amp;amp;=\mathbf{E}\left[(X-\mathbf{E}[X])(Y-\mathbf{E}[Y])\right]\\&lt;br /&gt;
&amp;amp;= \mathbf{E}\left[X-\mathbf{E}[X]\right]\mathbf{E}\left[Y-\mathbf{E}[Y]\right] &amp;amp;\qquad(\mbox{Independence})\\&lt;br /&gt;
&amp;amp;=0.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The variance of the sum of pairwise independent random variables is equal to the sum of variances. &lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:For &#039;&#039;&#039;pairwise&#039;&#039;&#039; independent random variables &amp;lt;math&amp;gt;X_1,X_2,\ldots,X_n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{Var}\left[\sum_{i=1}^n X_i\right]=\sum_{i=1}^n\mathbf{Var}[X_i].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;Remark&lt;br /&gt;
:The theorem holds for &#039;&#039;&#039;pairwise&#039;&#039;&#039; independent random variables, a much weaker independence requirement than the &#039;&#039;&#039;mutual&#039;&#039;&#039; independence. This makes the second-moment methods very useful for pairwise independent random variables.&lt;br /&gt;
&lt;br /&gt;
=== Variance of binomial distribution ===&lt;br /&gt;
For a Bernoulli trial with parameter &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X=\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{with probability }p\\&lt;br /&gt;
0&amp;amp; \mbox{with probability }1-p&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The variance is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{Var}[X]=\mathbf{E}[X^2]-(\mathbf{E}[X])^2=\mathbf{E}[X]-(\mathbf{E}[X])^2=p-p^2=p(1-p).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; be a binomial random variable with parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;, i.e. &amp;lt;math&amp;gt;Y=\sum_{i=1}^nY_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt;&#039;s are i.i.d. Bernoulli trials with parameter &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. The variance is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbf{Var}[Y] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{Var}\left[\sum_{i=1}^nY_i\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{i=1}^n\mathbf{Var}\left[Y_i\right] &amp;amp;\qquad (\mbox{Independence})\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{i=1}^np(1-p) &amp;amp;\qquad (\mbox{Bernoulli})\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
p(1-p)n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Chebyshev&#039;s inequality ==&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem (Chebyshev&#039;s Inequality)|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;t&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[|X-\mathbf{E}[X]| \ge t\right] \le \frac{\mathbf{Var}[X]}{t^2}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Observe that &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[|X-\mathbf{E}[X]| \ge t] = \Pr[(X-\mathbf{E}[X])^2 \ge t^2].&amp;lt;/math&amp;gt;&lt;br /&gt;
Since &amp;lt;math&amp;gt;(X-\mathbf{E}[X])^2&amp;lt;/math&amp;gt; is a nonnegative random variable, we can apply Markov&#039;s inequality, such that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[(X-\mathbf{E}[X])^2 \ge t^2] \le&lt;br /&gt;
\frac{\mathbf{E}[(X-\mathbf{E}[X])^2]}{t^2}&lt;br /&gt;
=\frac{\mathbf{Var}[X]}{t^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Limited_independence&amp;diff=13905</id>
		<title>高级算法 (Fall 2026)/Limited independence</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Limited_independence&amp;diff=13905"/>
		<updated>2026-09-07T10:13:26Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;= &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-wise  independence = Recall the definition of independence between events: {{Theorem |Definition (Independent events)| :Events &amp;lt;math&amp;gt;\mathcal{E}_1, \mathcal{E}_2, \ldots, \mathcal{E}_n&amp;lt;/math&amp;gt; are &amp;#039;&amp;#039;&amp;#039;mutually independent&amp;#039;&amp;#039;&amp;#039; if, for any subset &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, ::&amp;lt;math&amp;gt;\begin{align} \Pr\left[\bigwedge_{i\in I}\mathcal{E}_i\right] &amp;amp;= \prod_{i\in I}\Pr[\mathcal{E}_i]. \end{align}&amp;lt;/math&amp;gt; }} Similarly, we can define independence between...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-wise  independence =&lt;br /&gt;
Recall the definition of independence between events:&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (Independent events)|&lt;br /&gt;
:Events &amp;lt;math&amp;gt;\mathcal{E}_1, \mathcal{E}_2, \ldots, \mathcal{E}_n&amp;lt;/math&amp;gt; are &#039;&#039;&#039;mutually independent&#039;&#039;&#039; if, for any subset &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[\bigwedge_{i\in I}\mathcal{E}_i\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i\in I}\Pr[\mathcal{E}_i].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
Similarly, we can define independence between random variables:&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (Independent variables)|&lt;br /&gt;
:Random variables &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are &#039;&#039;&#039;mutually independent&#039;&#039;&#039; if, for any subset &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; and any values &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;i\in I&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[\bigwedge_{i\in I}(X_i=x_i)\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i\in I}\Pr[X_i=x_i].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Mutual independence is an ideal condition of independence. The limited notion of independence is usually defined by the &#039;&#039;&#039;k-wise independence&#039;&#039;&#039;.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (k-wise Independenc)|&lt;br /&gt;
:1. Events &amp;lt;math&amp;gt;\mathcal{E}_1, \mathcal{E}_2, \ldots, \mathcal{E}_n&amp;lt;/math&amp;gt; are &#039;&#039;&#039;k-wise independent&#039;&#039;&#039; if, for any subset &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;  with &amp;lt;math&amp;gt;|I|\le k&amp;lt;/math&amp;gt;&lt;br /&gt;
:::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[\bigwedge_{i\in I}\mathcal{E}_i\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i\in I}\Pr[\mathcal{E}_i].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
:2. Random variables &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are &#039;&#039;&#039;k-wise independent&#039;&#039;&#039; if, for any subset &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|I|\le k&amp;lt;/math&amp;gt; and any values &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;i\in I&amp;lt;/math&amp;gt;,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[\bigwedge_{i\in I}(X_i=x_i)\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i\in I}\Pr[X_i=x_i].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
A very common case is pairwise independence, i.e. the 2-wise independence.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (pairwise Independent random variables)|&lt;br /&gt;
:Random variables &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are &#039;&#039;&#039;pairwise independent&#039;&#039;&#039; if, for any &amp;lt;math&amp;gt;X_i,X_j&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;i\neq j&amp;lt;/math&amp;gt; and any values &amp;lt;math&amp;gt;a,b&amp;lt;/math&amp;gt;&lt;br /&gt;
:::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[X_i=a\wedge X_j=b\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr[X_i=a]\cdot\Pr[X_j=b].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Note that the definition of k-wise independence is hereditary:&lt;br /&gt;
* If &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are k-wise independent, then they are also &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;-wise independent for any &amp;lt;math&amp;gt;\ell&amp;lt;k&amp;lt;/math&amp;gt;.&lt;br /&gt;
* If &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are NOT k-wise independent, then they cannot be &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;-wise independent for any &amp;lt;math&amp;gt;\ell&amp;gt;k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Pairwise Independent Bits ==&lt;br /&gt;
Suppose we have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; mutually independent and uniform random bits &amp;lt;math&amp;gt;X_1,\ldots, X_m&amp;lt;/math&amp;gt;. We are going to extract &amp;lt;math&amp;gt;n=2^m-1&amp;lt;/math&amp;gt; pairwise independent bits from these &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; mutually independent bits.&lt;br /&gt;
&lt;br /&gt;
Enumerate all the nonempty subsets of &amp;lt;math&amp;gt;\{1,2,\ldots,m\}&amp;lt;/math&amp;gt; in some order. Let &amp;lt;math&amp;gt;S_j&amp;lt;/math&amp;gt;  be the &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;th subset. Let&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y_j=\bigoplus_{i\in S_j} X_i,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\oplus&amp;lt;/math&amp;gt; is the exclusive-or, whose truth table is as follows.&lt;br /&gt;
:{|cellpadding=&amp;quot;4&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
|&amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;&lt;br /&gt;
|&amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;&lt;br /&gt;
|&amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt;&amp;lt;math&amp;gt;\oplus&amp;lt;/math&amp;gt;&amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 0 || 0 ||align=&amp;quot;center&amp;quot;| 0&lt;br /&gt;
|-&lt;br /&gt;
| 0 || 1 ||align=&amp;quot;center&amp;quot;| 1&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 0 ||align=&amp;quot;center&amp;quot;| 1&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 ||align=&amp;quot;center&amp;quot;| 0&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
There are &amp;lt;math&amp;gt;n=2^m-1&amp;lt;/math&amp;gt; such &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt;, because there are &amp;lt;math&amp;gt;2^m-1&amp;lt;/math&amp;gt; nonempty subsets of &amp;lt;math&amp;gt;\{1,2,\ldots,m\}&amp;lt;/math&amp;gt;. An equivalent definition of &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;Y_j=\left(\sum_{i\in S_j}X_i\right)\bmod 2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Sometimes, &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt; is called the &#039;&#039;&#039;parity&#039;&#039;&#039; of the bits in &amp;lt;math&amp;gt;S_j&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We claim that &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt; are pairwise independent and uniform.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt; and any &amp;lt;math&amp;gt;b\in\{0,1\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[Y_j=b\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1}{2}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
:For any &amp;lt;math&amp;gt;Y_j,Y_\ell&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;j\neq\ell&amp;lt;/math&amp;gt; and any &amp;lt;math&amp;gt;a,b\in\{0,1\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[Y_j=a\wedge Y_\ell=b\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1}{4}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The proof is left for your exercise.&lt;br /&gt;
&lt;br /&gt;
Therefore, we extract exponentially many pairwise independent uniform random bits from a sequence of mutually independent uniform random bits.&lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt; are not 3-wise independent. For example, consider the subsets &amp;lt;math&amp;gt;S_1=\{1\},S_2=\{2\},S_3=\{1,2\}&amp;lt;/math&amp;gt; and the corresponding random bits &amp;lt;math&amp;gt;Y_1,Y_2,Y_3&amp;lt;/math&amp;gt;. Any two of &amp;lt;math&amp;gt;Y_1,Y_2,Y_3&amp;lt;/math&amp;gt; would decide the value of the third one.&lt;br /&gt;
&lt;br /&gt;
==  Pairwise Independent Variables ==&lt;br /&gt;
We now consider constructing pairwise independent random variables ranging over &amp;lt;math&amp;gt;[p]=\{0,1,2,\ldots,p-1\}&amp;lt;/math&amp;gt; for some prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. Unlike the above construction, now we only need two independent random sources &amp;lt;math&amp;gt;X_0,X_1&amp;lt;/math&amp;gt;, which are uniformly and independently distributed over &amp;lt;math&amp;gt;[p]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y_0,Y_1,\ldots, Y_{p-1}&amp;lt;/math&amp;gt; be defined as:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
Y_i=(X_0+i\cdot X_1)\bmod p &amp;amp;\quad \mbox{for }i\in[p].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
: The random variables &amp;lt;math&amp;gt;Y_0,Y_1,\ldots, Y_{p-1}&amp;lt;/math&amp;gt; are pairwise independent uniform random variables over &amp;lt;math&amp;gt;[p]&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| We first show that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; are uniform. That is, we will show that for any &amp;lt;math&amp;gt;i,a\in[p]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[(X_0+i\cdot X_1)\bmod p=a\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1}{p}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
Due to the law of total probability,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[(X_0+i\cdot X_1)\bmod p=a\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{j\in[p]}\Pr[X_1=j]\cdot\Pr\left[(X_0+ij)\bmod p=a\right]\\&lt;br /&gt;
&amp;amp;=\frac{1}{p}\sum_{j\in[p]}\Pr\left[X_0\equiv(a-ij)\pmod{p}\right].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
For prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;i,j,a\in[p]&amp;lt;/math&amp;gt;, there is exact one value in &amp;lt;math&amp;gt;[p]&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;X_0&amp;lt;/math&amp;gt; satisfying &amp;lt;math&amp;gt;X_0\equiv(a-ij)\pmod{p}&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;\Pr\left[X_0\equiv(a-ij)\pmod{p}\right]=1/p&amp;lt;/math&amp;gt; and the above probability is &amp;lt;math&amp;gt;\frac{1}{p}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We then show that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; are pairwise independent, i.e. we will show that for any &amp;lt;math&amp;gt;Y_i,Y_j&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;i\neq j&amp;lt;/math&amp;gt; and any &amp;lt;math&amp;gt;a,b\in[p]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr\left[Y_i=a\wedge Y_j=b\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1}{p^2}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The event &amp;lt;math&amp;gt;Y_i=a\wedge Y_j=b&amp;lt;/math&amp;gt; is equivalent to that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{cases}&lt;br /&gt;
(X_0+iX_1)\equiv a\pmod{p}\\&lt;br /&gt;
(X_0+jX_1)\equiv b\pmod{p}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Due to the [http://en.wikipedia.org/wiki/Chinese_remainder_theorem Chinese remainder theorem], there exists a unique solution of &amp;lt;math&amp;gt;X_0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;X_1&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;[p]&amp;lt;/math&amp;gt; to the above linear congruential system. Thus the probability of the event is &amp;lt;math&amp;gt;\frac{1}{p^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=Universal Hashing =&lt;br /&gt;
Hashing is one of the oldest tools in Computer Science. Knuth&#039;s memorandum in 1963 on analysis of hash tables is now considered to be the birth of the area of analysis of algorithms.&lt;br /&gt;
* Knuth. Notes on &amp;quot;open&amp;quot; addressing, July 22 1963. Unpublished memorandum.&lt;br /&gt;
&lt;br /&gt;
The idea of hashing is simple: an unknown set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; data &#039;&#039;&#039;items&#039;&#039;&#039; (or keys) are drawn from a large &#039;&#039;&#039;universe&#039;&#039;&#039; &amp;lt;math&amp;gt;U=[N]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;N\gg n&amp;lt;/math&amp;gt;; in order to store &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; in a table of &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; entries (slots), we assume a consistent mapping (called a &#039;&#039;&#039;hash function&#039;&#039;&#039;) from the universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; to a small range &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This idea seems clever: we use a consistent mapping to deal with an arbitrary unknown data set. However, there is a fundamental flaw for hashing.&lt;br /&gt;
* For sufficiently large universe (&amp;lt;math&amp;gt;N&amp;gt; M(n-1)&amp;lt;/math&amp;gt;), for any function, there exists a bad data set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, such that all items in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; are mapped to the same entry in the table.&lt;br /&gt;
&lt;br /&gt;
A simple use of pigeonhole principle can prove the above statement. &lt;br /&gt;
&lt;br /&gt;
To overcome this situation, randomization is introduced into hashing. We assume that the hash function is a random mapping from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;. In order to ease the analysis, the following ideal assumption is used:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Simple Uniform Hash Assumption&#039;&#039;&#039; (&#039;&#039;&#039;SUHA&#039;&#039;&#039; or &#039;&#039;&#039;UHA&#039;&#039;&#039;, a.k.a. the random oracle model): &lt;br /&gt;
:A &#039;&#039;uniform&#039;&#039; random function &amp;lt;math&amp;gt;h:[N]\rightarrow[M]&amp;lt;/math&amp;gt; is available and the computation of &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is efficient.&lt;br /&gt;
&lt;br /&gt;
== Families of universal hash functions ==&lt;br /&gt;
The assumption of completely random function simplifies the analysis.  However, in practice, truly uniform random hash function is extremely expensive to compute and store. Thus, this simple assumption can hardly represent the reality.&lt;br /&gt;
&lt;br /&gt;
There are two approaches for implementing practical hash functions. One is to use &#039;&#039;ad hoc&#039;&#039; implementations and wish they may work. The other approach is to construct class of hash functions which are efficient to compute and store but with weaker randomness guarantees, and then analyze the applications of hash functions based on this weaker assumption of randomness.&lt;br /&gt;
&lt;br /&gt;
This route was took by Carter and Wegman in 1977 while they introduced universal families of hash functions.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (universal hash families)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; be a universe with &amp;lt;math&amp;gt;N\ge M&amp;lt;/math&amp;gt;. A family of hash functions &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; is said to be &#039;&#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-universal&#039;&#039;&#039; if, for any items &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_k\in [N]&amp;lt;/math&amp;gt; and for a hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; chosen uniformly at random from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, we have &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)=\cdots=h(x_k)]\le\frac{1}{M^{k-1}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
:A family of hash functions &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; is said to be &#039;&#039;&#039;strongly &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-universal&#039;&#039;&#039; if, for any items &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_k\in [N]&amp;lt;/math&amp;gt;, any values &amp;lt;math&amp;gt;y_1,y_2,\ldots,y_k\in[M]&amp;lt;/math&amp;gt;, and for a hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; chosen uniformly at random from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, we have &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=y_1\wedge h(x_2)=y_2 \wedge \cdots \wedge h(x_k)=y_k]=\frac{1}{M^{k}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
In particular, for a 2-universal family &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, for any elements &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt;, a uniform random &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt; has&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)]\le\frac{1}{M}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For a strongly 2-universal family &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, for any elements &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; and any values &amp;lt;math&amp;gt;y_1,y_2\in[M]&amp;lt;/math&amp;gt;, a uniform random &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt; has&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=y_1\wedge h(x_2)=y_2]=\frac{1}{M^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This behavior is exactly the same as uniform random hash functions on any pair of inputs. For this reason, a strongly 2-universal hash family are also called pairwise independent hash functions.&lt;br /&gt;
&lt;br /&gt;
== 2-universal hash families ==&lt;br /&gt;
&lt;br /&gt;
The construction of pairwise independent random variables via modulo a prime introduced in Section 1 already provides a way of constructing a strongly 2-universal hash family.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; be a prime. The function &amp;lt;math&amp;gt;h_{a,b}:[p]\rightarrow [p]&amp;lt;/math&amp;gt; is defined by&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
h_{a,b}(x)=(ax+b)\bmod p,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and the family is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{H}=\{h_{a,b}\mid a,b\in[p]\}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lemma|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is strongly 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| In Section 1, we have proved the pairwise independence of the sequence of &amp;lt;math&amp;gt;(a i+b)\bmod p&amp;lt;/math&amp;gt;, for &amp;lt;math&amp;gt;i=0,1,\ldots, p-1&amp;lt;/math&amp;gt;, which directly implies that &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is strongly 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;The original construction of Carter-Wegman&lt;br /&gt;
What if we want to have hash functions from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; for non-prime &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt;? Carter and Wegman developed the following method.&lt;br /&gt;
&lt;br /&gt;
Suppose that the universe is &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt;, and the functions map &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;N\ge M&amp;lt;/math&amp;gt;. For some prime &amp;lt;math&amp;gt;p\ge N&amp;lt;/math&amp;gt;, let&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
h_{a,b}(x)=((ax+b)\bmod p)\bmod M,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and the family&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{H}=\{h_{a,b}\mid 1\le a\le p-1, b\in[p]\}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that unlike the first construction, now &amp;lt;math&amp;gt;a\neq 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lemma (Carter-Wegman)|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Due to the definition of &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;p(p-1)&amp;lt;/math&amp;gt; many different hash functions in &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, because each hash function in &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; corresponds to a pair of &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b\in[p]&amp;lt;/math&amp;gt;. We only need to count for any particular pair of &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, the number of hash functions that &amp;lt;math&amp;gt;h(x_1)=h(x_2)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We first note that for any &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;a x_1+b\not\equiv a x_2+b \pmod p&amp;lt;/math&amp;gt;. This is because &amp;lt;math&amp;gt;a x_1+b\equiv a x_2+b \pmod p&amp;lt;/math&amp;gt; would imply that &amp;lt;math&amp;gt;a(x_1-x_2)\equiv 0\pmod p&amp;lt;/math&amp;gt;, which can never happen since &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt; (note that &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; for an &amp;lt;math&amp;gt;N\le p&amp;lt;/math&amp;gt;).  Therefore, we can assume that &amp;lt;math&amp;gt;(a x_1+b)\bmod p=u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(a x_2+b)\bmod p=v&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;u\neq v&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By linear algebra (over finite field), for any &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;u,v\in[p]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;u\neq v&amp;lt;/math&amp;gt;, there is exact one solution to &amp;lt;math&amp;gt;(a,b)&amp;lt;/math&amp;gt; satisfying:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{cases}&lt;br /&gt;
a x_1+b \equiv u \pmod p\\&lt;br /&gt;
a x_2+b \equiv v \pmod p.&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
After modulo &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt;, every &amp;lt;math&amp;gt;u\in[p]&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;\lceil p/M\rceil -1&amp;lt;/math&amp;gt; many &amp;lt;math&amp;gt;v\in[p]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;v\neq u&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;v\equiv u\pmod M&amp;lt;/math&amp;gt;. Therefore, for every pair of &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, there exist at most &amp;lt;math&amp;gt;p(\lceil p/M\rceil -1)\le p(p-1)/M&amp;lt;/math&amp;gt; pairs of &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b\in[p]&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;((ax_1+b)\bmod p)\bmod M=((ax_2+b)\bmod p)\bmod M&amp;lt;/math&amp;gt;, which means there are at most &amp;lt;math&amp;gt; p(p-1)/M&amp;lt;/math&amp;gt; many hash functions &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt; having &amp;lt;math&amp;gt;h(x_1)=h(x_2)&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;. For &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; uniformly chosen from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)]\le \frac{p(p-1)/M}{p(p-1)}=\frac{1}{M}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
We prove that &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;A construction used in practice&lt;br /&gt;
The main issue of Carter-Wegman construction is the efficiency. The mod operation is very slow, and has been so for more than 30 years.&lt;br /&gt;
&lt;br /&gt;
The following construction is due to Dietzfelbinger &#039;&#039;et al&#039;&#039;. It was published in 1997 and has been practically used in various applications of universal hashing.&lt;br /&gt;
&lt;br /&gt;
The family of hash functions is from &amp;lt;math&amp;gt;[2^u]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[2^v]&amp;lt;/math&amp;gt;. With a binary representation, the functions map binary strings of length &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; to binary strings of length &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
h_{a}(x)=\left\lfloor\frac{a\cdot x\bmod 2^u}{2^{u-v}}\right\rfloor,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and the family&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{H}=\{h_{a}\mid a\in[2^v]\mbox{ and }a\mbox{ is odd}\}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This family of hash functions does not exactly meet the requirement of 2-universal family. However,  Dietzfelbinger &#039;&#039;et al&#039;&#039; proved that &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is close to a 2-universal family. Specifically, for any input values &amp;lt;math&amp;gt;x_1,x_2\in[2^u]&amp;lt;/math&amp;gt;, for a uniformly random &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)]\le\frac{1}{2^{v-1}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is within an approximation ratio of 2 to being 2-universal. The proof uses the fact that odd numbers are relative prime to a power of 2.&lt;br /&gt;
&lt;br /&gt;
The function is extremely simple to compute in c language.&lt;br /&gt;
We exploit that C-multiplication (*) of unsigned u-bit numbers is done &amp;lt;math&amp;gt;\bmod 2^u&amp;lt;/math&amp;gt;, and have a one-line C-code for computing the hash function:&lt;br /&gt;
 h_a(x) = (a*x)&amp;gt;&amp;gt;(u-v)&lt;br /&gt;
The bit-wise shifting is a lot faster than modular. It explains the popularity of this scheme in practice than the original Carter-Wegman construction.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Hashing_and_Sketching&amp;diff=13904</id>
		<title>高级算法 (Fall 2026)/Hashing and Sketching</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Hashing_and_Sketching&amp;diff=13904"/>
		<updated>2026-09-07T10:12:58Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;=Balls into Bins= The following is the so-called balls into bins model. Consider throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls into &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins uniformly and independently at random. This is equivalent to a random mapping &amp;lt;math&amp;gt;f:[m]\to[n]&amp;lt;/math&amp;gt;. Needless to say, random mapping is an important random model and may have many applications in Computer Science, e.g. hashing.  We are concerned with the following three questions regarding the balls into bins model: * birthday problem...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Balls into Bins=&lt;br /&gt;
The following is the so-called balls into bins model.&lt;br /&gt;
Consider throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls into &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins uniformly and independently at random. This is equivalent to a random mapping &amp;lt;math&amp;gt;f:[m]\to[n]&amp;lt;/math&amp;gt;. Needless to say, random mapping is an important random model and may have many applications in Computer Science, e.g. hashing.&lt;br /&gt;
&lt;br /&gt;
We are concerned with the following three questions regarding the balls into bins model:&lt;br /&gt;
* birthday problem: the probability that every bin contains at most one ball (the mapping is 1-1);&lt;br /&gt;
* coupon collector problem: the probability that every bin contains at least one ball (the mapping is on-to);&lt;br /&gt;
* occupancy problem: the maximum load of bins.&lt;br /&gt;
&lt;br /&gt;
== Birthday Problem==&lt;br /&gt;
We now consider the &#039;&#039;&#039;birthday problem&#039;&#039;&#039;.&lt;br /&gt;
There are &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; students in the class. Assume that for each student, his/her birthday is uniformly and independently distributed over the 365 days in a years. We wonder what the probability that no two students share a birthday.&lt;br /&gt;
&lt;br /&gt;
Due to the [http://en.wikipedia.org/wiki/Pigeonhole_principle pigeonhole principle], it is obvious that for &amp;lt;math&amp;gt;m&amp;gt;365&amp;lt;/math&amp;gt;, there must be two students with the same birthday. Surprisingly, for any &amp;lt;math&amp;gt;m&amp;gt;57&amp;lt;/math&amp;gt; this event occurs with more than 99% probability. This is called the [http://en.wikipedia.org/wiki/Birthday_problem &#039;&#039;&#039;birthday paradox&#039;&#039;&#039;]. Despite the name, the birthday paradox is not a real paradox.&lt;br /&gt;
&lt;br /&gt;
We can model this problem as a balls-into-bins problem. &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; different balls (students) are uniformly and independently thrown into 365 bins (days). More generally, let &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; be the number of bins. We ask for the probability of the following event &amp;lt;math&amp;gt;\mathcal{E}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}&amp;lt;/math&amp;gt;: there is no bin with more than one balls (i.e. no two students share birthday).&lt;br /&gt;
&lt;br /&gt;
We first analyze this by counting. There are totally &amp;lt;math&amp;gt;n^m&amp;lt;/math&amp;gt; ways of assigning &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins. The number of assignments that no two balls share a bin is &amp;lt;math&amp;gt;{n\choose m}m!&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Thus the probability is given by:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}]&lt;br /&gt;
=&lt;br /&gt;
\frac{{n\choose m}m!}{n^m}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;{n\choose m}=\frac{n!}{(n-m)!m!}&amp;lt;/math&amp;gt;. Then &lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}]&lt;br /&gt;
=&lt;br /&gt;
\frac{{n\choose m}m!}{n^m}&lt;br /&gt;
=&lt;br /&gt;
\frac{n!}{n^m(n-m)!}&lt;br /&gt;
=&lt;br /&gt;
\frac{n}{n}\cdot\frac{n-1}{n}\cdot\frac{n-2}{n}\cdots\frac{n-(m-1)}{n}&lt;br /&gt;
=&lt;br /&gt;
\prod_{k=1}^{m-1}\left(1-\frac{k}{n}\right).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
There is also a more &amp;quot;probabilistic&amp;quot; argument for the above equation. Consider again that &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; students are mapped to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; possible birthdays uniformly at random.&lt;br /&gt;
&lt;br /&gt;
The first student has a birthday for sure. The probability that the second student has a different birthday from the first student is &amp;lt;math&amp;gt;\left(1-\frac{1}{n}\right)&amp;lt;/math&amp;gt;. Given that the first two students have different birthdays, the probability that the third student has a different birthday from the first two students is &amp;lt;math&amp;gt;\left(1-\frac{2}{n}\right)&amp;lt;/math&amp;gt;. Continuing this on, assuming that the first &amp;lt;math&amp;gt;k-1&amp;lt;/math&amp;gt; students all have different birthdays, the probability that the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;th student has a different birthday than the first &amp;lt;math&amp;gt;k-1&amp;lt;/math&amp;gt;, is given by &amp;lt;math&amp;gt;\left(1-\frac{k-1}{n}\right)&amp;lt;/math&amp;gt;. By the chain rule, the probability that all &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; students have different birthdays is:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}]=\left(1-\frac{1}{n}\right)\cdot \left(1-\frac{2}{n}\right)\cdots \left(1-\frac{m-1}{n}\right)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{k=1}^{m-1}\left(1-\frac{k}{n}\right),&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is the same as what we got by the counting argument.&lt;br /&gt;
&lt;br /&gt;
[[File:Birthday.png|border|450px|right]]&lt;br /&gt;
&lt;br /&gt;
There are several ways of analyzing this formular. Here is a convenient one: Due to [http://en.wikipedia.org/wiki/Taylor_series Taylor&#039;s expansion], &amp;lt;math&amp;gt;e^{-k/n}\approx 1-k/n&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\prod_{k=1}^{m-1}\left(1-\frac{k}{n}\right)&lt;br /&gt;
&amp;amp;\approx&lt;br /&gt;
\prod_{k=1}^{m-1}e^{-\frac{k}{n}}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\sum_{k=1}^{m-1}\frac{k}{n}\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
e^{-m(m-1)/2n}\\&lt;br /&gt;
&amp;amp;\approx&lt;br /&gt;
e^{-m^2/2n}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
The quality of this approximation is shown in the Figure.&lt;br /&gt;
&lt;br /&gt;
Therefore, for &amp;lt;math&amp;gt;m=\sqrt{2n\ln \frac{1}{\epsilon}}&amp;lt;/math&amp;gt;, the probability that &amp;lt;math&amp;gt;\Pr[\mathcal{E}]\approx\epsilon&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==Universal Hashing ==&lt;br /&gt;
Hashing is one of the oldest tools in Computer Science. Knuth&#039;s memorandum in 1963 on analysis of hash tables is now considered to be the birth of the area of analysis of algorithms.&lt;br /&gt;
* Knuth. Notes on &amp;quot;open&amp;quot; addressing, July 22 1963. Unpublished memorandum.&lt;br /&gt;
&lt;br /&gt;
The idea of hashing is simple: an unknown set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; data &#039;&#039;&#039;items&#039;&#039;&#039; (or keys) are drawn from a large &#039;&#039;&#039;universe&#039;&#039;&#039; &amp;lt;math&amp;gt;U=[N]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;N\gg n&amp;lt;/math&amp;gt;; in order to store &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; in a table of &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; entries (slots), we assume a consistent mapping (called a &#039;&#039;&#039;hash function&#039;&#039;&#039;) from the universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; to a small range &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This idea seems clever: we use a consistent mapping to deal with an arbitrary unknown data set. However, there is a fundamental flaw for hashing.&lt;br /&gt;
* For sufficiently large universe (&amp;lt;math&amp;gt;N&amp;gt; M(n-1)&amp;lt;/math&amp;gt;), for any function, there exists a bad data set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, such that all items in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; are mapped to the same entry in the table.&lt;br /&gt;
&lt;br /&gt;
A simple use of pigeonhole principle can prove the above statement. &lt;br /&gt;
&lt;br /&gt;
To overcome this situation, randomization is introduced into hashing. We assume that the hash function is a random mapping from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;. In order to ease the analysis, the following ideal assumption is used:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Simple Uniform Hash Assumption&#039;&#039;&#039; (&#039;&#039;&#039;SUHA&#039;&#039;&#039; or &#039;&#039;&#039;UHA&#039;&#039;&#039;, a.k.a. the random oracle model): &lt;br /&gt;
:A &#039;&#039;uniform&#039;&#039; random function &amp;lt;math&amp;gt;h:[N]\rightarrow[M]&amp;lt;/math&amp;gt; is available and the computation of &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is efficient.&lt;br /&gt;
&lt;br /&gt;
=== Families of universal hash functions ===&lt;br /&gt;
The assumption of completely random function simplifies the analysis.  However, in practice, truly uniform random hash function is extremely expensive to compute and store. Thus, this simple assumption can hardly represent the reality.&lt;br /&gt;
&lt;br /&gt;
There are two approaches for implementing practical hash functions. One is to use &#039;&#039;ad hoc&#039;&#039; implementations and wish they may work. The other approach is to construct class of hash functions which are efficient to compute and store but with weaker randomness guarantees, and then analyze the applications of hash functions based on this weaker assumption of randomness.&lt;br /&gt;
&lt;br /&gt;
This route was took by Carter and Wegman in 1977 while they introduced universal families of hash functions.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (universal hash families)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; be a universe with &amp;lt;math&amp;gt;N\ge M&amp;lt;/math&amp;gt;. A family of hash functions &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; is said to be &#039;&#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-universal&#039;&#039;&#039; if, for any distinct items &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_k\in [N]&amp;lt;/math&amp;gt; and for a hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; chosen uniformly at random from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, we have &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)=\cdots=h(x_k)]\le\frac{1}{M^{k-1}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
:A family of hash functions &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; is said to be &#039;&#039;&#039;strongly &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-universal&#039;&#039;&#039; if, for any distinct items &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_k\in [N]&amp;lt;/math&amp;gt;, any values &amp;lt;math&amp;gt;y_1,y_2,\ldots,y_k\in[M]&amp;lt;/math&amp;gt;, and for a hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; chosen uniformly at random from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, we have &lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=y_1\wedge h(x_2)=y_2 \wedge \cdots \wedge h(x_k)=y_k]=\frac{1}{M^{k}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
In particular, for a 2-universal family &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, for any distinct elements &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt;, a uniform random &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt; has&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)]\le\frac{1}{M}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For a strongly 2-universal family &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, for any distinct elements &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; and any values &amp;lt;math&amp;gt;y_1,y_2\in[M]&amp;lt;/math&amp;gt;, a uniform random &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt; has&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=y_1\wedge h(x_2)=y_2]=\frac{1}{M^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This behavior is exactly the same as uniform random hash functions on any distinct pair of inputs. For this reason, a strongly 2-universal hash family are also called pairwise independent hash functions.&lt;br /&gt;
&lt;br /&gt;
=== 2-universal hash families ===&lt;br /&gt;
&lt;br /&gt;
The construction of pairwise independent random variables via modulo a prime introduced in Section 1 already provides a way of constructing a strongly 2-universal hash family.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; be a prime. The function &amp;lt;math&amp;gt;h_{a,b}:[p]\rightarrow [p]&amp;lt;/math&amp;gt; is defined by&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
h_{a,b}(x)=(ax+b)\bmod p,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and the family is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{H}=\{h_{a,b}\mid a,b\in[p]\}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lemma|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is strongly 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| In Section 1, we have proved the pairwise independence of the sequence of &amp;lt;math&amp;gt;(a i+b)\bmod p&amp;lt;/math&amp;gt;, for &amp;lt;math&amp;gt;i=0,1,\ldots, p-1&amp;lt;/math&amp;gt;, which directly implies that &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is strongly 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;The original construction of Carter-Wegman&lt;br /&gt;
What if we want to have hash functions from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; for non-prime &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt;? Carter and Wegman developed the following method.&lt;br /&gt;
&lt;br /&gt;
Suppose that the universe is &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt;, and the functions map &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;N\ge M&amp;lt;/math&amp;gt;. For some prime &amp;lt;math&amp;gt;p\ge N&amp;lt;/math&amp;gt;, let&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
h_{a,b}(x)=((ax+b)\bmod p)\bmod M,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and the family&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{H}=\{h_{a,b}\mid 1\le a\le p-1, b\in[p]\}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that unlike the first construction, now &amp;lt;math&amp;gt;a\neq 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lemma (Carter-Wegman)|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Due to the definition of &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;p(p-1)&amp;lt;/math&amp;gt; many different hash functions in &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, because each hash function in &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; corresponds to a pair of &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b\in[p]&amp;lt;/math&amp;gt;. We only need to count for any particular pair of &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, the number of hash functions that &amp;lt;math&amp;gt;h(x_1)=h(x_2)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We first note that for any &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;a x_1+b\not\equiv a x_2+b \pmod p&amp;lt;/math&amp;gt;. This is because &amp;lt;math&amp;gt;a x_1+b\equiv a x_2+b \pmod p&amp;lt;/math&amp;gt; would imply that &amp;lt;math&amp;gt;a(x_1-x_2)\equiv 0\pmod p&amp;lt;/math&amp;gt;, which can never happen since &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt; (note that &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; for an &amp;lt;math&amp;gt;N\le p&amp;lt;/math&amp;gt;).  Therefore, we can assume that &amp;lt;math&amp;gt;(a x_1+b)\bmod p=u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(a x_2+b)\bmod p=v&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;u\neq v&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By linear algebra (over finite field), for any &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;u,v\in[p]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;u\neq v&amp;lt;/math&amp;gt;, there is exact one solution to &amp;lt;math&amp;gt;(a,b)&amp;lt;/math&amp;gt; satisfying:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{cases}&lt;br /&gt;
a x_1+b \equiv u \pmod p\\&lt;br /&gt;
a x_2+b \equiv v \pmod p.&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
After modulo &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt;, every &amp;lt;math&amp;gt;u\in[p]&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;\lceil p/M\rceil -1&amp;lt;/math&amp;gt; many &amp;lt;math&amp;gt;v\in[p]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;v\neq u&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;v\equiv u\pmod M&amp;lt;/math&amp;gt;. Therefore, for every pair of &amp;lt;math&amp;gt;x_1,x_2\in[N]&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;, there exist at most &amp;lt;math&amp;gt;p(\lceil p/M\rceil -1)\le p(p-1)/M&amp;lt;/math&amp;gt; pairs of &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b\in[p]&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;((ax_1+b)\bmod p)\bmod M=((ax_2+b)\bmod p)\bmod M&amp;lt;/math&amp;gt;, which means there are at most &amp;lt;math&amp;gt; p(p-1)/M&amp;lt;/math&amp;gt; many hash functions &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt; having &amp;lt;math&amp;gt;h(x_1)=h(x_2)&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;. For &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; uniformly chosen from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;x_1\neq x_2&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)]\le \frac{p(p-1)/M}{p(p-1)}=\frac{1}{M}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
We prove that &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is 2-universal.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;A construction used in practice&lt;br /&gt;
The main issue of Carter-Wegman construction is the efficiency. The mod operation is very slow, and has been so for more than 30 years.&lt;br /&gt;
&lt;br /&gt;
The following construction is due to Dietzfelbinger &#039;&#039;et al&#039;&#039;. It was published in 1997 and has been practically used in various applications of universal hashing.&lt;br /&gt;
&lt;br /&gt;
The family of hash functions is from &amp;lt;math&amp;gt;[2^u]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[2^v]&amp;lt;/math&amp;gt;. With a binary representation, the functions map binary strings of length &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; to binary strings of length &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
h_{a}(x)=\left\lfloor\frac{a\cdot x\bmod 2^u}{2^{u-v}}\right\rfloor,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and the family&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{H}=\{h_{a}\mid a\in[2^v]\mbox{ and }a\mbox{ is odd}\}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This family of hash functions does not exactly meet the requirement of 2-universal family. However,  Dietzfelbinger &#039;&#039;et al&#039;&#039; proved that &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is close to a 2-universal family. Specifically, for any distinct input values &amp;lt;math&amp;gt;x_1,x_2\in[2^u]&amp;lt;/math&amp;gt;, for a uniformly random &amp;lt;math&amp;gt;h\in\mathcal{H}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h(x_1)=h(x_2)]\le\frac{1}{2^{v-1}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is within an approximation ratio of 2 to being 2-universal. The proof uses the fact that odd numbers are relative prime to a power of 2.&lt;br /&gt;
&lt;br /&gt;
The function is extremely simple to compute in c language.&lt;br /&gt;
We exploit that C-multiplication (*) of unsigned u-bit numbers is done &amp;lt;math&amp;gt;\bmod 2^u&amp;lt;/math&amp;gt;, and have a one-line C-code for computing the hash function:&lt;br /&gt;
 h_a(x) = (a*x)&amp;gt;&amp;gt;(u-v)&lt;br /&gt;
The bit-wise shifting is a lot faster than modular. It explains the popularity of this scheme in practice than the original Carter-Wegman construction.&lt;br /&gt;
&lt;br /&gt;
== Collision number ==&lt;br /&gt;
Consider a 2-universal family &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; of hash functions from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; be a hash function chosen uniformly from &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;. For a fixed set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct elements from &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt;, say &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt;, the elements are mapped to the hash values &amp;lt;math&amp;gt;h(x_1), h(x_2), \ldots, h(x_n)&amp;lt;/math&amp;gt;. This can be seen as throwing &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; balls to &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; bins, with pairwise independent choices of bins.&lt;br /&gt;
&lt;br /&gt;
As in the balls-into-bins with full independence, we are curious about the questions such as the birthday problem or the maximum load. These questions are interesting not only because they are natural to ask in a balls-into-bins setting, but in the context of hashing, they are closely related to the performance of hash functions.&lt;br /&gt;
&lt;br /&gt;
The old techniques for analyzing balls-into-bins rely too much on the independence of the choice of the bin for each ball, therefore can hardly be extended to the setting of 2-universal hash families. However, it turns out several balls-into-bins questions can somehow be answered by analyzing a very natural quantity: the number of &#039;&#039;&#039;collision pairs&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
A collision pair for hashing is a pair of elements &amp;lt;math&amp;gt;x_1,x_2\in S&amp;lt;/math&amp;gt; which are mapped to the same hash value, i.e. &amp;lt;math&amp;gt;h(x_1)=h(x_2)&amp;lt;/math&amp;gt;. Formally, for a fixed set of elements &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;1\le i,j\le n&amp;lt;/math&amp;gt;, let the random variable&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X_{ij}&lt;br /&gt;
=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1 &amp;amp; \text{if }h(x_i)=h(x_j),\\&lt;br /&gt;
0 &amp;amp; \text{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The total number of collision pairs among the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;X=\sum_{i&amp;lt;j} X_{ij}.\,&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Since &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt; is 2-universal, for any &amp;lt;math&amp;gt;i\neq j&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[X_{ij}=1]=\Pr[h(x_i)=h(x_j)]\le\frac{1}{M}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The expected number of collision pairs is &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X]=\mathbf{E}\left[\sum_{i&amp;lt;j}X_{ij}\right]=\sum_{i&amp;lt;j}\mathbf{E}[X_{ij}]=\sum_{i&amp;lt;j}\Pr[X_{ij}=1]\le{n\choose 2}\frac{1}{M}&amp;lt;\frac{n^2}{2M}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
In particular, for &amp;lt;math&amp;gt;n=M&amp;lt;/math&amp;gt;, i.e. &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items are mapped to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; hash values by a pairwise independent hash function, the expected collision number is &amp;lt;math&amp;gt;\mathbf{E}[X]&amp;lt;\frac{n^2}{2M}=\frac{n}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The above analysis gives us an estimation on the expected number of collision pairs, such that &amp;lt;math&amp;gt;\mathbf{E}[X]&amp;lt;\frac{n^2}{2M}&amp;lt;/math&amp;gt;. Apply the Markov&#039;s inequality, for &amp;lt;math&amp;gt;0&amp;lt;\epsilon&amp;lt;1&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[X\ge \frac{n^2}{2\epsilon M}\right]\le\Pr\left[X\ge \frac{1}{\epsilon}\mathbf{E}[X]\right]\le\epsilon.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
When &amp;lt;math&amp;gt;n\le\sqrt{2\epsilon M}&amp;lt;/math&amp;gt;, the number of collision pairs is &amp;lt;math&amp;gt;X\ge1&amp;lt;/math&amp;gt; with probability at most &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;, therefore with probability at least &amp;lt;math&amp;gt;1-\epsilon&amp;lt;/math&amp;gt;, there is no collision at all. Therefore, we have the following theorem.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:If &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is chosen uniformly from a 2-universal family of hash functions mapping the universe &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[M]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;N\ge M&amp;lt;/math&amp;gt;, then for any set &amp;lt;math&amp;gt;S\subset [N]&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items, where &amp;lt;math&amp;gt;n\le\sqrt{2\epsilon M}&amp;lt;/math&amp;gt;, the probability that there exits a collision pair is&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\mbox{collision occurs}]\le\epsilon.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Recall that for mutually independent choices of bins, for some &amp;lt;math&amp;gt;n=\sqrt{2M\ln(1/\epsilon)}&amp;lt;/math&amp;gt;, the probability that a collision occurs is about &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;. For constant &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;, this gives an essentially same bound as the pairwise independent setting. Therefore, &lt;br /&gt;
the behavior of pairwise independent hash function is essentially the same as the uniform random hash function for the birthday problem. This is easy to understand, because birthday problem is about the behavior of collisions, and the definition of 2-universal hash function can be interpreted as &amp;quot;functions that the probability of collision is as low as a uniform random function&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
= Set  Membership=&lt;br /&gt;
A basic question in Computer Science is:&lt;br /&gt;
:&amp;quot;&amp;lt;math&amp;gt;\mbox{Is }x\in S?&amp;lt;/math&amp;gt;&amp;quot;&lt;br /&gt;
for a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; and an element &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;. This is the &#039;&#039;&#039;set membership&#039;&#039;&#039; problem.&lt;br /&gt;
&lt;br /&gt;
Formally, given an arbitrary set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements from a universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt;, we want to use a succinct &#039;&#039;&#039;data structure&#039;&#039;&#039; to represent this set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, so that upon each &#039;&#039;&#039;query&#039;&#039;&#039; of any element &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; from the universe &amp;lt;math&amp;gt;[N]&amp;lt;/math&amp;gt;, the question of whether &amp;lt;math&amp;gt;x\in S&amp;lt;/math&amp;gt; is efficiently answered. The complexity of such data structure is measured in two-fold:&lt;br /&gt;
* &#039;&#039;&#039;space cost&#039;&#039;&#039;: size of the data structure to represent a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &#039;&#039;&#039;time cost&#039;&#039;&#039;: time complexity of answering each query by accessing to the data structure.&lt;br /&gt;
&lt;br /&gt;
Suppose that the universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; is of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. Clearly, the membership problem can be solved by a &#039;&#039;&#039;dictionary data structure&#039;&#039;&#039;, e.g.:&lt;br /&gt;
* &#039;&#039;&#039;sorted table / balanced search tree&#039;&#039;&#039;: with space cost &amp;lt;math&amp;gt;O(n\log N)&amp;lt;/math&amp;gt; bits and time cost &amp;lt;math&amp;gt;O(\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;\log{N\choose n}=\Theta\left(n\log \frac{N}{n}\right)&amp;lt;/math&amp;gt; is the entropy of sets &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements from a universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;. Therefore it is necessary to use so many bits to represent a set without losing any information. &lt;br /&gt;
With hashing, we can solve this fundamental problem with asymptotic optimal space cost and time cost at the same time.&lt;br /&gt;
&lt;br /&gt;
== Perfect hashing using quadratic space==&lt;br /&gt;
The idea of perfect hashing is that we use a hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; to map the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items to distinct entries of the table; store every item &amp;lt;math&amp;gt;x\in S&amp;lt;/math&amp;gt; in the entry &amp;lt;math&amp;gt;h(x)&amp;lt;/math&amp;gt;; and also store the hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; in a fixed location in the table (usually the beginning of the table). The algorithm for searching for an item is as follows:&lt;br /&gt;
&lt;br /&gt;
:search for &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; in table &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt;:&lt;br /&gt;
# retrieve &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; from a fixed location in the table;&lt;br /&gt;
# if &amp;lt;math&amp;gt;x=T[h(x)]&amp;lt;/math&amp;gt; return &amp;lt;math&amp;gt;h(x)&amp;lt;/math&amp;gt;; else return NOT_FOUND;&lt;br /&gt;
&lt;br /&gt;
This scheme works as long as that the hash function satisfies the following two conditions:&lt;br /&gt;
* The description of &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is sufficiently short, so that &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; can be stored in one entry (or in constant many entries) of the table.&lt;br /&gt;
* &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; has no collisions on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, i.e. there is no pair of items &amp;lt;math&amp;gt;x_1,x_2\in S&amp;lt;/math&amp;gt; that are mapped to the same value by &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first condition is easy to guarantee for 2-universal hash families. As shown by Carter-Wegman construction, a 2-universal hash function can be uniquely represented by two integers &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, which can be stored in two entries (or just one, if the word length is sufficiently large) of the table.&lt;br /&gt;
&lt;br /&gt;
Our discussion is now focused on the second condition. We find that it relies on the &#039;&#039;perfectness&#039;&#039; of the hash function for a data set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is &#039;&#039;&#039;perfect&#039;&#039;&#039; for a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of items if &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; maps all items in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; to different values, i.e. there is no collision.&lt;br /&gt;
&lt;br /&gt;
We have shown by the birthday problem for 2-universal hashing that when &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items are mapped to &amp;lt;math&amp;gt;n^2&amp;lt;/math&amp;gt; values, for an &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; chosen uniformly from a 2-universal family of hash functions, the probability that a collision occurs is at most 1/2. Thus&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[h\mbox{ is perfect for }S]\ge\frac{1}{2}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
for a table of &amp;lt;math&amp;gt;n^2&amp;lt;/math&amp;gt; entries.&lt;br /&gt;
&lt;br /&gt;
The construction of perfect hashing is straightforward then:&lt;br /&gt;
:For a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements:&lt;br /&gt;
# uniformly choose an &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; from a 2-universal family &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;; (for Carter-Wegman&#039;s construction, it means uniformly choose two integer &amp;lt;math&amp;gt;1\le a\le p-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b\in[p]&amp;lt;/math&amp;gt; for a sufficiently large prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;.)&lt;br /&gt;
# check whether &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is perfect for &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;;&lt;br /&gt;
# if &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; is NOT perfect for &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, start over again; otherwise, construct the table;&lt;br /&gt;
&lt;br /&gt;
This is a Las Vegas randomized algorithm, which construct a perfect hashing for a fixed set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with expectedly at most two trials (due to geometric distribution). The resulting data structure is a &amp;lt;math&amp;gt;O(n^2)&amp;lt;/math&amp;gt;-size static dictionary of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements which answers every search in deterministic &amp;lt;math&amp;gt;O(1)&amp;lt;/math&amp;gt; time.&lt;br /&gt;
&lt;br /&gt;
== FKS perfect hashing ==&lt;br /&gt;
In the last section we see how to use &amp;lt;math&amp;gt;O(n^2)&amp;lt;/math&amp;gt; space and constant time for answering search in a set. Now we see how to do it with linear space and constant time. This solves the problem of searching asymptotically optimal for both time and space.&lt;br /&gt;
&lt;br /&gt;
This was once seemingly impossible, until Yao&#039;s seminal paper:&lt;br /&gt;
*Yao. Should tables be sorted? &#039;&#039;Journal of the ACM (JACM)&#039;&#039;, 1981.&lt;br /&gt;
&lt;br /&gt;
Yao&#039;s paper shows a possibility of achieving linear space and constant time at the same time by exploiting the power of hashing, but assumes an unrealistically large universe. &lt;br /&gt;
&lt;br /&gt;
Inspired by Yao&#039;s work, Fredman, Komlós, and Szemerédi discover the first linear-space and constant-time static dictionary in a realistic setting: &lt;br /&gt;
* Fredman, Komlós, and Szemerédi. Storing a sparse table with O(1) worst case access time. &#039;&#039;Journal of the ACM (JACM)&#039;&#039;, 1984.&lt;br /&gt;
&lt;br /&gt;
The idea of FKS hashing is to arrange hash table in two levels:&lt;br /&gt;
* In the first level, &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items are hashed to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;buckets&#039;&#039; by a 2-universal hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt;. &lt;br /&gt;
: Let &amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt; be the set of items hashed to the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th bucket.&lt;br /&gt;
* In the second level, construct a &amp;lt;math&amp;gt;|B_i|^2&amp;lt;/math&amp;gt;-size perfect hashing for each bucket &amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The data structure can be stored in a table. The first few entries are reserved to store the primary hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt;. To help the searching algorithm locate a bucket, we use the next &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; entries of the table as the &amp;quot;pointers&amp;quot; to the bucket: each entry stores the address of the first entry of the space to store a bucket. In the rest of table, the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; buckets are stored in order, each using a &amp;lt;math&amp;gt;|B_i|^2&amp;lt;/math&amp;gt; space as required by perfect hashing.&lt;br /&gt;
&lt;br /&gt;
::[[File:FKS.png|600px]]&lt;br /&gt;
&lt;br /&gt;
It is easy to see that the search time is constant. To search for an item &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;, the algorithm does the followings:&lt;br /&gt;
* Retrieve &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Retrieve the address for bucket &amp;lt;math&amp;gt;h(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Search by perfect hashing within bucket &amp;lt;math&amp;gt;h(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Each line takes constant time. So the worst-case search time is constant.&lt;br /&gt;
&lt;br /&gt;
We then need to guarantee that the space is linear to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. At the first glance, this seems impossible because each instance of perfect hashing for a bucket costs a square-size of space. We will prove that although the individual buckets use square-sized spaces, the sum of the them is still linear.&lt;br /&gt;
&lt;br /&gt;
For a fixed set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items, for a hash function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; chosen uniformly from a 2-universe family which maps the items to &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;, called &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;buckets&#039;&#039;,  let &amp;lt;math&amp;gt;Y_i=|B_i|&amp;lt;/math&amp;gt; be the number of items in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; mapped to the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th bucket.&lt;br /&gt;
We are going to bound the following quantity:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y=\sum_{i=1}^n Y_i^2.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Since each bucket &amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt; use a space of &amp;lt;math&amp;gt;Y_i^2&amp;lt;/math&amp;gt; for perfect hashing. &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; gives the size of the space for storing the buckets. &lt;br /&gt;
&lt;br /&gt;
We will show that &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is related to the total number of collision pairs. (Indeed, the number of collision pairs can be computed by a degree-2 polynomial, just like &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;.)&lt;br /&gt;
&lt;br /&gt;
Note that a bucket of &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; items contributes &amp;lt;math&amp;gt;{Y_i\choose 2}&amp;lt;/math&amp;gt; collision pairs. Let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be the total number of collision pairs.&lt;br /&gt;
&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; can be computed by summing over the collision pairs in every bucket:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X=\sum_{i=1}^n{Y_i\choose 2}=\sum_{i=1}^n\frac{Y_i(Y_i-1)}{2}=\frac{1}{2}\left(\sum_{i=1}^nY_i^2-\sum_{i=1}^nY_i\right)=\frac{1}{2}\left(\sum_{i=1}^nY_i^2-n\right).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Therefore, the sum of squares of the sizes of buckets is related to collision number by:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\sum_{i=1}^nY_i^2=2X+n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By our analysis of the collision number, we know that for &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; items mapped to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; buckets, the expected number of collision pairs is: &amp;lt;math&amp;gt;\mathbf{E}[X]\le \frac{n}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}\left[\sum_{i=1}^nY_i^2\right]=\mathbf{E}[2X+n]\le 2n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Due to Markov&#039;s inequality, &amp;lt;math&amp;gt;\sum_{i=1}^nY_i^2=O(n)&amp;lt;/math&amp;gt; with a constant probability. For any set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, we can find a suitable &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; after expected constant number of trials, and FKS can be constructed with guaranteed (instead of expected) linear-size which answers each search in constant time.&lt;br /&gt;
&lt;br /&gt;
== Bloom filter ==&lt;br /&gt;
Now we consider the lossy representation of the original data set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, to further save the space usage. Such lossy data structure is sometimes called a &#039;&#039;&#039;&#039;&#039;sketch&#039;&#039;&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
The Bloom filter is such a lossy data structure. It is a space-efficient hash table that solves the &#039;&#039;&#039;approximate membership&#039;&#039;&#039; problem with one-sided error (&#039;&#039;false positive&#039;&#039;).&lt;br /&gt;
&lt;br /&gt;
Given a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements from a universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt;, a Bloom filter consists of an array &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;cn&amp;lt;/math&amp;gt; bits, and &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; hash functions &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k&amp;lt;/math&amp;gt; map &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[cn]&amp;lt;/math&amp;gt;, where both &amp;lt;math&amp;gt;c&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; are parameters that we can try to optimize later.&lt;br /&gt;
&lt;br /&gt;
As before, we assume the &#039;&#039;&#039;Uniform Hash Assumption (UHA)&#039;&#039;&#039;: &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k&amp;lt;/math&amp;gt; are mutually independent hash function where each &amp;lt;math&amp;gt;h_i&amp;lt;/math&amp;gt; is a uniform random hash function &amp;lt;math&amp;gt;h_i:U\to[cn]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The Bloom filter works as follows:&lt;br /&gt;
{{Theorem|&#039;&#039;Bloom filter&#039;&#039; (Bloom 1970)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k:U\to[cn]&amp;lt;/math&amp;gt; are uniform and independent random hash functions.&lt;br /&gt;
-----&lt;br /&gt;
:&#039;&#039;&#039;Data structure construction:&#039;&#039;&#039; Given a set &amp;lt;math&amp;gt;S\subset U&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;n=|S|&amp;lt;/math&amp;gt;, the data structure is a Boolean array &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;cn&amp;lt;/math&amp;gt; bits constructed as&lt;br /&gt;
:* initialize all &amp;lt;math&amp;gt;cn&amp;lt;/math&amp;gt; bits of the Boolean array &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; to 0;&lt;br /&gt;
:* for each &amp;lt;math&amp;gt;x\in S&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;A[h_i(x)]=1&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;1\le i\le k&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
:&#039;&#039;&#039;Query resolution:&#039;&#039;&#039; Upon each query of an arbitrary &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt;,&lt;br /&gt;
:* answer &amp;quot;yes&amp;quot; if &amp;lt;math&amp;gt;A[h_i(x)]=1&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;1\le i\le k&amp;lt;/math&amp;gt; and &amp;quot;no&amp;quot; if otherwise.&lt;br /&gt;
}}&lt;br /&gt;
The Boolean array is our data structure, whose size is &amp;lt;math&amp;gt;cn&amp;lt;/math&amp;gt; bits. With Uniform Hash Assumption (UHA), the time cost of the data structure for answering each query is &amp;lt;math&amp;gt;O(k)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
When the answer returned by the algorithm is &amp;quot;no&amp;quot;, it holds that &amp;lt;math&amp;gt;A[h_i(x)]=0&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;1\le i\le k&amp;lt;/math&amp;gt;, in which case the query &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; must not belong to the set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Thus, the Bloom filter has no false negatives.&lt;br /&gt;
&lt;br /&gt;
On the other hand, when the answer returned by the algorithm is &amp;quot;yes&amp;quot;, &amp;lt;math&amp;gt;A[h_i(x)]=1&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;1\le i\le k&amp;lt;/math&amp;gt;. It is still possible for some  &amp;lt;math&amp;gt;x\not\in S&amp;lt;/math&amp;gt; that all bits  &amp;lt;math&amp;gt;A[h_i(x)]&amp;lt;/math&amp;gt; are set by elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. We want to bound such false positive, that is, the following probability for an  &amp;lt;math&amp;gt;x\not\in S&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\,\forall 1\le i\le k, A[h_i(x)]=1\,]&amp;lt;/math&amp;gt;,&lt;br /&gt;
which by independence between different hash functions and by symmetry is equal to:&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\, A[h_1(x)]=1\,]^k=(1-\Pr[\, A[h_1(x)]=0\,])^k&amp;lt;/math&amp;gt;.&lt;br /&gt;
For an element &amp;lt;math&amp;gt;x\not\in S&amp;lt;/math&amp;gt;, its hash value &amp;lt;math&amp;gt;h_1(x)&amp;lt;/math&amp;gt; is independent of all hash values &amp;lt;math&amp;gt;h_i(y)&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;1\le i\le k&amp;lt;/math&amp;gt; and all &amp;lt;math&amp;gt;y\in S&amp;lt;/math&amp;gt;. This is due to the Uniform Hash Assumption. The hash value &amp;lt;math&amp;gt;h_1(x)&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;x\not\in S&amp;lt;/math&amp;gt; is then independent of the content of the array &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. Therefore, the probability of this position &amp;lt;math&amp;gt;A[h_1(x)]&amp;lt;/math&amp;gt; missed by all &amp;lt;math&amp;gt;kn&amp;lt;/math&amp;gt; updates to the Boolean array &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; caused by all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\, A[h_1(x)]=0\,]=\left(1-\frac{1}{cn}\right)^{kn}\approx e^{-k/c}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Putting everything together, for any &amp;lt;math&amp;gt;x\not\in S&amp;lt;/math&amp;gt;, the false positive is bounded as:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\,\text{wrongly answer &#039;&#039;yes&#039;&#039;}\,]&lt;br /&gt;
&amp;amp;=\Pr[\,\forall 1\le i\le k, A[h_i(x)]=1\,]\\&lt;br /&gt;
&amp;amp;=\Pr[\, A[h_1(x)]=1\,]^k=(1-\Pr[\, A[h_1(x)]=0\,])^k\\&lt;br /&gt;
&amp;amp;=\left(1-\left(1-\frac{1}{cn}\right)^{kn}\right)^k\\&lt;br /&gt;
&amp;amp;\approx \left(1- e^{-k/c}\right)^k&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is &amp;lt;math&amp;gt;(0.6185)^c&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;k=c\ln 2&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Bloom filter solves the membership query with a small constant error of false positives using a data structure of &amp;lt;math&amp;gt;O(n)&amp;lt;/math&amp;gt; bits which answers each query with &amp;lt;math&amp;gt;O(1)&amp;lt;/math&amp;gt; time cost.&lt;br /&gt;
&lt;br /&gt;
=Distinct Elements=&lt;br /&gt;
Consider the following problem of &#039;&#039;&#039;counting distinct elements&#039;&#039;&#039;: Suppose that &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; is a sufficiently large universe.&lt;br /&gt;
*&#039;&#039;&#039;Input:&#039;&#039;&#039; a sequence of (not necessarily distinct) elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in U&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&#039;&#039;&#039;Output:&#039;&#039;&#039; an estimation of the total number of distinct elements &amp;lt;math&amp;gt;z=|\{x_1,x_2,\ldots,x_n\}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A straightforward way of solving this problem is to maintain a dictionary data structure, which costs at least linear (&amp;lt;math&amp;gt;O(n)&amp;lt;/math&amp;gt;) space. For &#039;&#039;big data&#039;&#039;, where &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is very large, this is still too expensive. However, due to an information-theoretical argument, linear space is necessary if you want to compute the &#039;&#039;exact&#039;&#039; value of &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Our goal is to relax the problem a little bit to significantly reduce the space cost by tolerating &#039;&#039;approximate&#039;&#039; answers. The form of approximation we consider is &#039;&#039;&#039;&amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator&#039;&#039;&#039;.&lt;br /&gt;
{{Theorem|&amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator|&lt;br /&gt;
: A random variable &amp;lt;math&amp;gt;\widehat{Z}&amp;lt;/math&amp;gt; is an &#039;&#039;&#039;&amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator&#039;&#039;&#039; of a quantity &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; if&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[\,(1-\epsilon)z\le \widehat{Z}\le (1+\epsilon)z\,]\ge 1-\delta&amp;lt;/math&amp;gt;.&lt;br /&gt;
: &amp;lt;math&amp;gt;\widehat{Z}&amp;lt;/math&amp;gt; is said to be an &#039;&#039;&#039;unbiased estimator&#039;&#039;&#039; of &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;\mathbb{E}[\widehat{Z}]=z&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
Usually &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; is called &#039;&#039;&#039;approximation error&#039;&#039;&#039; and &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt; is called &#039;&#039;&#039;confidence error&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
We now present an elegant algorithm. The algorithm can be implemented in [https://en.wikipedia.org/wiki/Streaming_algorithm &#039;&#039;&#039;data stream model&#039;&#039;&#039;]: The input elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n&amp;lt;/math&amp;gt; is presented to the algorithm one at a time, where the size of data &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is unknown to the algorithm. The algorithm maintains a value &amp;lt;math&amp;gt;\widehat{Z}&amp;lt;/math&amp;gt; which is an &amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator of the total number of distinct elements &amp;lt;math&amp;gt;z=|\{x_1,x_2,\ldots,x_n\}|&amp;lt;/math&amp;gt;, using only a small amount of memory space to memorize (with loss) the data set &amp;lt;math&amp;gt;\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A famous quotation of Flajolet describes the performance of this algorithm as:&lt;br /&gt;
&lt;br /&gt;
 &amp;quot;Using only memory equivalent to 5 lines of printed text, you can estimate with a typical accuracy of 5% and in a single pass the total vocabulary of Shakespeare.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
== The &amp;lt;math&amp;gt;\min&amp;lt;/math&amp;gt;-sketch ==&lt;br /&gt;
Suppose that we can access to an idealized random hash function &amp;lt;math&amp;gt;h:U\to[0,1]&amp;lt;/math&amp;gt; which is uniformly distributed over all mappings from the universe &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; to unit interval &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Recall that the input sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in U&amp;lt;/math&amp;gt; consists of &amp;lt;math&amp;gt;z=|\{x_1,x_2,\ldots,x_n\}|&amp;lt;/math&amp;gt; distinct elements. These elements are mapped by the random function &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; hash values uniformly and independently distributed in &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;. We could maintain these hash values instead of the original elements, but this would still be too expensive because in the worst case we still have up to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct values to maintain. However, due to the idealized random hash function, the unit interval &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt; will be partitioned into &amp;lt;math&amp;gt;z+1&amp;lt;/math&amp;gt; subintervals by these &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; uniform and independent hash values. The typical length of the subinterval gives an estimation of the number &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}\left[\min_{1\le i\le n}h(x_i)\right]=\frac{1}{z+1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
The input sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in U&amp;lt;/math&amp;gt; consisting of &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; distinct elements are mapped to &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; random hash values uniformly and independently distributed in &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;. These &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; hash values partition the unit interval &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;z+1&amp;lt;/math&amp;gt; subintervals &amp;lt;math&amp;gt;[0,v_1],[v_1,v_2],[v_2,v_3]\ldots,[v_{z-1},v_z],[v_z,1]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;v_i&amp;lt;/math&amp;gt; denotes the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th smallest value among all hash values &amp;lt;math&amp;gt;\{h(x_1),h(x_2),\ldots,h(x_n)\}&amp;lt;/math&amp;gt;. Clearly we have &lt;br /&gt;
:&amp;lt;math&amp;gt;v_1=\min_{1\le i\le n}h(x_i)&amp;lt;/math&amp;gt;. &lt;br /&gt;
Meanwhile, since all hash values are uniformly and independently distributed in &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;, the lengths of all subintervals &amp;lt;math&amp;gt;v_1, v_2-v_1, v_3-v_2,\ldots, v_z-v_{z-1}, 1-v_z&amp;lt;/math&amp;gt; are identically distributed. By symmetry, they have the same expectation, therefore&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(z+1)\mathbb{E}[v_1]=&lt;br /&gt;
\mathbb{E}[v_1]+\sum_{i=1}^{z-1}\mathbb{E}[v_{i+1}-v_i]+\mathbb{E}[1-v_z]&lt;br /&gt;
=\mathbb{E}\left[v_1+(v_2-v_1)+(v_3-v_2)+\cdots+(v_{z}-v_{z-1})+1-v_z\right]&lt;br /&gt;
=1,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which implies that&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}\left[\min_{1\le i\le n}h(x_i)\right]=\mathbb{E}[v_1]=\frac{1}{z+1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\min_{1\le i\le n}h(x_i)&amp;lt;/math&amp;gt; can be computed with small space cost (for storing the current smallest hash value) by scan the input sequence in a single pass. Because as we proved its expectation is &amp;lt;math&amp;gt;\frac{1}{z+1}&amp;lt;/math&amp;gt;, the smallest hash value &amp;lt;math&amp;gt;Y=\min_{1\le i\le n}h(x_i)&amp;lt;/math&amp;gt; gives an unbiased estimator for &amp;lt;math&amp;gt;\frac{1}{z+1}&amp;lt;/math&amp;gt;. However, &amp;lt;math&amp;gt;\frac{1}{Y}-1&amp;lt;/math&amp;gt; is not necessarily a good estimator for &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt;. Actually, it is a rather poor estimator. Consider for example when &amp;lt;math&amp;gt;z=1&amp;lt;/math&amp;gt;, all input elements are the same. In this case, there is only one hash value and &amp;lt;math&amp;gt;Y=\min_{1\le i\le n}h(x_i)&amp;lt;/math&amp;gt; is distributed uniformly over &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;\frac{1}{Y}-1&amp;lt;/math&amp;gt; fails to be close enough to the correct answer 1 with high probability.&lt;br /&gt;
&lt;br /&gt;
==Apply the mean trick to the &amp;lt;math&amp;gt;\min&amp;lt;/math&amp;gt;-sketch==&lt;br /&gt;
The reason that the above estimator of a single hash function performs poorly is that the unbiased estimator &amp;lt;math&amp;gt;\min_{1\le i\le n}h(x_i)&amp;lt;/math&amp;gt; has large variance. So a natural way to reduce this variance is to have multiple independent hash functions and take the average. This generic approach for reducing the variance is called &#039;&#039;&#039;the mean trick&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Suppose that we can access to &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; independent random hash functions &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k&amp;lt;/math&amp;gt;, where each &amp;lt;math&amp;gt;h_j: U\to[0,1]&amp;lt;/math&amp;gt; is uniformly and independently distributed over all functions mapping &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;. Here &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; is a parameter to be fixed by the desired approximation error &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; and confidence error &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt;. The &#039;&#039;&amp;lt;math&amp;gt;\min&amp;lt;/math&amp;gt;-sketch algorithm&#039;&#039; (using the mean trick) is given by the following pseudocode.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|The &amp;lt;math&amp;gt;\min&amp;lt;/math&amp;gt;-sketch|&lt;br /&gt;
:Suppose that &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k: U\to[0,1]&amp;lt;/math&amp;gt; are &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; uniform and independent random hash functions, where &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; is a parameter to be fixed later.&lt;br /&gt;
-----&lt;br /&gt;
:Scan the input sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in U&amp;lt;/math&amp;gt; in a single pass to compute:&lt;br /&gt;
::* &amp;lt;math&amp;gt;Y_j=\min_{1\le i\le n}h_j(x_i)&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;j=1,2,\ldots,k&amp;lt;/math&amp;gt;;&lt;br /&gt;
::* average value &amp;lt;math&amp;gt;\overline{Y}=\frac{1}{k}\sum_{j=1}^kY_j&amp;lt;/math&amp;gt;;&lt;br /&gt;
:return &amp;lt;math&amp;gt;\widehat{Z}=\frac{1}{\overline{Y}}-1&amp;lt;/math&amp;gt; as the estimator.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The algorithm is easy to implement in data stream model, with a space cost of storing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; hash values. The following theorem guarantees that the algorithm returns an &amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator of the total number of distinct elements for a suitable &amp;lt;math&amp;gt;k=O\left(\frac{1}{\epsilon^2\delta}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;\epsilon,\delta&amp;lt;1/2&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;k\ge\left\lceil\frac{4}{\epsilon^2\delta}\right\rceil&amp;lt;/math&amp;gt; then the output &amp;lt;math&amp;gt;\widehat{Z}&amp;lt;/math&amp;gt; always gives an &amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator of the correct answer &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
In the following we prove this main theorem for &amp;lt;math&amp;gt;\min&amp;lt;/math&amp;gt;-sketch algorithm. &lt;br /&gt;
&lt;br /&gt;
An obstacle to analyze the estimator &amp;lt;math&amp;gt;\widehat{Z}=\frac{1}{\overline{Y}}-1&amp;lt;/math&amp;gt; is that it is a nonlinear function of &amp;lt;math&amp;gt;\overline{Y}&amp;lt;/math&amp;gt; who is easier to analyze. Nevertheless, we observe that &amp;lt;math&amp;gt;\widehat{Z}&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator of &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; as long as  &amp;lt;math&amp;gt;\overline{Y}&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;(\epsilon/2,\delta)&amp;lt;/math&amp;gt;-estimator of &amp;lt;math&amp;gt;\frac{1}{z+1}&amp;lt;/math&amp;gt;. This can be deduced by just verifying the following:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1-\epsilon/2}{z+1}\le \overline{Y}\le \frac{1+\epsilon/2}{z+1} \implies (1-\epsilon)z\le\frac{1}{\overline{Y}}-1\le (1+\epsilon)z&amp;lt;/math&amp;gt;,&lt;br /&gt;
for &amp;lt;math&amp;gt;\epsilon&amp;lt;\frac{1}{2}&amp;lt;/math&amp;gt;. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,(1-\epsilon)z\le \widehat{Z} \le (1+\epsilon)z\,\right]\ge \Pr\left[\,\frac{1-\epsilon/2}{z+1}\le \overline{Y}\le \frac{1+\epsilon/2}{z+1}\,\right]&lt;br /&gt;
=\Pr\left[\,\left|\overline{Y}-\frac{1}{z+1}\right|\le \frac{\epsilon/2}{z+1}\,\right]&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is then sufficient to show that &amp;lt;math&amp;gt;\Pr\left[\,\left|\overline{Y}-\frac{1}{z+1}\right|\le \frac{\epsilon/2}{z+1}\,\right]\ge 1-\delta&amp;lt;/math&amp;gt; for proving the main theorem above. We will see that this is equivalent to show the concentration inequality &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,\left|\overline{Y}-\mathbb{E}\left[\overline{Y}\right]\right|\le \frac{\epsilon/2}{z+1}\,\right]\ge 1-\delta\quad\qquad({\color{red}*})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:The followings hold for each &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;j=1,2\ldots,k&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;\overline{Y}=\frac{1}{k}\sum_{j=1}^kY_j&amp;lt;/math&amp;gt;:&lt;br /&gt;
:*&amp;lt;math&amp;gt;\mathbb{E}\left[\overline{Y}\right]=\mathbb{E}\left[Y_j\right]=\frac{1}{z+1}&amp;lt;/math&amp;gt;;&lt;br /&gt;
:*&amp;lt;math&amp;gt;\mathbf{Var}\left[Y_j\right]\le\frac{1}{(z+1)^2}&amp;lt;/math&amp;gt;, and consequently &amp;lt;math&amp;gt;\mathbf{Var}\left[\overline{Y}\right]\le\frac{1}{k(z+1)^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
As in the case of single hash function, by symmetry it holds that &amp;lt;math&amp;gt;\mathbb{E}[Y_j]=\frac{1}{z+1}&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;j=1,2,\ldots,k&amp;lt;/math&amp;gt;. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}\left[\overline{Y}\right]=\frac{1}{k}\sum_{j=1}^k\mathbb{E}[Y_j]=\frac{1}{z+1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Recall that each &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt; is the minimum of &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; random hash values uniformly and independently distributed over &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;. By geometry probability, it holds that for any &amp;lt;math&amp;gt;y\in[0,1]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[Y_j&amp;gt;y]=(1-y)^z&amp;lt;/math&amp;gt;,&lt;br /&gt;
which means &amp;lt;math&amp;gt;\Pr[Y_j\le y]=1-(1-y)^z&amp;lt;/math&amp;gt;. Taking the derivative with respect to &amp;lt;math&amp;gt;y&amp;lt;/math&amp;gt;, we obtain the probability density function of random variable &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt;, which is &amp;lt;math&amp;gt;z(1-y)^{z-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We then compute the second moment.&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}[Y_j^2]=\int^{1}_0y^2z(1-y)^{z-1}\,\mathrm{d}y=\frac{2}{(z+1)(z+2)}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The variance is bounded as&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{Var}\left[Y_j\right]=\mathbb{E}\left[Y_j^2\right]-\mathbb{E}\left[Y_j\right]^2=\frac{2}{(z+1)(z+2)}-\frac{1}{(z+1)^2}\le\frac{1}{(z+1)^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the (pairwise) independence between &amp;lt;math&amp;gt;Y_j&amp;lt;/math&amp;gt;&#039;s,&lt;br /&gt;
::&amp;lt;math&amp;gt;\mathbf{Var}\left[\overline{Y}\right]=\mathbf{Var}\left[\frac{1}{k}\sum_{j=1}^kY_j\right]=\frac{1}{k^2}\sum_{j=1}^k\mathbf{Var}\left[Y_j\right]\le \frac{1}{k(z+1)^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We resume to prove the inequality &amp;lt;math&amp;gt;({\color{red}*})&amp;lt;/math&amp;gt;. By [[高级算法_(Fall 2023)/Basic_deviation_inequalities#Chebyshev.27s_inequality|Chebyshev&#039;s inequality]], it holds that &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,\left|\overline{Y}-\mathbb{E}\left[\overline{Y}\right]\right|&amp;gt; \frac{\epsilon/2}{z+1}\,\right]&lt;br /&gt;
\le\frac{4}{\epsilon^2}(z+1)^2\mathbf{Var}\left[\overline{Y}\right]&lt;br /&gt;
\le\frac{4}{\epsilon^2k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
When &amp;lt;math&amp;gt;k\ge\left\lceil\frac{4}{\epsilon^2\delta}\right\rceil&amp;lt;/math&amp;gt;, this probability is at most &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt;. The inequality &amp;lt;math&amp;gt;({\color{red}*})&amp;lt;/math&amp;gt; is proved. As we discussed above, this proves the above main theorem &amp;lt;math&amp;gt;\min&amp;lt;/math&amp;gt;-sketch algorithm improved by the mean trick.&lt;br /&gt;
&lt;br /&gt;
= Frequency Estimation=&lt;br /&gt;
Suppose that &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; is the data universe. The &#039;&#039;&#039;frequency estimation&#039;&#039;&#039; problem is defined as follows.&lt;br /&gt;
*&#039;&#039;&#039;Data:&#039;&#039;&#039; a sequence of (not necessarily distinct) elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in U&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&#039;&#039;&#039;Query:&#039;&#039;&#039; an element &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&#039;&#039;&#039;Output:&#039;&#039;&#039; an estimation &amp;lt;math&amp;gt;\hat{f}_x&amp;lt;/math&amp;gt; of the frequency &amp;lt;math&amp;gt;f_x\triangleq|\{i\mid x_i=x\}|&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; in input data.&lt;br /&gt;
&lt;br /&gt;
We still want to give an algorithm in the data stream model: the algorithm scan the input sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n&amp;lt;/math&amp;gt; to construct a succinct data structure, such that upon each query of &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt;, the algorithm returns an estimation of the frequency &amp;lt;math&amp;gt;f_x&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Clearly this problem can always be solved by storing all appeared distinct elements along with their frequencies. However, the space cost of this straightforward solution is rather high. Instead, we want to use a lossy representation (a &#039;&#039;sketch&#039;&#039;) of input data which uses significantly less space but can still answer queries with tolarable accuracy. &lt;br /&gt;
&lt;br /&gt;
Formally, upon each query of &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt;, the algorithm should return an answer &amp;lt;math&amp;gt;\hat{f}_x&amp;lt;/math&amp;gt; satisfying:&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,\left|\hat{f}_x-f_x\right|\le \epsilon n\,\right]\ge 1-\delta&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that this notion of approximation is with bounded &#039;&#039;additive&#039;&#039; error which is weaker than the notion of &amp;lt;math&amp;gt;(\epsilon,\delta)&amp;lt;/math&amp;gt;-estimator, whose error bound is &#039;&#039;multiplicative&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
With such weak accuracy guarantee, its is possible to give a succinct data structure whose size is determined only by the error bounds &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt; but independent of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, because only the frequencies of those &#039;&#039;&#039;heavy hitters&#039;&#039;&#039; (elements &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; with high frequencies &amp;lt;math&amp;gt;f_x&amp;gt;\epsilon n&amp;lt;/math&amp;gt;) need to be memorized, and there are at most &amp;lt;math&amp;gt;1/\epsilon&amp;lt;/math&amp;gt; many such heavy hitters.&lt;br /&gt;
&lt;br /&gt;
== Count-min sketch==&lt;br /&gt;
The [https://en.wikipedia.org/wiki/Count–min_sketch count-min sketch] given by Cormode and Muthukrishnan is an elegant data structure for frequency estimation.&lt;br /&gt;
&lt;br /&gt;
The data structure is a two-dimensional &amp;lt;math&amp;gt;k\times m&amp;lt;/math&amp;gt; integer array, where &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; are two parameters to be determined by the error bounds &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt;. We still adopt the Uniform Hash Assumption to assume that we have access to &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; mutually independent uniform random hash functions &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k: U\to[m]&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem|&#039;&#039;Count-min sketch&#039;&#039; (Cormode and Muthukrishnan 2003)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k: U\to[m]&amp;lt;/math&amp;gt; are uniform and independent random hash functions.&lt;br /&gt;
-----&lt;br /&gt;
:&#039;&#039;&#039;Data structure construction:&#039;&#039;&#039; Given a sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in U&amp;lt;/math&amp;gt;, the data structure is a two-dimensional &amp;lt;math&amp;gt;k\times m&amp;lt;/math&amp;gt; integer array &amp;lt;math&amp;gt;CMS[k][m]&amp;lt;/math&amp;gt; constructed as&lt;br /&gt;
:*initialize all entries of &amp;lt;math&amp;gt;CMS[k][m]&amp;lt;/math&amp;gt; to 0;&lt;br /&gt;
:*for &amp;lt;math&amp;gt;i=1,2,\ldots,n&amp;lt;/math&amp;gt;, upon receiving &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt;:&lt;br /&gt;
::: for every &amp;lt;math&amp;gt;1\le j\le k&amp;lt;/math&amp;gt;, evaluate &amp;lt;math&amp;gt;h_j(x_i)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;CMS[j][h_j(x_i)]++&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
:&#039;&#039;&#039;Query resolution:&#039;&#039;&#039; Upon each query of an arbitrary &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt;,&lt;br /&gt;
:* return &amp;lt;math&amp;gt;\hat{f}_x=\min_{1\le j\le k}CMS[j][h_j(x)]&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to see that the space cost of count-min sketch is &amp;lt;math&amp;gt;O(km)&amp;lt;/math&amp;gt; memory words, or &amp;lt;math&amp;gt;O(km\log n)&amp;lt;/math&amp;gt; bits. Each query is answered within time cost &amp;lt;math&amp;gt;O(k)&amp;lt;/math&amp;gt;, assuming that an evaluation of hash function can be done in unit or constant time. We then analyze the error bounds.&lt;br /&gt;
&lt;br /&gt;
First, it is easy to observe that for any query &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt; and every hash function &amp;lt;math&amp;gt;1\le j\le k&amp;lt;/math&amp;gt;, it always holds for the corresponding entry in the count-min sketch&lt;br /&gt;
:&amp;lt;math&amp;gt;CMS[j][h_j(x)]\ge f_x&amp;lt;/math&amp;gt;,&lt;br /&gt;
because the appearances of element &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; in the input sequence contribute at least &amp;lt;math&amp;gt;f_x&amp;lt;/math&amp;gt; to the value of &amp;lt;math&amp;gt;CMS[j][h_j(x)]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, for any query &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt; it always holds for the answer &amp;lt;math&amp;gt;\hat{f}_x=\min_{1\le j\le k}CMS[j][h_j(x)]\ge f_x&amp;lt;/math&amp;gt;, which means&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,\left|\hat{f}_x- f_x\right|\ge\epsilon n\,\right]=\Pr\left[\,\hat{f}_x- f_x\ge\epsilon n\,\right]=\prod_{j=1}^k\Pr[\,CMS[j][h_j(x)]-f_x\ge\epsilon n\,],\quad\qquad({\color{red}\diamondsuit})&amp;lt;/math&amp;gt;&lt;br /&gt;
where the second equation is due to the mutual independence of random hash functions &amp;lt;math&amp;gt;h_1,h_2,\ldots,h_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
It remains to upper bound the probability &amp;lt;math&amp;gt;\Pr[\,CMS[j][h_j(x)]-f_x\ge\epsilon n\,]&amp;lt;/math&amp;gt;, which can be done by calculating the expectation of &amp;lt;math&amp;gt;CMS[j][h_j(x)]&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt; and every &amp;lt;math&amp;gt;1\le j\le k&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;\mathbb{E}\left[CMS[j][h_j(x)]\right]\le f_x+\frac{n}{m}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
The value of &amp;lt;math&amp;gt;CMS[j][h_j(x)]&amp;lt;/math&amp;gt; is constituted by the frequency &amp;lt;math&amp;gt;f_x&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; and the frequencies &amp;lt;math&amp;gt;f_y&amp;lt;/math&amp;gt; of all other elements &amp;lt;math&amp;gt;y\neq x&amp;lt;/math&amp;gt; among &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
CMS[j][h_j(x)]&lt;br /&gt;
&amp;amp;=f_x+\sum_{\scriptstyle y\in\{x_1,\ldots,x_n\}\setminus\{x\}\atop\scriptstyle h_j(y)=h_j(x)} f_y\\&lt;br /&gt;
&amp;amp;=f_x+\sum_{y\in\{x_1,\ldots,x_n\}\setminus\{x\}} f_y \cdot I[h_j(y)=h_j(x)]&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;I[h_j(y)=h_j(x)]&amp;lt;/math&amp;gt; denotes the Boolean random variable that indicates the occurrence of event &amp;lt;math&amp;gt;h_j(y)=h_j(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}[CMS[j][h_j(x)]]=f_x+\sum_{y\in\{x_1,x_2,\ldots,x_n\}\setminus\{x\}} f_y \cdot \Pr[h_j(y)=h_j(x)]&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to Uniform Hash Assumption (UHA), &amp;lt;math&amp;gt;h_j: U\to[m]&amp;lt;/math&amp;gt; is a uniform random function. For any &amp;lt;math&amp;gt;y\neq x&amp;lt;/math&amp;gt;, the probability of hash collision is&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[h_j(y)=h_j(x)]=\frac{1}{m}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathbb{E}[CMS[j][h_j(x)]]&lt;br /&gt;
&amp;amp;=f_x+\frac{1}{m}\sum_{y\in\{x_1,\ldots,x_n\}\setminus\{x\}} f_y \\&lt;br /&gt;
&amp;amp;\le f_x+\frac{1}{m}\sum_{y\in\{x_1,\ldots,x_n\}} f_y\\&lt;br /&gt;
&amp;amp;=f_x+\frac{n}{m},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to the obvious identity &amp;lt;math&amp;gt;\sum_{y\in\{x_1,\ldots,x_n\}}f_y=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The above proposition shows that for any &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt; and every &amp;lt;math&amp;gt;1\le j\le k&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}\left[CMS[j][h_j(x)]-f_x\right]\le \frac{n}{m}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;CMS[j][h_j(x)]\ge f_x&amp;lt;/math&amp;gt; always holds, thus &amp;lt;math&amp;gt;CMS[j][h_j(x)]-f_x&amp;lt;/math&amp;gt; is a positive random variable. By Markov&#039;s inequality, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\,CMS[j][h_j(x)]-f_x\ge\epsilon n\,]\le \frac{1}{\epsilon m}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Combining with above equation &amp;lt;math&amp;gt;({\color{red}\diamondsuit})&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,\left|\hat{f}_x- f_x\right|\ge\epsilon n\,\right]=(\Pr[\,CMS[j][h_j(x)]-f_x\ge\epsilon n\,])^k\le \frac{1}{(\epsilon m)^k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By setting &amp;lt;math&amp;gt;m=\left\lceil\frac{\mathrm{e}}{\epsilon}\right\rceil&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;k=\left\lceil\ln\frac{1}{\delta}\right\rceil&amp;lt;/math&amp;gt;, the above error probability is bounded as &amp;lt;math&amp;gt;\frac{1}{(\epsilon m)^k}\le\delta&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For any positive &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt;, the count-min sketch gives a data structure of size &amp;lt;math&amp;gt;O(km)=O\left(\frac{1}{\epsilon}\log\frac{1}{\delta}\right)&amp;lt;/math&amp;gt; (in memory words) and answering each query &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt; in time &amp;lt;math&amp;gt;O(k)=O\left(\frac{1}{\epsilon}\right)&amp;lt;/math&amp;gt; with the following accuracy guarantee:&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\,\left|\hat{f}_x- f_x\right|\le\epsilon n\,\right]\ge 1-\delta&amp;lt;/math&amp;gt;.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13903</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13903"/>
		<updated>2026-09-07T10:12:32Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Hashing and Sketching|Hashing and Sketching]]   &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Limited independence|Limited independence]]&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Basic deviation inequalities|Basic deviation inequalities]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Fingerprinting&amp;diff=13895</id>
		<title>高级算法 (Fall 2026)/Fingerprinting</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Fingerprinting&amp;diff=13895"/>
		<updated>2026-09-02T16:03:28Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Freivalds Algorithm */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=  Checking Matrix Multiplication=&lt;br /&gt;
[[File: matrix_multiplication.png|thumb|360px|right|The evolution of time complexity &amp;lt;math&amp;gt;O(n^{\omega})&amp;lt;/math&amp;gt; for matrix multiplication.]]&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; be a feild (you may think of it as the filed &amp;lt;math&amp;gt;\mathbb{Q}&amp;lt;/math&amp;gt; of rational numbers, or the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; of integers modulo prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;). We suppose that each field operation (addition, subtraction, multiplication, division) has unit cost. This model is called the &#039;&#039;&#039;unit-cost RAM&#039;&#039;&#039; model, which is an ideal abstraction of a computer.&lt;br /&gt;
&lt;br /&gt;
Consider the following problem:&lt;br /&gt;
* &#039;&#039;&#039;Input&#039;&#039;&#039;: Three &amp;lt;math&amp;gt;n\times n&amp;lt;/math&amp;gt; matrices &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; over the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* &#039;&#039;&#039;Output&#039;&#039;&#039;: &amp;quot;yes&amp;quot; if &amp;lt;math&amp;gt;C=AB&amp;lt;/math&amp;gt; and &amp;quot;no&amp;quot; if otherwise.&lt;br /&gt;
&lt;br /&gt;
A naive way to solve this is to multiply &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; and compare the result with &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;. &lt;br /&gt;
The straightforward algorithm for matrix multiplication takes &amp;lt;math&amp;gt;O(n^3)&amp;lt;/math&amp;gt; time, assuming that each arithmetic operation takes unit time.&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Strassen_algorithm Strassen&#039;s algorithm] discovered in 1969 now implemented by many numerical libraries runs in time &amp;lt;math&amp;gt;O(n^{\log_2 7})\approx O(n^{2.81})&amp;lt;/math&amp;gt;. Strassen&#039;s algorithm starts the search for fast matrix multiplication algorithms. The [http://en.wikipedia.org/wiki/Coppersmith%E2%80%93Winograd_algorithm Coppersmith–Winograd algorithm] discovered in 1987 runs in time &amp;lt;math&amp;gt;O(n^{2.376})&amp;lt;/math&amp;gt; but is only faster than Strassens&#039; algorithm on extremely large matrices due to the very large constant coefficient. This has been the best known for decades, until recently Stothers got an &amp;lt;math&amp;gt;O(n^{2.374})&amp;lt;/math&amp;gt; algorithm in his PhD thesis in 2010, and independently Vassilevska Williams got an &amp;lt;math&amp;gt;O(n^{2.373})&amp;lt;/math&amp;gt; algorithm in 2012. Both these improvements are based on generalization of Coppersmith–Winograd algorithm. It is unknown whether the matrix multiplication can be done in time &amp;lt;math&amp;gt;O(n^{2+o(1)})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Freivalds Algorithm ==&lt;br /&gt;
The following is a very simple randomized algorithm due to Freivalds, running in &amp;lt;math&amp;gt;O(n^2)&amp;lt;/math&amp;gt; time:&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Algorithm (Freivalds, 1977)|&lt;br /&gt;
*pick a vector &amp;lt;math&amp;gt;r \in\{0, 1\}^n&amp;lt;/math&amp;gt; uniformly at random;&lt;br /&gt;
*if &amp;lt;math&amp;gt;A(Br) = Cr&amp;lt;/math&amp;gt; then return &amp;quot;yes&amp;quot; else return &amp;quot;no&amp;quot;;&lt;br /&gt;
}}&lt;br /&gt;
The product &amp;lt;math&amp;gt;A(Br)&amp;lt;/math&amp;gt; is computed by first multiplying &amp;lt;math&amp;gt;Br&amp;lt;/math&amp;gt; and then &amp;lt;math&amp;gt;A(Br)&amp;lt;/math&amp;gt;.&lt;br /&gt;
The running time of Freivalds algorithm is &amp;lt;math&amp;gt;O(n^2)&amp;lt;/math&amp;gt; because the algorithm computes 3 matrix-vector multiplications. &lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;A(Br) = Cr&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;r \in\{0, 1\}^n&amp;lt;/math&amp;gt;, thus the algorithm will return a &amp;quot;yes&amp;quot; for any positive instance (&amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt;). &lt;br /&gt;
But if &amp;lt;math&amp;gt;AB \neq C&amp;lt;/math&amp;gt; then the algorithm will make a mistake if it chooses such an &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;ABr = Cr&amp;lt;/math&amp;gt;. However, the following lemma states that the probability of this event is bounded.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:If &amp;lt;math&amp;gt;AB\neq C&amp;lt;/math&amp;gt; then for a uniformly random &amp;lt;math&amp;gt;r \in\{0, 1\}^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[ABr = Cr]\le \frac{1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Let &amp;lt;math&amp;gt;D=AB-C&amp;lt;/math&amp;gt;. The event &amp;lt;math&amp;gt;ABr=Cr&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;Dr=0&amp;lt;/math&amp;gt;. It is then sufficient to show that for a &amp;lt;math&amp;gt;D\neq \boldsymbol{0}&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;\Pr[Dr = \boldsymbol{0}]\le \frac{1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Since &amp;lt;math&amp;gt;D\neq \boldsymbol{0}&amp;lt;/math&amp;gt;, it must have at least one non-zero entry. Suppose that &amp;lt;math&amp;gt;D_{ij}\neq 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We assume the event that &amp;lt;math&amp;gt;Dr=\boldsymbol{0}&amp;lt;/math&amp;gt;. In particular, the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th entry of &amp;lt;math&amp;gt;Dr&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;(Dr)_{i}=\sum_{k=1}^n D_{ik}r_k=0.&amp;lt;/math&amp;gt; &lt;br /&gt;
The &amp;lt;math&amp;gt;r_j&amp;lt;/math&amp;gt; can be calculated by&lt;br /&gt;
:&amp;lt;math&amp;gt;r_j=-\frac{1}{D_{ij}}\sum_{k\neq j}^n D_{ik}r_k.&amp;lt;/math&amp;gt;&lt;br /&gt;
Once all other entries &amp;lt;math&amp;gt;r_k&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;k\neq j&amp;lt;/math&amp;gt; are fixed, there is a unique solution of &amp;lt;math&amp;gt;r_j&amp;lt;/math&amp;gt;. Therefore, the number of &amp;lt;math&amp;gt;r\in\{0,1\}^n&amp;lt;/math&amp;gt; satisfying &amp;lt;math&amp;gt;Dr=\boldsymbol{0}&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;2^{n-1}&amp;lt;/math&amp;gt;. The probability that &amp;lt;math&amp;gt;ABr=Cr&amp;lt;/math&amp;gt; is bounded as&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[ABr=Cr]=\Pr[Dr=\boldsymbol{0}]\le\frac{2^{n-1}}{2^n}=\frac{1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When &amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt;, Freivalds algorithm always returns &amp;quot;yes&amp;quot;; and when &amp;lt;math&amp;gt;AB\neq C&amp;lt;/math&amp;gt;, Freivalds algorithm returns &amp;quot;no&amp;quot; with probability at least 1/2.&lt;br /&gt;
&lt;br /&gt;
To improve its accuracy, we can run Freivalds algorithm for &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; times, each time with an &#039;&#039;independent&#039;&#039; &amp;lt;math&amp;gt;r\in\{0,1\}^n&amp;lt;/math&amp;gt;, and return &amp;quot;yes&amp;quot; if and only if all running instances returns &amp;quot;yes&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Freivalds&#039; Algorithm (multi-round)|&lt;br /&gt;
*pick &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vectors &amp;lt;math&amp;gt;r_1,r_2,\ldots,r_k \in\{0, 1\}^n&amp;lt;/math&amp;gt; uniformly and independently at random;&lt;br /&gt;
*if &amp;lt;math&amp;gt;A(Br_i) = Cr_i&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,\ldots,k&amp;lt;/math&amp;gt; then return &amp;quot;yes&amp;quot; else return &amp;quot;no&amp;quot;;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt;, then the algorithm returns a &amp;quot;yes&amp;quot; with probability 1. If &amp;lt;math&amp;gt;AB\neq C&amp;lt;/math&amp;gt;, then due to the independence, the probability that all &amp;lt;math&amp;gt;r_i&amp;lt;/math&amp;gt; have &amp;lt;math&amp;gt;ABr_i=C_i&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;2^{-k}&amp;lt;/math&amp;gt;, so the algorithm returns &amp;quot;no&amp;quot; with probability at least &amp;lt;math&amp;gt;1-2^{-k}&amp;lt;/math&amp;gt;. For any &amp;lt;math&amp;gt;0&amp;lt;\epsilon&amp;lt;1&amp;lt;/math&amp;gt;, choose &amp;lt;math&amp;gt;k=\log_2 \frac{1}{\epsilon}&amp;lt;/math&amp;gt;. The algorithm runs in time &amp;lt;math&amp;gt;O(n^2\log_2\frac{1}{\epsilon})&amp;lt;/math&amp;gt; and has a one-sided error (false positive) bounded by &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=Polynomial Identity Testing (PIT) =&lt;br /&gt;
The  &#039;&#039;&#039;Polynomial Identity Testing (PIT)&#039;&#039;&#039; is such a problem: given as input two polynomials, determine whether they are identical. It plays a fundamental role in &#039;&#039;Identity Testing&#039;&#039; problems.&lt;br /&gt;
&lt;br /&gt;
First, let&#039;s consider the univariate (&amp;quot;one variable&amp;quot;) case:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; two polynomials &amp;lt;math&amp;gt;f, g\in\mathbb{F}[x]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv g&amp;lt;/math&amp;gt; (&amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; are identical).&lt;br /&gt;
Here the &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; denotes the [http://en.wikipedia.org/wiki/Polynomial_ring ring of univariate polynomials] on a field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. More precisely, a polynomial &amp;lt;math&amp;gt;f\in\mathbb{F}[x]&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x)=\sum_{i=0}^\infty a_ix^i&amp;lt;/math&amp;gt;,&lt;br /&gt;
where the coefficients &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; are taken from the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;, and the addition and multiplication are also defined over the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. &lt;br /&gt;
And:&lt;br /&gt;
* the &#039;&#039;&#039;degree&#039;&#039;&#039; of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is the highest &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; with non-zero &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;;&lt;br /&gt;
* a polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is a &#039;&#039;&#039;zero-polynomial&#039;&#039;&#039;, denoted as &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;, if all coefficients &amp;lt;math&amp;gt;a_i=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Alternatively, we can consider the following equivalent problem by comparing the polynomial &amp;lt;math&amp;gt;f-g&amp;lt;/math&amp;gt; (whose degree is at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;) with the zero-polynomial:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; a polynomial &amp;lt;math&amp;gt;f\in\mathbb{F}[x]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt; (&amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is the 0 polynomial).&lt;br /&gt;
&lt;br /&gt;
The problem is trivial if the input polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is given explicitly: one can trivially solve the problem by checking whether all &amp;lt;math&amp;gt;d+1&amp;lt;/math&amp;gt; coefficients are &amp;lt;math&amp;gt;0&amp;lt;/math&amp;gt;. To make the problem nontrivial, we assume that the input polynomial is given implicitly as a &#039;&#039;black box&#039;&#039; (also called an &#039;&#039;oracle&#039;&#039;): the only way the algorithm can access to &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is to evaluate &amp;lt;math&amp;gt;f(x)&amp;lt;/math&amp;gt; over some &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; from the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; chosen by the algorithm.&lt;br /&gt;
&lt;br /&gt;
A straightforward deterministic algorithm is to evaluate &amp;lt;math&amp;gt;f(x_1),f(x_2),\ldots,f(x_{d+1})&amp;lt;/math&amp;gt; over &amp;lt;math&amp;gt;d+1&amp;lt;/math&amp;gt; &#039;&#039;distinct&#039;&#039; elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_{d+1}&amp;lt;/math&amp;gt; from the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; and check whether they are all zero. By the [https://en.wikipedia.org/wiki/Fundamental_theorem_of_algebra fundamental theorem of algebra], also known as polynomial interpolations, this guarantees to verify whether a degree-&amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; univariate polynomial &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem|Fundamental Theorem of Algebra|&lt;br /&gt;
:Any non-zero univariate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; roots.&lt;br /&gt;
}}&lt;br /&gt;
The reason for this fundamental theorem holding generally over any field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; is that any univariate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; factors uniquely into at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; irreducible polynomials, each of which has at most one root.&lt;br /&gt;
&lt;br /&gt;
The following simple randomized algorithm is natural:&lt;br /&gt;
{{Theorem|Algorithm for PIT|&lt;br /&gt;
*suppose we have a finite subset &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; (to be specified later);&lt;br /&gt;
*pick &amp;lt;math&amp;gt;r\in S&amp;lt;/math&amp;gt; &#039;&#039;uniformly&#039;&#039; at random;&lt;br /&gt;
*if &amp;lt;math&amp;gt;f(r) = 0&amp;lt;/math&amp;gt; then return “yes” else return “no”;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This algorithm evaluates &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; at one point chosen uniformly at random from a finite subset &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt;. It is easy to see the followings:&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;, the algorithm always returns &amp;quot;yes&amp;quot;, so it is always correct.&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;, the algorithm may wrongly return &amp;quot;yes&amp;quot; (a &#039;&#039;&#039;false positive&#039;&#039;&#039;). But this happens only when the random &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is a root of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;. By the fundamental theorem of algebra, &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; roots, so the probability that the algorithm is wrong is bounded as&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[f(r)=0]\le\frac{d}{|S|}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By fixing &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; to be an arbitrary subset of size &amp;lt;math&amp;gt;|S|=2d&amp;lt;/math&amp;gt;, this probability of false positive is at most &amp;lt;math&amp;gt;1/2&amp;lt;/math&amp;gt;. We can reduce it to an arbitrarily small constant &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt; by repeat the above testing &#039;&#039;independently&#039;&#039; for &amp;lt;math&amp;gt;\log_2 \frac{1}{\delta}&amp;lt;/math&amp;gt; times, since the error probability decays geometrically as we repeat the algorithm independently. &lt;br /&gt;
&lt;br /&gt;
== Communication Complexity of Equality ==&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Communication_complexity communication complexity] is introduced by Andrew Chi-Chih Yao as a model of computation with more than one entities, each with partial information about the input.&lt;br /&gt;
&lt;br /&gt;
Assume that there are two entities, say Alice and Bob. Alice has a private input &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; and Bob has a private input &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;. Together they want to compute a function &amp;lt;math&amp;gt;f(a,b)&amp;lt;/math&amp;gt; by communicating with each other. The communication follows a predefined &#039;&#039;&#039;communication protocol&#039;&#039;&#039; (the &amp;quot;algorithm&amp;quot; in this model). The complexity of a communication protocol is measured by the number of bits communicated between Alice and Bob in the worst case.&lt;br /&gt;
&lt;br /&gt;
The problem of checking identity is formally defined by the function EQ as follows: &amp;lt;math&amp;gt;\mathrm{EQ}:\{0,1\}^n\times\{0,1\}^n\rightarrow\{0,1\}&amp;lt;/math&amp;gt; and for any &amp;lt;math&amp;gt;a,b\in\{0,1\}^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathrm{EQ}(a,b)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{if } a=b,\\&lt;br /&gt;
0&amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A trivial way to solve EQ is to let Bob send his entire input string &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; to Alice and let Alice check whether &amp;lt;math&amp;gt;a=b&amp;lt;/math&amp;gt;. This costs &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bits of communications.&lt;br /&gt;
&lt;br /&gt;
It is known that for deterministic communication protocols, this is the best we can get for computing EQ.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Yao 1979)|&lt;br /&gt;
:Any deterministic communication protocol computing EQ on two &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-bit strings costs &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bits of communication in the worst-case.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This theorem is much more nontrivial to prove than it looks, because Alice and Bob are allowed to interact with each other in arbitrary ways. The proof of this theorem is in Yao&#039;s [http://math.ucla.edu/~znorwood/290d.2.14s/papers/yao.pdf  celebrated paper in 1979 with a humble title]. It pioneered the field of communication complexity.&lt;br /&gt;
&lt;br /&gt;
If we allow randomness in protocols, and also tolerate a small probabilistic error, the problem can be solved with significantly less communications. To present this randomized protocol, we need a few preparations:&lt;br /&gt;
* We represent the inputs  &amp;lt;math&amp;gt;a,b \in\{0,1\}^{n}&amp;lt;/math&amp;gt; of Alice and Bob as two univariate polynomials of degree at most &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, respectively &lt;br /&gt;
::&amp;lt;math&amp;gt;f(x)=\sum_{i=0}^{n-1}a_ix^{i}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(x)=\sum_{i=0}^{n-1}b_ix^{i}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* The two polynomials &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; are defined over finite field &amp;lt;math&amp;gt;\mathbb{Z}_p=\{0,1,\ldots,p-1\}&amp;lt;/math&amp;gt; for some suitable prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; (to be specified later), which means the additions and multiplications are modulo &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;.&lt;br /&gt;
The randomized communication protocol is then as follows:&lt;br /&gt;
{{Theorem|A randomized protocol for EQ|&lt;br /&gt;
&#039;&#039;&#039;Bob does&#039;&#039;&#039;:&lt;br /&gt;
:* pick &amp;lt;math&amp;gt;r\in\mathbb{Z}_p&amp;lt;/math&amp;gt; uniformly at random;&lt;br /&gt;
:* send &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(r)&amp;lt;/math&amp;gt; to Alice;&lt;br /&gt;
&#039;&#039;&#039;Upon receiving&#039;&#039;&#039; &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(r)&amp;lt;/math&amp;gt; &#039;&#039;&#039;Alice does&#039;&#039;&#039;:&lt;br /&gt;
:* compute &amp;lt;math&amp;gt;f(r)&amp;lt;/math&amp;gt;;&lt;br /&gt;
:* If &amp;lt;math&amp;gt;f(r)= g(r)&amp;lt;/math&amp;gt; return &amp;quot;&#039;&#039;&#039;yes&#039;&#039;&#039;&amp;quot;; else return &amp;quot;&#039;&#039;&#039;no&#039;&#039;&#039;&amp;quot;.&lt;br /&gt;
}}&lt;br /&gt;
The communication complexity of the protocol is given by the number of bits used to represent the values of &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(r)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since the polynomials are defined over finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; and the random number &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is also chosen from &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, this is bounded by &amp;lt;math&amp;gt;O(\log p)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
On the other hand the protocol makes mistakes only when &amp;lt;math&amp;gt;a\neq b&amp;lt;/math&amp;gt; but wrongly answers &amp;quot;yes&amp;quot;. This happens only when &amp;lt;math&amp;gt;f\not\equiv g&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;f(r)=g(r)&amp;lt;/math&amp;gt;. The degrees of &amp;lt;math&amp;gt;f, g&amp;lt;/math&amp;gt; are at most &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen among &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; distinct values, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[f(r)=g(r)]\le \frac{n-1}{p}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By choosing &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; to be a prime in the interval &amp;lt;math&amp;gt;[n^2, 2n^2]&amp;lt;/math&amp;gt; (by [https://en.wikipedia.org/wiki/Bertrand%27s_postulate Chebyshev&#039;s theorem], such prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; always exists), the above randomized communication protocol solves the Equality function EQ with an error probability of false positive at most &amp;lt;math&amp;gt;O(1/n)&amp;lt;/math&amp;gt;, with communication complexity &amp;lt;math&amp;gt;O(\log n)&amp;lt;/math&amp;gt;, an EXPONENTIAL improvement to ANY deterministic communication protocol!&lt;br /&gt;
&lt;br /&gt;
== Schwartz-Zippel Theorem ==&lt;br /&gt;
Now let&#039;s see the the true form of &#039;&#039;&#039;Polynomial Identity Testing (PIT)&#039;&#039;&#039;, for multivariate polynomials:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; two &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomials &amp;lt;math&amp;gt;f, g\in\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv g&amp;lt;/math&amp;gt;.&lt;br /&gt;
The &amp;lt;math&amp;gt;\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; is the [http://en.wikipedia.org/wiki/Polynomial_ring#The_polynomial_ring_in_several_variables ring of multivariate polynomials] over field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. An &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, written as a sum of monomials, is:&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)=\sum_{i_1,i_2,\ldots,i_n\ge 0\atop i_1+i_2+\cdots+i_n\le d}a_{i_1,i_2,\ldots,i_n}x_{1}^{i_1}x_2^{i_2}\cdots x_{n}^{i_n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The &#039;&#039;&#039;degree&#039;&#039;&#039; or &#039;&#039;&#039;total degree&#039;&#039;&#039; of a monomial &amp;lt;math&amp;gt;a_{i_1,i_2,\ldots,i_n}x_{1}^{i_1}x_2^{i_2}\cdots x_{n}^{i_n}&amp;lt;/math&amp;gt; is given by &amp;lt;math&amp;gt;i_1+i_2+\cdots+i_n&amp;lt;/math&amp;gt; and the degree of a polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is the maximum degree of monomials of nonzero coefficients.&lt;br /&gt;
&lt;br /&gt;
As before, we also consider the following equivalent problem:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; a polynomial &amp;lt;math&amp;gt;f\in\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is written explicitly as a sum of monomials, then the problem can be solved by checking whether all coefficients, and there at most &amp;lt;math&amp;gt;{n+d\choose d}\le (n+d)^{d}&amp;lt;/math&amp;gt; coefficients in an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
A multivariate polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; can also be presented in its &#039;&#039;&#039;product form&#039;&#039;&#039;, for example:&lt;br /&gt;
{{Theorem|Example|&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Vandermonde_matrix Vandermonde matrix] &amp;lt;math&amp;gt;M=M(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;M_{ij}=x_i^{j-1}&amp;lt;/math&amp;gt;, that is&lt;br /&gt;
:&amp;lt;math&amp;gt;M=\begin{bmatrix}&lt;br /&gt;
1 &amp;amp; x_1 &amp;amp; x_1^2 &amp;amp; \dots &amp;amp; x_1^{n-1}\\&lt;br /&gt;
1 &amp;amp; x_2 &amp;amp; x_2^2 &amp;amp; \dots &amp;amp; x_2^{n-1}\\&lt;br /&gt;
1 &amp;amp; x_3 &amp;amp; x_3^2 &amp;amp; \dots &amp;amp; x_3^{n-1}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp;\vdots \\&lt;br /&gt;
1 &amp;amp; x_n &amp;amp; x_n^2 &amp;amp; \dots &amp;amp; x_n^{n-1}&lt;br /&gt;
\end{bmatrix}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be the polynomial defined as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
f(x_1,\ldots,x_n)=\det(M)=\prod_{j&amp;lt;i}(x_i-x_j).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
For polynomials in product form, it is quite efficient to &#039;&#039;&#039;evaluate&#039;&#039;&#039; the polynomial at any specific point from the field over which the polynomial is defined, however, &#039;&#039;&#039;expanding&#039;&#039;&#039; the polynomial to a sum of monomials can be very expensive.  &lt;br /&gt;
&lt;br /&gt;
The following is a simple randomized algorithm for testing identity of multivariate polynomials:&lt;br /&gt;
{{Theorem|Randomized algorithm for multivariate PIT|&lt;br /&gt;
*suppose we have a finite subset &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; (to be specified later);&lt;br /&gt;
* pick &amp;lt;math&amp;gt;r_1,r_2,\ldots,r_n\in S&amp;lt;/math&amp;gt; &#039;&#039;uniformly&#039;&#039; and &#039;&#039;independently&#039;&#039; at random;&lt;br /&gt;
* if &amp;lt;math&amp;gt;f(\vec{r})=f(r_1,r_2,\ldots,r_n) = 0&amp;lt;/math&amp;gt; then return “yes” else return “no”;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This algorithm evaluates &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; at one point chosen uniformly from an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-dimensional cube &amp;lt;math&amp;gt;S^n&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; is a finite subset. And:&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;, the algorithm always returns &amp;quot;yes&amp;quot;, so it is always correct.&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;, the algorithm may wrongly return &amp;quot;yes&amp;quot; (a &#039;&#039;&#039;false positive&#039;&#039;&#039;). But this happens only when the random &amp;lt;math&amp;gt;\vec{r}=(r_1,r_2,\ldots,r_n)&amp;lt;/math&amp;gt; is a root of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;. The probability of this bad event is upper bounded by the following famous result due to [https://pdfs.semanticscholar.org/b913/cf330852035f49b4ec5fe2db86c47d8a98fd.pdf Schwartz (1980)] and [http://www.cecm.sfu.ca/~monaganm/teaching/TopicsinCA15/zippel79.pdf Zippel (1979)]. &lt;br /&gt;
&lt;br /&gt;
{{Theorem|Schwartz-Zippel Theorem|&lt;br /&gt;
: Let &amp;lt;math&amp;gt;f\in\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; be a multivariate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; over a field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;, then for any finite set &amp;lt;math&amp;gt;S\subset\mathbb{F}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;r_1,r_2\ldots,r_n\in S&amp;lt;/math&amp;gt; chosen uniformly and independently at random, &lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[f(r_1,r_2,\ldots,r_n)=0]\le\frac{d}{|S|}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
The Schwartz-Zippel Theorem states that for any nonzero &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, the number of roots in any cube &amp;lt;math&amp;gt;S^n&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;d\cdot |S|^{n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Dana Moshkovitz gave a surprisingly simply and elegant [http://eccc.hpi-web.de/report/2010/096/ proof] of Schwartz-Zippel Theorem, using some advanced ideas. Now we introduce the standard proof by induction.&lt;br /&gt;
&lt;br /&gt;
{{Proof| &lt;br /&gt;
The theorem is proved by induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
;Induction basis&#039;&#039;&#039;:&#039;&#039;&#039;&lt;br /&gt;
For &amp;lt;math&amp;gt;n=1&amp;lt;/math&amp;gt;, this is the univariate case. Assume that &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;. Due to the fundamental theorem of algebra, any polynomial &amp;lt;math&amp;gt;f(x)&amp;lt;/math&amp;gt; of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; must have at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; roots, thus &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[f(r)=0]\le\frac{d}{|S|}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
;Induction hypothesis&#039;&#039;&#039;:&#039;&#039;&#039; &lt;br /&gt;
Assume the theorem holds for any &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-variate polynomials for all &amp;lt;math&amp;gt;m&amp;lt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
;Induction step&#039;&#039;&#039;:&#039;&#039;&#039;&lt;br /&gt;
For any &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial &amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt;  of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, we write &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; as&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)=\sum_{i=0}^kx_n^{i}f_i(x_1,x_2,\ldots,x_{n-1})&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; is the highest degree of &amp;lt;math&amp;gt;x_n&amp;lt;/math&amp;gt;, which means the degree of &amp;lt;math&amp;gt;f_k&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;d-k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_k\not\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In particular, we write &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; as a sum of two parts:&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)=x_n^k f_k(x_1,x_2,\ldots,x_{n-1})+\bar{f}(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt;,&lt;br /&gt;
where both &amp;lt;math&amp;gt;f_k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\bar{f}&amp;lt;/math&amp;gt; are polynomials, such that &lt;br /&gt;
* &amp;lt;math&amp;gt;f_k\not\equiv 0&amp;lt;/math&amp;gt; is as above, whose degree is  at most &amp;lt;math&amp;gt;d-k&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\bar{f}(x_1,x_2,\ldots,x_n)=\sum_{i=0}^{k-1}x_n^i f_i(x_1,x_2,\ldots,x_{n-1})&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;\bar{f}(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;x_n^{k}&amp;lt;/math&amp;gt; factor in any term.&lt;br /&gt;
&lt;br /&gt;
By the law of total probability, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\Pr[f(r_1,r_2,\ldots,r_n)=0]\\&lt;br /&gt;
=&lt;br /&gt;
&amp;amp;\Pr[f(\vec{r})=0\mid f_k(r_1,r_2,\ldots,r_{n-1})=0]\cdot\Pr[f_k(r_1,r_2,\ldots,r_{n-1})=0]\\&lt;br /&gt;
&amp;amp;+\Pr[f(\vec{r})=0\mid f_k(r_1,r_2,\ldots,r_{n-1})\neq0]\cdot\Pr[f_k(r_1,r_2,\ldots,r_{n-1})\neq0].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that &amp;lt;math&amp;gt;f_k(r_1,r_2,\ldots,r_{n-1})&amp;lt;/math&amp;gt; is a polynomial on &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; variables of degree &amp;lt;math&amp;gt;d-k&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;f_k\not\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the induction hypothesis, we have &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(*)&lt;br /&gt;
&amp;amp;\qquad&lt;br /&gt;
&amp;amp;\Pr[f_k(r_1,r_2,\ldots,r_{n-1})=0]\le\frac{d-k}{|S|}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Now we look at the case conditioning on &amp;lt;math&amp;gt;f_k(r_1,r_2,\ldots,r_{n-1})\neq0&amp;lt;/math&amp;gt;. Recall that &amp;lt;math&amp;gt;\bar{f}(x_1,\ldots,x_n)&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;x_n^k&amp;lt;/math&amp;gt; factor in any term, thus the condition &amp;lt;math&amp;gt;f_k(r_1,r_2,\ldots,r_{n-1})\neq0&amp;lt;/math&amp;gt; guarantees that &lt;br /&gt;
:&amp;lt;math&amp;gt;f(r_1,\ldots,r_{n-1},x_n)=x_n^k f_k(r_1,r_2,\ldots,r_{n-1})+\bar{f}(r_1,r_2,\ldots,r_{n-1},x_n)=g_{r_1,\ldots,r_{n-1}}(x_n)&amp;lt;/math&amp;gt;&lt;br /&gt;
is a nonzero univariate polynomial of &amp;lt;math&amp;gt;x_n&amp;lt;/math&amp;gt; such that the degree of &amp;lt;math&amp;gt;g_{r_1,\ldots,r_{n-1}}(x_n)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g_{r_1,\ldots,r_{n-1}}\not\equiv 0&amp;lt;/math&amp;gt;, for which we already known that the probability &amp;lt;math&amp;gt;g_{r_1,\ldots,r_{n-1}}(r_n)=0&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;\frac{k}{|S|}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(**)&lt;br /&gt;
&amp;amp;\qquad&lt;br /&gt;
&amp;amp;\Pr[f(\vec{r})=0\mid f_k(r_1,r_2,\ldots,r_{n-1})\neq0]=\Pr[g_{r_1,\ldots,r_{n-1}}(r_n)=0\mid f_k(r_1,r_2,\ldots,r_{n-1})\neq0]\le\frac{k}{|S|}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
Substituting both &amp;lt;math&amp;gt;(*)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(**)&amp;lt;/math&amp;gt; back in the total probability, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[f(r_1,r_2,\ldots,r_n)=0]&lt;br /&gt;
\le\frac{d-k}{|S|}+\frac{k}{|S|}=\frac{d}{|S|},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which proves the theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Detecting perfect matching ==&lt;br /&gt;
&lt;br /&gt;
TBA&lt;br /&gt;
&lt;br /&gt;
;Edmonds matrix&lt;br /&gt;
&lt;br /&gt;
= Fingerprinting =&lt;br /&gt;
The polynomial identity testing algorithm in the Schwartz-Zippel theorem can be abstracted as the following framework:&lt;br /&gt;
Suppose we want to compare two objects &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;. Instead of comparing them directly, we compute random &#039;&#039;&#039;fingerprints&#039;&#039;&#039; &amp;lt;math&amp;gt;\mathrm{FING}(X)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathrm{FING}(Y)&amp;lt;/math&amp;gt; of them and compare the fingerprints. &lt;br /&gt;
&lt;br /&gt;
The fingerprints has the following properties:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; is a function, meaning that if &amp;lt;math&amp;gt;X= Y&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;\mathrm{FING}(X)=\mathrm{FING}(Y)&amp;lt;/math&amp;gt;.&lt;br /&gt;
* It is much easier to compute and compare the fingerprints.&lt;br /&gt;
* Ideally, the domain of fingerprints is much smaller than the domain of original objects, so storing and comparing fingerprints are easy. This means the fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; cannot be an injection (one-to-one mapping), so it&#039;s possible that different &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; are mapped to the same fingerprint. We resolve this by making fingerprint function &#039;&#039;randomized&#039;&#039;, and for &amp;lt;math&amp;gt;X\neq Y&amp;lt;/math&amp;gt;, we want the probability &amp;lt;math&amp;gt;\Pr[\mathrm{FING}(X)=\mathrm{FING}(Y)]&amp;lt;/math&amp;gt; to be small.&lt;br /&gt;
&lt;br /&gt;
In Schwartz-Zippel theorem, the objects to compare are polynomials from &amp;lt;math&amp;gt;\mathbb{F}[x_1,\ldots,x_n]&amp;lt;/math&amp;gt;. Given a polynomial &amp;lt;math&amp;gt;f\in \mathbb{F}[x_1,\ldots,x_n]&amp;lt;/math&amp;gt;, its fingerprint is computed as &amp;lt;math&amp;gt;\mathrm{FING}(f)=f(r_1,\ldots,r_n)&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;r_i&amp;lt;/math&amp;gt; chosen independently and uniformly at random from some fixed set &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
With this generic framework, for various identity testing problems, we may design different fingerprints &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Communication protocols for Equality ==&lt;br /&gt;
Now consider again the communication model where the two players Alice with a private input &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt; and Bob with a private input &amp;lt;math&amp;gt;y\in\{0,1\}^n&amp;lt;/math&amp;gt; together compute a function &amp;lt;math&amp;gt;f(x,y)&amp;lt;/math&amp;gt; by running a communication protocol. &lt;br /&gt;
&lt;br /&gt;
We still consider the communication protocols for the equality function EQ&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathrm{EQ}(x,y)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{if } x=y,\\&lt;br /&gt;
0&amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
With the language of fingerprinting, this communication problem can be solved by the following generic scheme:&lt;br /&gt;
{{Theorem|Communication protocol for EQ by fingerprinting|&lt;br /&gt;
&#039;&#039;&#039;Bob does&#039;&#039;&#039;:&lt;br /&gt;
:* choose a random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; and compute the fingerprint of her input &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt;;&lt;br /&gt;
:* sends both the description of &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; and the value of &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt; to Alice;&lt;br /&gt;
&#039;&#039;&#039;Upon receiving&#039;&#039;&#039; the description of &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; and the value of &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt;, &#039;&#039;&#039;Alice does&#039;&#039;&#039;:&lt;br /&gt;
:* computes &amp;lt;math&amp;gt;\mathrm{FING}(x)&amp;lt;/math&amp;gt; and check whether &amp;lt;math&amp;gt;\mathrm{FING}(x)=\mathrm{FING}(y)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
In this way we have a randomized communication protocol for the equality function EQ with false positive. The communication cost as well as the error probability are reduced to the question of how to design this random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; to guarantee:&lt;br /&gt;
# A random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; can be described succinctly.&lt;br /&gt;
# The range of &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; is small, so the fingerprints are succinct.&lt;br /&gt;
# If &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;, the probability &amp;lt;math&amp;gt;\Pr[\mathrm{FING}(x)=\mathrm{FING}(y)]&amp;lt;/math&amp;gt; is small.&lt;br /&gt;
&lt;br /&gt;
=== Fingerprinting by PIT===&lt;br /&gt;
As before, we can define the fingerprint function as: for any bit-string &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt;,  its random fingerprint is &amp;lt;math&amp;gt;\mathrm{FING}(x)=\sum_{i=1}^n x_i r^{i}&amp;lt;/math&amp;gt;, where the additions and multiplications are defined over a finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen uniformly at random from &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; is some suitable prime which can be represented in &amp;lt;math&amp;gt;\Theta(\log n)&amp;lt;/math&amp;gt; bits. More specifically, we can choose &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; to be any prime from the interval &amp;lt;math&amp;gt;[n^2, 2n^2]&amp;lt;/math&amp;gt;. Due to Chebyshev&#039;s theorem, such prime must exist. &lt;br /&gt;
&lt;br /&gt;
As we have shown before, it takes &amp;lt;math&amp;gt;O(\log p)=O(\log n)&amp;lt;/math&amp;gt; bits to represent &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt; and to describe the random function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; (since it a random function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; from this family is uniquely identified by a random &amp;lt;math&amp;gt;r\in\mathbb{Z}_p&amp;lt;/math&amp;gt;, which can be represented within &amp;lt;math&amp;gt;\log p=O(\log n)&amp;lt;/math&amp;gt; bits). And it follows easily from the fundamental theorem of algebra that for any distinct &amp;lt;math&amp;gt;x, y\in\{0,1\}^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\mathrm{FING}(x)=\mathrm{FING}(y)] \le \frac{n-1}{p}\le \frac{1}{n}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Fingerprinting by randomized checksum===&lt;br /&gt;
Now we consider a new fingerprint function: We treat each input string &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt; as the binary representation of a number, and let &amp;lt;math&amp;gt;\mathrm{FING}(x)=x\bmod p&amp;lt;/math&amp;gt; for some random prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; chosen from &amp;lt;math&amp;gt;[k]=\{0,1,\ldots,k-1\}&amp;lt;/math&amp;gt;, for some &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; to be specified later. &lt;br /&gt;
&lt;br /&gt;
Now a random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; can be uniquely identified by this random prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. The new communication protocol for EQ with this fingerprint is as follows:&lt;br /&gt;
{{Theorem|Communication protocol for EQ by random checksum|&lt;br /&gt;
&#039;&#039;&#039;Bob does&#039;&#039;&#039;:&lt;br /&gt;
:for some parameter &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; (to be specified), &lt;br /&gt;
:* choose a prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt; uniformly at random;&lt;br /&gt;
:* send &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x\bmod p&amp;lt;/math&amp;gt; to Alice;&lt;br /&gt;
&#039;&#039;&#039;Upon receiving&#039;&#039;&#039; &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x\bmod p&amp;lt;/math&amp;gt;, &#039;&#039;&#039;Alice does&#039;&#039;&#039;:&lt;br /&gt;
:* check whether &amp;lt;math&amp;gt;x\bmod p=y\bmod p&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The number of bits to be communicated is obviously &amp;lt;math&amp;gt;O(\log k)&amp;lt;/math&amp;gt;. When &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;, we want to upper bound the error probability &amp;lt;math&amp;gt;\Pr[x\bmod p=y\bmod p]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose without loss of generality &amp;lt;math&amp;gt;x&amp;gt;y&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;z=x-y&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;z&amp;lt;2^n&amp;lt;/math&amp;gt; since &amp;lt;math&amp;gt;x,y\in[2^n]&amp;lt;/math&amp;gt;, and &lt;br /&gt;
&amp;lt;math&amp;gt;z\neq 0&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;.  It holds that &amp;lt;math&amp;gt;x\equiv y\pmod p&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;p\mid z&amp;lt;/math&amp;gt;. Therefore, we only need to upper bound the probability&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[z\bmod p=0]&amp;lt;/math&amp;gt;&lt;br /&gt;
for an arbitrarily fixed &amp;lt;math&amp;gt;0&amp;lt;z&amp;lt;2^n&amp;lt;/math&amp;gt;, and a uniform random prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The probability &amp;lt;math&amp;gt;\Pr[z\bmod p=0]&amp;lt;/math&amp;gt; is computed directly as&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[z\bmod p=0]\le\frac{\mbox{the number of prime divisors of }z}{\mbox{the number of primes in }[k]}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For the numerator, any positive &amp;lt;math&amp;gt;z&amp;lt;2^n&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; prime factors. To see this, by contradiction assume that &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; has more than &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; prime factors. Note that any prime number is at least 2. Then &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; must be greater than &amp;lt;math&amp;gt;2^n&amp;lt;/math&amp;gt;, contradicting the fact that &amp;lt;math&amp;gt;z&amp;lt;2^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For the denominator, we need to lower bound the number of primes in &amp;lt;math&amp;gt;[k]&amp;lt;/math&amp;gt;. This is given by the celebrated [http://en.wikipedia.org/wiki/Prime_number_theorem &#039;&#039;&#039;Prime Number Theorem (PNT)&#039;&#039;&#039;].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Prime Number Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\pi(k)&amp;lt;/math&amp;gt; denote the number of primes less than &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\pi(k)\sim\frac{k}{\ln k}&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;k\rightarrow\infty&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Therefore, by choosing &amp;lt;math&amp;gt;k=2n^2\ln n&amp;lt;/math&amp;gt;, we have that for a &amp;lt;math&amp;gt;0&amp;lt;z&amp;lt;2^n&amp;lt;/math&amp;gt;, and a random prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[z\bmod p=0]\le\frac{n}{\pi(k)}\sim\frac{1}{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
which means the for any inputs &amp;lt;math&amp;gt;x,y\in\{0,1\}^n&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;, then the false positive is bounded as&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathrm{FING}(x)=\mathrm{FING}(y)]\le\Pr[|x-y|\bmod p=0]\le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Moreover, by this choice of parameter &amp;lt;math&amp;gt;k=2n^2\ln n&amp;lt;/math&amp;gt;, the communication complexity of the protocol is bounded by &amp;lt;math&amp;gt;O(\log k)=O(\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Pattern matching ==&lt;br /&gt;
Consider the following problem of pattern matching, which has nothing to do with communication complexity.&lt;br /&gt;
&lt;br /&gt;
*Input: a string &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt; and a &amp;quot;pattern&amp;quot; &amp;lt;math&amp;gt;y\in\{0,1\}^m&amp;lt;/math&amp;gt;.&lt;br /&gt;
*Determine whether the pattern &amp;lt;math&amp;gt;y&amp;lt;/math&amp;gt; is a contiguous  substring of &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;. Usually, we are also asked to find the location of the substring.&lt;br /&gt;
&lt;br /&gt;
A naive algorithm trying every possible match runs in &amp;lt;math&amp;gt;O(nm)&amp;lt;/math&amp;gt; time. The more sophisticated KMP algorithm inspired by automaton theory runs in &amp;lt;math&amp;gt;O(n+m)&amp;lt;/math&amp;gt; time.&lt;br /&gt;
&lt;br /&gt;
A simple randomized algorithm, due to Karp and Rabin, uses the idea of fingerprinting and also runs in &amp;lt;math&amp;gt;O(n + m)&amp;lt;/math&amp;gt; time.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X(j)=x_jx_{j+1}\cdots x_{j+m-1}&amp;lt;/math&amp;gt; denote the substring of &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; of length &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; starting at position &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Algorithm (Karp-Rabin)|&lt;br /&gt;
:pick a random prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt;;&lt;br /&gt;
:&#039;&#039;&#039;for&#039;&#039;&#039; &amp;lt;math&amp;gt;j = 1&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;n -m + 1&amp;lt;/math&amp;gt; &#039;&#039;&#039;do&#039;&#039;&#039;&lt;br /&gt;
::&#039;&#039;&#039;if&#039;&#039;&#039; &amp;lt;math&amp;gt;X(j)\bmod p = y \bmod p&amp;lt;/math&amp;gt; then report a match;&lt;br /&gt;
:&#039;&#039;&#039;return&#039;&#039;&#039; &amp;quot;no match&amp;quot;;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
So the algorithm just compares the &amp;lt;math&amp;gt;\mathrm{FING}(X(j))&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;, with the same definition of fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; as in the communication protocol for EQ.&lt;br /&gt;
&lt;br /&gt;
By the same analysis, by choosing &amp;lt;math&amp;gt;k=n^2m\ln (n^2m)&amp;lt;/math&amp;gt;, the probability of a single false match is &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X(j)\bmod p=y\bmod p\mid X(j)\neq y ]=O\left(\frac{1}{n^2}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the union bound, the probability that a false match occurs is &amp;lt;math&amp;gt;O\left(\frac{1}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The algorithm runs in linear time if we assume that we can compute &amp;lt;math&amp;gt;X(j)\bmod p &amp;lt;/math&amp;gt; for each &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; in constant time. This outrageous assumption can be made realistic by the following observation.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathrm{FING}(a)=a\bmod p&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;\mathrm{FING}(X(j+1))\equiv2(\mathrm{FING}(X(j))-2^{m-1}x_j)+x_{j+m}\pmod p\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| It holds that &lt;br /&gt;
:&amp;lt;math&amp;gt;X(j+1)=2(X(j)-2^{m-1}x_j)+x_{j+m}\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
So the equation holds on the finite field modulo &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Due to this lemma, each fingerprint &amp;lt;math&amp;gt;\mathrm{FING}(X(j))&amp;lt;/math&amp;gt; can be computed in an incremental way, each in constant time. The running time of the algorithm is &amp;lt;math&amp;gt;O(n+m)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Checking distinctness ==&lt;br /&gt;
Consider the following problem of &#039;&#039;&#039;checking distinctness&#039;&#039;&#039;:&lt;br /&gt;
*Given a sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, check whether every element of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; appears &#039;&#039;&#039;exactly&#039;&#039;&#039; once.&lt;br /&gt;
&lt;br /&gt;
Obviously this problem can be solved in linear time and linear space (in addition to the space for storing the input) by maintaining a &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-bit vector that indicates which numbers among &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; have appeared.&lt;br /&gt;
&lt;br /&gt;
When this &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is enormously large, &amp;lt;math&amp;gt;\Omega(n)&amp;lt;/math&amp;gt; space cost is too expensive. We wonder whether we could solve this problem with a space cost (in addition to the space for storing the input) much less than &amp;lt;math&amp;gt;O(n)&amp;lt;/math&amp;gt;. This can be done by fingerprinting if we tolerate a certain degree of inaccuracy.&lt;br /&gt;
&lt;br /&gt;
We consider the following more generalized problem, &#039;&#039;&#039;checking identity of multisets&#039;&#039;&#039;:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; two multisets &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots, a_n\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B=\{b_1,b_2,\ldots, b_n\}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;a_1,a_2,\ldots,b_1,b_2,\ldots,b_n\in \{1,2,\ldots,n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt; (multiset equivalence).&lt;br /&gt;
&lt;br /&gt;
Here for a &#039;&#039;&#039;multiset&#039;&#039;&#039; &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots, a_n\}&amp;lt;/math&amp;gt;, its elements &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; are not necessarily distinct. The &#039;&#039;&#039;multiplicity&#039;&#039;&#039; of an element &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; in a multiset &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is the number of times &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; appears in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. Two multisets &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are equivalent if they contain the same set of elements and the multiplicities of every element in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are equal. &lt;br /&gt;
&lt;br /&gt;
Obviously the above problem of checking distinctness can be treated as a special case of checking identity of multisets: by checking the identity of the multiset &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and set &amp;lt;math&amp;gt;\{1,2,\ldots, n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The following fingerprinting function for multisets was introduced by Lipton for solving multiset identity testing.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fingerprint for multiset|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; be a uniform random prime chosen from the interval &amp;lt;math&amp;gt;[(n\log n)^2,2(n\log n)^2]&amp;lt;/math&amp;gt;. By Chebyshev&#039;s theorem, such prime must exist. And consider the the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p=[p]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
:Given a multiset &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots,a_n\}&amp;lt;/math&amp;gt;, we define a univariate polynomial &amp;lt;math&amp;gt;f_A\in\mathbb{Z}_p[x]&amp;lt;/math&amp;gt; over the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
::&amp;lt;math&amp;gt;f_A(x)=\prod_{i=1}^n(x-a_i)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; are defined over the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
:We then define the random fingerprinting function as:&lt;br /&gt;
::&amp;lt;math&amp;gt;\mathrm{FING}(A)=f_A(r)=\prod_{i=1}^n(r-a_i)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen uniformly at random from &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Since all computations of &amp;lt;math&amp;gt;\mathrm{FING}(A)=\prod_{i=1}^n(r-a_i)&amp;lt;/math&amp;gt; are over the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, the space cost for computing the fingerprint &amp;lt;math&amp;gt;\mathrm{FING}(A)&amp;lt;/math&amp;gt; is only &amp;lt;math&amp;gt;O(\log p)=O(\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Moreover, the fingerprinting function &amp;lt;math&amp;gt;\mathrm{FING}(A)&amp;lt;/math&amp;gt; is invariant under permutation of elements of the multiset &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots,a_n\}&amp;lt;/math&amp;gt;, thus it is indeed a function of multisets (meaning every multiset has only one fingerprint). Therefore, if &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;\mathrm{FING}(A)=\mathrm{FING}(B)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For two distinct multisets &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;, it is possible that &amp;lt;math&amp;gt;\mathrm{FING}(A)=\mathrm{FING}(B)&amp;lt;/math&amp;gt;, but the following theorem due to Lipton bounds this error probability of false positive.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Lipton 1989)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots,a_n\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B=\{b_1,b_2,\ldots,b_n\}&amp;lt;/math&amp;gt; be two multisets whose elements are from &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[\mathrm{FING}(A)= \mathrm{FING}(B)]=O\left(\frac{1}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| &lt;br /&gt;
Let &amp;lt;math&amp;gt;\tilde{f}_A(x)=\prod_{i=1}^n(x-a_i)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\tilde{f}_B(x)=\prod_{i=1}^n(x-b_i)&amp;lt;/math&amp;gt; be two univariate polynomials defined over reals &amp;lt;math&amp;gt;\mathbb{R}&amp;lt;/math&amp;gt;. Note that in contrast to &amp;lt;math&amp;gt;f_A(x)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_B(x)&amp;lt;/math&amp;gt;,  the &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;\tilde{f}_A(x), \tilde{f}_B(x)&amp;lt;/math&amp;gt; do not modulo &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. It is easy to verify that the polynomials &amp;lt;math&amp;gt;\tilde{f}_A(x), \tilde{f}_B(x)&amp;lt;/math&amp;gt; have the following properties:&lt;br /&gt;
*&amp;lt;math&amp;gt;\tilde{f}_A\equiv \tilde{f}_B&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt;. Here &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt; means the multiset equivalence.&lt;br /&gt;
*By the properties of finite field, for any value &amp;lt;math&amp;gt;r\in\mathbb{Z}_p&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;f_A(r)=\tilde{f}_A(r)\bmod p&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_B(r)=\tilde{f}_B(r)\bmod p&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, assuming that &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;, we must have &amp;lt;math&amp;gt;\tilde{f}_A(x)\not\equiv \tilde{f}_B(x)&amp;lt;/math&amp;gt;. Then by the law of total probability:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathrm{FING}(A)= \mathrm{FING}(B)]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[f_A(r)=f_B(r)\mid f_A\not\equiv f_B\right]\Pr[f_A\not\equiv f_B]\\&lt;br /&gt;
&amp;amp;\quad\,\,+\Pr\left[f_A(r)=f_B(r)\mid f_A\equiv f_B\right]\Pr[f_A\equiv f_B]\\&lt;br /&gt;
&amp;amp;\le \Pr\left[f_A(r)=f_B(r)\mid f_A\not\equiv f_B\right]+\Pr[f_A\equiv f_B].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that the degrees of &amp;lt;math&amp;gt;f_A,f_B&amp;lt;/math&amp;gt; are at most &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen uniformly from &amp;lt;math&amp;gt;[p]&amp;lt;/math&amp;gt;. By the Schwartz-Zippel theorem for univariate polynomials, the first probability&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[f_A(r)=f_B(r)\mid f_A\not\equiv f_B\right]\le \frac{n}{p}=o\left(\frac{1}{n}\right),&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
since &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; is chosen from the interval &amp;lt;math&amp;gt;[(n\log n)^2,2(n\log n)^2]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For the second probability &amp;lt;math&amp;gt;\Pr[f_A\equiv f_B]&amp;lt;/math&amp;gt;, recall that &amp;lt;math&amp;gt;\tilde{f}_A\not\equiv \tilde{f}_B&amp;lt;/math&amp;gt;, therefore there is at least a non-zero coefficient &amp;lt;math&amp;gt;c\le n^n&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;\tilde{f}_A-\tilde{f}_B&amp;lt;/math&amp;gt;. The event &amp;lt;math&amp;gt;f_A\equiv f_B&amp;lt;/math&amp;gt; occurs only if &amp;lt;math&amp;gt;c\bmod p=0&amp;lt;/math&amp;gt;, which means&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[f_A\equiv f_B]&lt;br /&gt;
&amp;amp;\le \Pr[c\bmod p=0]\\&lt;br /&gt;
&amp;amp;=\frac{\text{number of prime factors of }c}{\text{number of primes in }[(n\log n)^2,2(n\log n)^2]}\\&lt;br /&gt;
&amp;amp;\le \frac{n\log_2n}{\pi(2(n\log n)^2)-\pi((n\log n)^2)}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By the prime number theorem, &amp;lt;math&amp;gt;\pi(N)\rightarrow \frac{N}{\ln N}&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;N\to\infty&amp;lt;/math&amp;gt;. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[f_A\equiv f_B]=O\left(\frac{n\log n}{n^2\log n}\right)=O\left(\frac{1}{n}\right).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Combining everything together, we have&lt;br /&gt;
&amp;lt;math&amp;gt;\Pr[\mathrm{FING}(A)= \mathrm{FING}(B)]=O\left(\frac{1}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13894</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13894"/>
		<updated>2026-09-02T16:01:20Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-cuts-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) ([http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-en.pdf AI-generated lecture note], [http://tcs.nju.edu.cn/slides/aa2026/AA26-note-fingerprinting-zh.pdf AI生成讲义])&lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13893</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13893"/>
		<updated>2026-09-02T09:19:28Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Finite_Field_Basics&amp;diff=13892</id>
		<title>高级算法 (Fall 2026)/Finite Field Basics</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Finite_Field_Basics&amp;diff=13892"/>
		<updated>2026-09-02T09:19:10Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;=Field= Let &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; be a set, &amp;#039;&amp;#039;&amp;#039;closed&amp;#039;&amp;#039;&amp;#039; under binary operations &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; (addition) and &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; (multiplication). It gives us the following algebraic structures if the corresponding set of axioms are satisfied. {|class=&amp;quot;wikitable&amp;quot; !colspan=&amp;quot;7&amp;quot;|Structures  !Axioms !Operations |- |rowspan=&amp;quot;9&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&amp;#039;&amp;#039;&amp;#039;&amp;#039;&amp;#039;field&amp;#039;&amp;#039;&amp;#039;&amp;#039;&amp;#039; |rowspan=&amp;quot;8&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&amp;#039;&amp;#039;&amp;#039;&amp;#039;&amp;#039;commutative&amp;lt;br&amp;gt;rin...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Field=&lt;br /&gt;
Let &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; be a set, &#039;&#039;&#039;closed&#039;&#039;&#039; under binary operations &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; (addition) and &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; (multiplication). It gives us the following algebraic structures if the corresponding set of axioms are satisfied.&lt;br /&gt;
{|class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!colspan=&amp;quot;7&amp;quot;|Structures &lt;br /&gt;
!Axioms&lt;br /&gt;
!Operations&lt;br /&gt;
|-&lt;br /&gt;
|rowspan=&amp;quot;9&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;field&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
|rowspan=&amp;quot;8&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;commutative&amp;lt;br&amp;gt;ring&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
|rowspan=&amp;quot;7&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;ring&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
|rowspan=&amp;quot;4&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;abelian&amp;lt;br&amp;gt;group&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
|rowspan=&amp;quot;3&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;group&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
| rowspan=&amp;quot;2&amp;quot; style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;monoid&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
|style=&amp;quot;background-color:#ffffcc;text-align:center;&amp;quot;|&#039;&#039;&#039;&#039;&#039;semigroup&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
|1. &#039;&#039;&#039;Addition&#039;&#039;&#039; is &#039;&#039;&#039;associative&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x,y,z\in S, (x+y)+z= x+(y+z).&amp;lt;/math&amp;gt;&lt;br /&gt;
|rowspan=&amp;quot;4&amp;quot; style=&amp;quot;text-align:center;&amp;quot;|&amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|&lt;br /&gt;
|2. Existence of &#039;&#039;&#039;additive identity 0&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x\in S, x+0= 0+x=x.&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|colspan=&amp;quot;2&amp;quot;|&lt;br /&gt;
|3. Everyone has an &#039;&#039;&#039;additive inverse&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x\in S, \exists -x\in S, \text{ s.t. } x+(-x)= (-x)+x=0.&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|colspan=&amp;quot;3&amp;quot;|&lt;br /&gt;
|4. &#039;&#039;&#039;Addition&#039;&#039;&#039; is &#039;&#039;&#039;commutative&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x,y\in S, x+y= y+x.&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|colspan=&amp;quot;4&amp;quot; rowspan=&amp;quot;3&amp;quot;|&lt;br /&gt;
|5. Multiplication &#039;&#039;&#039;distributes&#039;&#039;&#039; over addition: &amp;lt;math&amp;gt;\forall x,y,z\in S, x\cdot(y+z)= x\cdot y+x\cdot z&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(y+z)\cdot x= y\cdot x+z\cdot x.&amp;lt;/math&amp;gt;&lt;br /&gt;
|style=&amp;quot;text-align:center;&amp;quot;|&amp;lt;math&amp;gt;+,\cdot&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|6. &#039;&#039;&#039;Multiplication&#039;&#039;&#039; is &#039;&#039;&#039;associative&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x,y,z\in S, (x\cdot y)\cdot z= x\cdot (y\cdot z).&amp;lt;/math&amp;gt;&lt;br /&gt;
|rowspan=&amp;quot;4&amp;quot; style=&amp;quot;text-align:center;&amp;quot;|&amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|7. Existence of &#039;&#039;&#039;multiplicative identity 1&#039;&#039;&#039;:  &amp;lt;math&amp;gt;\forall x\in S, x\cdot 1= 1\cdot x=x.&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|colspan=&amp;quot;5&amp;quot;|&lt;br /&gt;
|8. &#039;&#039;&#039;Multiplication&#039;&#039;&#039; is &#039;&#039;&#039;commutative&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x,y\in S, x\cdot y= y\cdot x.&amp;lt;/math&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
|colspan=&amp;quot;6&amp;quot;|&lt;br /&gt;
|9. Every non-zero element has a &#039;&#039;&#039;multiplicative inverse&#039;&#039;&#039;: &amp;lt;math&amp;gt;\forall x\in S\setminus\{0\}, \exists x^{-1}\in S, \text{ s.t. } x\cdot x^{-1}= x^{-1}\cdot x=1.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
The semigroup, monoid, group and abelian group are given by &amp;lt;math&amp;gt;(S,+)&amp;lt;/math&amp;gt;, and the ring, commutative ring, and field are given by &amp;lt;math&amp;gt;(S,+,\cdot)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Examples:&lt;br /&gt;
* &#039;&#039;&#039;Infinite fields&#039;&#039;&#039;: &amp;lt;math&amp;gt;\mathbb{Q}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathbb{R}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathbb{C}&amp;lt;/math&amp;gt; are fields. The integer set &amp;lt;math&amp;gt;\mathbb{Z}&amp;lt;/math&amp;gt; is a commutative ring but is not a field.&lt;br /&gt;
* &#039;&#039;&#039;Finite fields&#039;&#039;&#039;: Finite fields are called &#039;&#039;&#039;Galois fields&#039;&#039;&#039;. The number of elements of a finite field is called its &#039;&#039;&#039;order&#039;&#039;&#039;. A finite field of order &amp;lt;math&amp;gt;q&amp;lt;/math&amp;gt;, is usually denoted as &amp;lt;math&amp;gt;\mathsf{GF}(q)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;\mathbb{F}_q&amp;lt;/math&amp;gt;.&lt;br /&gt;
** &#039;&#039;&#039;Prime field&#039;&#039;&#039; &amp;lt;math&amp;gt;{\mathbb{Z}_p}&amp;lt;/math&amp;gt;: For any integer &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,  &amp;lt;math&amp;gt;\mathbb{Z}_n=\{0,1,\ldots,n-1\}&amp;lt;/math&amp;gt; under modulo-&amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; addition &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; and multiplication &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; forms a commutative ring. It is called &#039;&#039;&#039;quotient ring&#039;&#039;&#039;, and is sometimes denoted as &amp;lt;math&amp;gt;\mathbb{Z}/n\mathbb{Z}&amp;lt;/math&amp;gt;.  In particular, for &#039;&#039;&#039;prime&#039;&#039;&#039; &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; is a field. This can be verified by [http://en.wikipedia.org/wiki/Fermat%27s_little_theorem Fermat&#039;s little theorem].&lt;br /&gt;
** &#039;&#039;&#039;Boolean arithmetics&#039;&#039;&#039; &amp;lt;math&amp;gt;\mathsf{GF}(2)&amp;lt;/math&amp;gt;: The finite field of order 2 &amp;lt;math&amp;gt;\mathsf{GF}(2)&amp;lt;/math&amp;gt; contains only two elements 0 and 1, with bit-wise XOR as addition and bit-wise AND as multiplication. &amp;lt;math&amp;gt;\mathsf{GF}(2^n)&amp;lt;/math&amp;gt;&lt;br /&gt;
** Other examples: There are other examples of finite fields, for instance &amp;lt;math&amp;gt;\{a+bi\mid a,b\in \mathbb{Z}_3\}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;i=\sqrt{-1}&amp;lt;/math&amp;gt;. This field is isomorphic to &amp;lt;math&amp;gt;\mathsf{GF}(9)&amp;lt;/math&amp;gt;. In fact, the following theorem holds for finite fields of given order.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:A finite field of order &amp;lt;math&amp;gt;q&amp;lt;/math&amp;gt; exists if and only if &amp;lt;math&amp;gt;q=p^k&amp;lt;/math&amp;gt; for some prime number &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; and positive integer &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. Moreover, all fields of a given order are isomorphic.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=Polynomial over a field=&lt;br /&gt;
Given a field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;, the &#039;&#039;&#039;polynomial ring&#039;&#039;&#039; &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; consists of all polynomials in the variable &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; with coefficients in &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. Addition and multiplication of polynomials are naturally defined by applying the distributive law and combining like terms.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition (polynomial ring)|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; is a ring.&lt;br /&gt;
}}&lt;br /&gt;
The &#039;&#039;&#039;degree&#039;&#039;&#039; &amp;lt;math&amp;gt;\mathrm{deg}(f)&amp;lt;/math&amp;gt; of a polynomial &amp;lt;math&amp;gt;f\in \mathbb{F}[x]&amp;lt;/math&amp;gt; is the exponent on the &#039;&#039;&#039;leading term&#039;&#039;&#039;, the term with a nonzero coefficient that has the largest exponent.&lt;br /&gt;
&lt;br /&gt;
Because &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; is a ring, we cannot do division the way we do it in a field like &amp;lt;math&amp;gt;\mathbb{R}&amp;lt;/math&amp;gt;, but we can do division the way we do it in a ring like &amp;lt;math&amp;gt;\mathbb{Z}&amp;lt;/math&amp;gt;, leaving a &#039;&#039;&#039;remainder&#039;&#039;&#039;. The equivalent of the &#039;&#039;&#039;integer division&#039;&#039;&#039; for &amp;lt;math&amp;gt;\mathbb{Z}&amp;lt;/math&amp;gt; is as follows.&lt;br /&gt;
{{Theorem|Proposition (division for polynomials)|&lt;br /&gt;
:Given a polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and a nonzero polynomial &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt;, there are unique polynomials &amp;lt;math&amp;gt;q&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;f =q\cdot g+r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathrm{deg}(r)&amp;lt;\mathrm{deg}(g)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The proof of this is by induction on &amp;lt;math&amp;gt;\mathrm{deg}(f)&amp;lt;/math&amp;gt;, with the basis &amp;lt;math&amp;gt;\mathrm{deg}(f)&amp;lt;\mathrm{deg}(g)&amp;lt;/math&amp;gt;, in which case the theorem holds trivially by letting &amp;lt;math&amp;gt;q=0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r=f&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
As we turn &amp;lt;math&amp;gt;\mathbb{Z}&amp;lt;/math&amp;gt; (a ring) into a finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; by taking quotients &amp;lt;math&amp;gt;\bmod p&amp;lt;/math&amp;gt;, we can turn a polynomial ring &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; into a finite field by taking &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; modulo a &amp;quot;prime-like&amp;quot; polynomial, using the division of polynomials above.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition (irreducible polynomial)|&lt;br /&gt;
:An &#039;&#039;&#039;irreducible polynomial&#039;&#039;&#039;, or a &#039;&#039;&#039;prime polynomial&#039;&#039;&#039;, is a non-constant polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; that &#039;&#039;cannot&#039;&#039; be factored as &amp;lt;math&amp;gt;f=g\cdot h&amp;lt;/math&amp;gt; for any non-constant polynomials &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;h&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Fingerprinting&amp;diff=13891</id>
		<title>高级算法 (Fall 2026)/Fingerprinting</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)/Fingerprinting&amp;diff=13891"/>
		<updated>2026-09-02T09:18:54Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;=  Checking Matrix Multiplication= The evolution of time complexity &amp;lt;math&amp;gt;O(n^{\omega})&amp;lt;/math&amp;gt; for matrix multiplication.  Let &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; be a feild (you may think of it as the filed &amp;lt;math&amp;gt;\mathbb{Q}&amp;lt;/math&amp;gt; of rational numbers, or the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; of integers modulo prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;). We suppose that each field operation (addition, subtraction, multiplication, division) has u...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=  Checking Matrix Multiplication=&lt;br /&gt;
[[File: matrix_multiplication.png|thumb|360px|right|The evolution of time complexity &amp;lt;math&amp;gt;O(n^{\omega})&amp;lt;/math&amp;gt; for matrix multiplication.]]&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; be a feild (you may think of it as the filed &amp;lt;math&amp;gt;\mathbb{Q}&amp;lt;/math&amp;gt; of rational numbers, or the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; of integers modulo prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;). We suppose that each field operation (addition, subtraction, multiplication, division) has unit cost. This model is called the &#039;&#039;&#039;unit-cost RAM&#039;&#039;&#039; model, which is an ideal abstraction of a computer.&lt;br /&gt;
&lt;br /&gt;
Consider the following problem:&lt;br /&gt;
* &#039;&#039;&#039;Input&#039;&#039;&#039;: Three &amp;lt;math&amp;gt;n\times n&amp;lt;/math&amp;gt; matrices &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; over the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* &#039;&#039;&#039;Output&#039;&#039;&#039;: &amp;quot;yes&amp;quot; if &amp;lt;math&amp;gt;C=AB&amp;lt;/math&amp;gt; and &amp;quot;no&amp;quot; if otherwise.&lt;br /&gt;
&lt;br /&gt;
A naive way to solve this is to multiply &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; and compare the result with &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;. &lt;br /&gt;
The straightforward algorithm for matrix multiplication takes &amp;lt;math&amp;gt;O(n^3)&amp;lt;/math&amp;gt; time, assuming that each arithmetic operation takes unit time.&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Strassen_algorithm Strassen&#039;s algorithm] discovered in 1969 now implemented by many numerical libraries runs in time &amp;lt;math&amp;gt;O(n^{\log_2 7})\approx O(n^{2.81})&amp;lt;/math&amp;gt;. Strassen&#039;s algorithm starts the search for fast matrix multiplication algorithms. The [http://en.wikipedia.org/wiki/Coppersmith%E2%80%93Winograd_algorithm Coppersmith–Winograd algorithm] discovered in 1987 runs in time &amp;lt;math&amp;gt;O(n^{2.376})&amp;lt;/math&amp;gt; but is only faster than Strassens&#039; algorithm on extremely large matrices due to the very large constant coefficient. This has been the best known for decades, until recently Stothers got an &amp;lt;math&amp;gt;O(n^{2.374})&amp;lt;/math&amp;gt; algorithm in his PhD thesis in 2010, and independently Vassilevska Williams got an &amp;lt;math&amp;gt;O(n^{2.373})&amp;lt;/math&amp;gt; algorithm in 2012. Both these improvements are based on generalization of Coppersmith–Winograd algorithm. It is unknown whether the matrix multiplication can be done in time &amp;lt;math&amp;gt;O(n^{2+o(1)})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Freivalds Algorithm ==&lt;br /&gt;
The following is a very simple randomized algorithm due to Freivalds, running in &amp;lt;math&amp;gt;O(n^2)&amp;lt;/math&amp;gt; time:&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Algorithm (Freivalds, 1979)|&lt;br /&gt;
*pick a vector &amp;lt;math&amp;gt;r \in\{0, 1\}^n&amp;lt;/math&amp;gt; uniformly at random;&lt;br /&gt;
*if &amp;lt;math&amp;gt;A(Br) = Cr&amp;lt;/math&amp;gt; then return &amp;quot;yes&amp;quot; else return &amp;quot;no&amp;quot;;&lt;br /&gt;
}}&lt;br /&gt;
The product &amp;lt;math&amp;gt;A(Br)&amp;lt;/math&amp;gt; is computed by first multiplying &amp;lt;math&amp;gt;Br&amp;lt;/math&amp;gt; and then &amp;lt;math&amp;gt;A(Br)&amp;lt;/math&amp;gt;.&lt;br /&gt;
The running time of Freivalds algorithm is &amp;lt;math&amp;gt;O(n^2)&amp;lt;/math&amp;gt; because the algorithm computes 3 matrix-vector multiplications. &lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;A(Br) = Cr&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;r \in\{0, 1\}^n&amp;lt;/math&amp;gt;, thus the algorithm will return a &amp;quot;yes&amp;quot; for any positive instance (&amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt;). &lt;br /&gt;
But if &amp;lt;math&amp;gt;AB \neq C&amp;lt;/math&amp;gt; then the algorithm will make a mistake if it chooses such an &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;ABr = Cr&amp;lt;/math&amp;gt;. However, the following lemma states that the probability of this event is bounded.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:If &amp;lt;math&amp;gt;AB\neq C&amp;lt;/math&amp;gt; then for a uniformly random &amp;lt;math&amp;gt;r \in\{0, 1\}^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[ABr = Cr]\le \frac{1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Let &amp;lt;math&amp;gt;D=AB-C&amp;lt;/math&amp;gt;. The event &amp;lt;math&amp;gt;ABr=Cr&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;Dr=0&amp;lt;/math&amp;gt;. It is then sufficient to show that for a &amp;lt;math&amp;gt;D\neq \boldsymbol{0}&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;\Pr[Dr = \boldsymbol{0}]\le \frac{1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Since &amp;lt;math&amp;gt;D\neq \boldsymbol{0}&amp;lt;/math&amp;gt;, it must have at least one non-zero entry. Suppose that &amp;lt;math&amp;gt;D_{ij}\neq 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We assume the event that &amp;lt;math&amp;gt;Dr=\boldsymbol{0}&amp;lt;/math&amp;gt;. In particular, the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th entry of &amp;lt;math&amp;gt;Dr&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;(Dr)_{i}=\sum_{k=1}^n D_{ik}r_k=0.&amp;lt;/math&amp;gt; &lt;br /&gt;
The &amp;lt;math&amp;gt;r_j&amp;lt;/math&amp;gt; can be calculated by&lt;br /&gt;
:&amp;lt;math&amp;gt;r_j=-\frac{1}{D_{ij}}\sum_{k\neq j}^n D_{ik}r_k.&amp;lt;/math&amp;gt;&lt;br /&gt;
Once all other entries &amp;lt;math&amp;gt;r_k&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;k\neq j&amp;lt;/math&amp;gt; are fixed, there is a unique solution of &amp;lt;math&amp;gt;r_j&amp;lt;/math&amp;gt;. Therefore, the number of &amp;lt;math&amp;gt;r\in\{0,1\}^n&amp;lt;/math&amp;gt; satisfying &amp;lt;math&amp;gt;Dr=\boldsymbol{0}&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;2^{n-1}&amp;lt;/math&amp;gt;. The probability that &amp;lt;math&amp;gt;ABr=Cr&amp;lt;/math&amp;gt; is bounded as&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[ABr=Cr]=\Pr[Dr=\boldsymbol{0}]\le\frac{2^{n-1}}{2^n}=\frac{1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When &amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt;, Freivalds algorithm always returns &amp;quot;yes&amp;quot;; and when &amp;lt;math&amp;gt;AB\neq C&amp;lt;/math&amp;gt;, Freivalds algorithm returns &amp;quot;no&amp;quot; with probability at least 1/2.&lt;br /&gt;
&lt;br /&gt;
To improve its accuracy, we can run Freivalds algorithm for &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; times, each time with an &#039;&#039;independent&#039;&#039; &amp;lt;math&amp;gt;r\in\{0,1\}^n&amp;lt;/math&amp;gt;, and return &amp;quot;yes&amp;quot; if and only if all running instances returns &amp;quot;yes&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Freivalds&#039; Algorithm (multi-round)|&lt;br /&gt;
*pick &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vectors &amp;lt;math&amp;gt;r_1,r_2,\ldots,r_k \in\{0, 1\}^n&amp;lt;/math&amp;gt; uniformly and independently at random;&lt;br /&gt;
*if &amp;lt;math&amp;gt;A(Br_i) = Cr_i&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,\ldots,k&amp;lt;/math&amp;gt; then return &amp;quot;yes&amp;quot; else return &amp;quot;no&amp;quot;;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;AB=C&amp;lt;/math&amp;gt;, then the algorithm returns a &amp;quot;yes&amp;quot; with probability 1. If &amp;lt;math&amp;gt;AB\neq C&amp;lt;/math&amp;gt;, then due to the independence, the probability that all &amp;lt;math&amp;gt;r_i&amp;lt;/math&amp;gt; have &amp;lt;math&amp;gt;ABr_i=C_i&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;2^{-k}&amp;lt;/math&amp;gt;, so the algorithm returns &amp;quot;no&amp;quot; with probability at least &amp;lt;math&amp;gt;1-2^{-k}&amp;lt;/math&amp;gt;. For any &amp;lt;math&amp;gt;0&amp;lt;\epsilon&amp;lt;1&amp;lt;/math&amp;gt;, choose &amp;lt;math&amp;gt;k=\log_2 \frac{1}{\epsilon}&amp;lt;/math&amp;gt;. The algorithm runs in time &amp;lt;math&amp;gt;O(n^2\log_2\frac{1}{\epsilon})&amp;lt;/math&amp;gt; and has a one-sided error (false positive) bounded by &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Polynomial Identity Testing (PIT) =&lt;br /&gt;
The  &#039;&#039;&#039;Polynomial Identity Testing (PIT)&#039;&#039;&#039; is such a problem: given as input two polynomials, determine whether they are identical. It plays a fundamental role in &#039;&#039;Identity Testing&#039;&#039; problems.&lt;br /&gt;
&lt;br /&gt;
First, let&#039;s consider the univariate (&amp;quot;one variable&amp;quot;) case:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; two polynomials &amp;lt;math&amp;gt;f, g\in\mathbb{F}[x]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv g&amp;lt;/math&amp;gt; (&amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; are identical).&lt;br /&gt;
Here the &amp;lt;math&amp;gt;\mathbb{F}[x]&amp;lt;/math&amp;gt; denotes the [http://en.wikipedia.org/wiki/Polynomial_ring ring of univariate polynomials] on a field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. More precisely, a polynomial &amp;lt;math&amp;gt;f\in\mathbb{F}[x]&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x)=\sum_{i=0}^\infty a_ix^i&amp;lt;/math&amp;gt;,&lt;br /&gt;
where the coefficients &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; are taken from the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;, and the addition and multiplication are also defined over the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. &lt;br /&gt;
And:&lt;br /&gt;
* the &#039;&#039;&#039;degree&#039;&#039;&#039; of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is the highest &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; with non-zero &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;;&lt;br /&gt;
* a polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is a &#039;&#039;&#039;zero-polynomial&#039;&#039;&#039;, denoted as &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;, if all coefficients &amp;lt;math&amp;gt;a_i=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Alternatively, we can consider the following equivalent problem by comparing the polynomial &amp;lt;math&amp;gt;f-g&amp;lt;/math&amp;gt; (whose degree is at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;) with the zero-polynomial:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; a polynomial &amp;lt;math&amp;gt;f\in\mathbb{F}[x]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt; (&amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is the 0 polynomial).&lt;br /&gt;
&lt;br /&gt;
The problem is trivial if the input polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is given explicitly: one can trivially solve the problem by checking whether all &amp;lt;math&amp;gt;d+1&amp;lt;/math&amp;gt; coefficients are &amp;lt;math&amp;gt;0&amp;lt;/math&amp;gt;. To make the problem nontrivial, we assume that the input polynomial is given implicitly as a &#039;&#039;black box&#039;&#039; (also called an &#039;&#039;oracle&#039;&#039;): the only way the algorithm can access to &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is to evaluate &amp;lt;math&amp;gt;f(x)&amp;lt;/math&amp;gt; over some &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; from the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; chosen by the algorithm.&lt;br /&gt;
&lt;br /&gt;
A straightforward deterministic algorithm is to evaluate &amp;lt;math&amp;gt;f(x_1),f(x_2),\ldots,f(x_{d+1})&amp;lt;/math&amp;gt; over &amp;lt;math&amp;gt;d+1&amp;lt;/math&amp;gt; &#039;&#039;distinct&#039;&#039; elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_{d+1}&amp;lt;/math&amp;gt; from the field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; and check whether they are all zero. By the [https://en.wikipedia.org/wiki/Fundamental_theorem_of_algebra fundamental theorem of algebra], also known as polynomial interpolations, this guarantees to verify whether a degree-&amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; univariate polynomial &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem|Fundamental Theorem of Algebra|&lt;br /&gt;
:Any non-zero univariate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; roots.&lt;br /&gt;
}}&lt;br /&gt;
The reason for this fundamental theorem holding generally over any field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt; is that any univariate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; factors uniquely into at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; irreducible polynomials, each of which has at most one root.&lt;br /&gt;
&lt;br /&gt;
The following simple randomized algorithm is natural:&lt;br /&gt;
{{Theorem|Algorithm for PIT|&lt;br /&gt;
*suppose we have a finite subset &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; (to be specified later);&lt;br /&gt;
*pick &amp;lt;math&amp;gt;r\in S&amp;lt;/math&amp;gt; &#039;&#039;uniformly&#039;&#039; at random;&lt;br /&gt;
*if &amp;lt;math&amp;gt;f(r) = 0&amp;lt;/math&amp;gt; then return “yes” else return “no”;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This algorithm evaluates &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; at one point chosen uniformly at random from a finite subset &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt;. It is easy to see the followings:&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;, the algorithm always returns &amp;quot;yes&amp;quot;, so it is always correct.&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;, the algorithm may wrongly return &amp;quot;yes&amp;quot; (a &#039;&#039;&#039;false positive&#039;&#039;&#039;). But this happens only when the random &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is a root of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;. By the fundamental theorem of algebra, &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; roots, so the probability that the algorithm is wrong is bounded as&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[f(r)=0]\le\frac{d}{|S|}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By fixing &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; to be an arbitrary subset of size &amp;lt;math&amp;gt;|S|=2d&amp;lt;/math&amp;gt;, this probability of false positive is at most &amp;lt;math&amp;gt;1/2&amp;lt;/math&amp;gt;. We can reduce it to an arbitrarily small constant &amp;lt;math&amp;gt;\delta&amp;lt;/math&amp;gt; by repeat the above testing &#039;&#039;independently&#039;&#039; for &amp;lt;math&amp;gt;\log_2 \frac{1}{\delta}&amp;lt;/math&amp;gt; times, since the error probability decays geometrically as we repeat the algorithm independently. &lt;br /&gt;
&lt;br /&gt;
== Communication Complexity of Equality ==&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Communication_complexity communication complexity] is introduced by Andrew Chi-Chih Yao as a model of computation with more than one entities, each with partial information about the input.&lt;br /&gt;
&lt;br /&gt;
Assume that there are two entities, say Alice and Bob. Alice has a private input &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; and Bob has a private input &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;. Together they want to compute a function &amp;lt;math&amp;gt;f(a,b)&amp;lt;/math&amp;gt; by communicating with each other. The communication follows a predefined &#039;&#039;&#039;communication protocol&#039;&#039;&#039; (the &amp;quot;algorithm&amp;quot; in this model). The complexity of a communication protocol is measured by the number of bits communicated between Alice and Bob in the worst case.&lt;br /&gt;
&lt;br /&gt;
The problem of checking identity is formally defined by the function EQ as follows: &amp;lt;math&amp;gt;\mathrm{EQ}:\{0,1\}^n\times\{0,1\}^n\rightarrow\{0,1\}&amp;lt;/math&amp;gt; and for any &amp;lt;math&amp;gt;a,b\in\{0,1\}^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathrm{EQ}(a,b)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{if } a=b,\\&lt;br /&gt;
0&amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A trivial way to solve EQ is to let Bob send his entire input string &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; to Alice and let Alice check whether &amp;lt;math&amp;gt;a=b&amp;lt;/math&amp;gt;. This costs &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bits of communications.&lt;br /&gt;
&lt;br /&gt;
It is known that for deterministic communication protocols, this is the best we can get for computing EQ.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Yao 1979)|&lt;br /&gt;
:Any deterministic communication protocol computing EQ on two &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-bit strings costs &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bits of communication in the worst-case.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This theorem is much more nontrivial to prove than it looks, because Alice and Bob are allowed to interact with each other in arbitrary ways. The proof of this theorem is in Yao&#039;s [http://math.ucla.edu/~znorwood/290d.2.14s/papers/yao.pdf  celebrated paper in 1979 with a humble title]. It pioneered the field of communication complexity.&lt;br /&gt;
&lt;br /&gt;
If we allow randomness in protocols, and also tolerate a small probabilistic error, the problem can be solved with significantly less communications. To present this randomized protocol, we need a few preparations:&lt;br /&gt;
* We represent the inputs  &amp;lt;math&amp;gt;a,b \in\{0,1\}^{n}&amp;lt;/math&amp;gt; of Alice and Bob as two univariate polynomials of degree at most &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, respectively &lt;br /&gt;
::&amp;lt;math&amp;gt;f(x)=\sum_{i=0}^{n-1}a_ix^{i}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(x)=\sum_{i=0}^{n-1}b_ix^{i}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* The two polynomials &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g&amp;lt;/math&amp;gt; are defined over finite field &amp;lt;math&amp;gt;\mathbb{Z}_p=\{0,1,\ldots,p-1\}&amp;lt;/math&amp;gt; for some suitable prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; (to be specified later), which means the additions and multiplications are modulo &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;.&lt;br /&gt;
The randomized communication protocol is then as follows:&lt;br /&gt;
{{Theorem|A randomized protocol for EQ|&lt;br /&gt;
&#039;&#039;&#039;Bob does&#039;&#039;&#039;:&lt;br /&gt;
:* pick &amp;lt;math&amp;gt;r\in\mathbb{Z}_p&amp;lt;/math&amp;gt; uniformly at random;&lt;br /&gt;
:* send &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(r)&amp;lt;/math&amp;gt; to Alice;&lt;br /&gt;
&#039;&#039;&#039;Upon receiving&#039;&#039;&#039; &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(r)&amp;lt;/math&amp;gt; &#039;&#039;&#039;Alice does&#039;&#039;&#039;:&lt;br /&gt;
:* compute &amp;lt;math&amp;gt;f(r)&amp;lt;/math&amp;gt;;&lt;br /&gt;
:* If &amp;lt;math&amp;gt;f(r)= g(r)&amp;lt;/math&amp;gt; return &amp;quot;&#039;&#039;&#039;yes&#039;&#039;&#039;&amp;quot;; else return &amp;quot;&#039;&#039;&#039;no&#039;&#039;&#039;&amp;quot;.&lt;br /&gt;
}}&lt;br /&gt;
The communication complexity of the protocol is given by the number of bits used to represent the values of &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(r)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since the polynomials are defined over finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; and the random number &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is also chosen from &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, this is bounded by &amp;lt;math&amp;gt;O(\log p)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
On the other hand the protocol makes mistakes only when &amp;lt;math&amp;gt;a\neq b&amp;lt;/math&amp;gt; but wrongly answers &amp;quot;yes&amp;quot;. This happens only when &amp;lt;math&amp;gt;f\not\equiv g&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;f(r)=g(r)&amp;lt;/math&amp;gt;. The degrees of &amp;lt;math&amp;gt;f, g&amp;lt;/math&amp;gt; are at most &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen among &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; distinct values, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[f(r)=g(r)]\le \frac{n-1}{p}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By choosing &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; to be a prime in the interval &amp;lt;math&amp;gt;[n^2, 2n^2]&amp;lt;/math&amp;gt; (by [https://en.wikipedia.org/wiki/Bertrand%27s_postulate Chebyshev&#039;s theorem], such prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; always exists), the above randomized communication protocol solves the Equality function EQ with an error probability of false positive at most &amp;lt;math&amp;gt;O(1/n)&amp;lt;/math&amp;gt;, with communication complexity &amp;lt;math&amp;gt;O(\log n)&amp;lt;/math&amp;gt;, an EXPONENTIAL improvement to ANY deterministic communication protocol!&lt;br /&gt;
&lt;br /&gt;
== Schwartz-Zippel Theorem ==&lt;br /&gt;
Now let&#039;s see the the true form of &#039;&#039;&#039;Polynomial Identity Testing (PIT)&#039;&#039;&#039;, for multivariate polynomials:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; two &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomials &amp;lt;math&amp;gt;f, g\in\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv g&amp;lt;/math&amp;gt;.&lt;br /&gt;
The &amp;lt;math&amp;gt;\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; is the [http://en.wikipedia.org/wiki/Polynomial_ring#The_polynomial_ring_in_several_variables ring of multivariate polynomials] over field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. An &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, written as a sum of monomials, is:&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)=\sum_{i_1,i_2,\ldots,i_n\ge 0\atop i_1+i_2+\cdots+i_n\le d}a_{i_1,i_2,\ldots,i_n}x_{1}^{i_1}x_2^{i_2}\cdots x_{n}^{i_n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The &#039;&#039;&#039;degree&#039;&#039;&#039; or &#039;&#039;&#039;total degree&#039;&#039;&#039; of a monomial &amp;lt;math&amp;gt;a_{i_1,i_2,\ldots,i_n}x_{1}^{i_1}x_2^{i_2}\cdots x_{n}^{i_n}&amp;lt;/math&amp;gt; is given by &amp;lt;math&amp;gt;i_1+i_2+\cdots+i_n&amp;lt;/math&amp;gt; and the degree of a polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is the maximum degree of monomials of nonzero coefficients.&lt;br /&gt;
&lt;br /&gt;
As before, we also consider the following equivalent problem:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; a polynomial &amp;lt;math&amp;gt;f\in\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is written explicitly as a sum of monomials, then the problem can be solved by checking whether all coefficients, and there at most &amp;lt;math&amp;gt;{n+d\choose d}\le (n+d)^{d}&amp;lt;/math&amp;gt; coefficients in an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
A multivariate polynomial &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; can also be presented in its &#039;&#039;&#039;product form&#039;&#039;&#039;, for example:&lt;br /&gt;
{{Theorem|Example|&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Vandermonde_matrix Vandermonde matrix] &amp;lt;math&amp;gt;M=M(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;M_{ij}=x_i^{j-1}&amp;lt;/math&amp;gt;, that is&lt;br /&gt;
:&amp;lt;math&amp;gt;M=\begin{bmatrix}&lt;br /&gt;
1 &amp;amp; x_1 &amp;amp; x_1^2 &amp;amp; \dots &amp;amp; x_1^{n-1}\\&lt;br /&gt;
1 &amp;amp; x_2 &amp;amp; x_2^2 &amp;amp; \dots &amp;amp; x_2^{n-1}\\&lt;br /&gt;
1 &amp;amp; x_3 &amp;amp; x_3^2 &amp;amp; \dots &amp;amp; x_3^{n-1}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp;\vdots \\&lt;br /&gt;
1 &amp;amp; x_n &amp;amp; x_n^2 &amp;amp; \dots &amp;amp; x_n^{n-1}&lt;br /&gt;
\end{bmatrix}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be the polynomial defined as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
f(x_1,\ldots,x_n)=\det(M)=\prod_{j&amp;lt;i}(x_i-x_j).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
For polynomials in product form, it is quite efficient to &#039;&#039;&#039;evaluate&#039;&#039;&#039; the polynomial at any specific point from the field over which the polynomial is defined, however, &#039;&#039;&#039;expanding&#039;&#039;&#039; the polynomial to a sum of monomials can be very expensive.  &lt;br /&gt;
&lt;br /&gt;
The following is a simple randomized algorithm for testing identity of multivariate polynomials:&lt;br /&gt;
{{Theorem|Randomized algorithm for multivariate PIT|&lt;br /&gt;
*suppose we have a finite subset &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; (to be specified later);&lt;br /&gt;
* pick &amp;lt;math&amp;gt;r_1,r_2,\ldots,r_n\in S&amp;lt;/math&amp;gt; &#039;&#039;uniformly&#039;&#039; and &#039;&#039;independently&#039;&#039; at random;&lt;br /&gt;
* if &amp;lt;math&amp;gt;f(\vec{r})=f(r_1,r_2,\ldots,r_n) = 0&amp;lt;/math&amp;gt; then return “yes” else return “no”;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This algorithm evaluates &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; at one point chosen uniformly from an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-dimensional cube &amp;lt;math&amp;gt;S^n&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt; is a finite subset. And:&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\equiv 0&amp;lt;/math&amp;gt;, the algorithm always returns &amp;quot;yes&amp;quot;, so it is always correct.&lt;br /&gt;
* If &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;, the algorithm may wrongly return &amp;quot;yes&amp;quot; (a &#039;&#039;&#039;false positive&#039;&#039;&#039;). But this happens only when the random &amp;lt;math&amp;gt;\vec{r}=(r_1,r_2,\ldots,r_n)&amp;lt;/math&amp;gt; is a root of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;. The probability of this bad event is upper bounded by the following famous result due to [https://pdfs.semanticscholar.org/b913/cf330852035f49b4ec5fe2db86c47d8a98fd.pdf Schwartz (1980)] and [http://www.cecm.sfu.ca/~monaganm/teaching/TopicsinCA15/zippel79.pdf Zippel (1979)]. &lt;br /&gt;
&lt;br /&gt;
{{Theorem|Schwartz-Zippel Theorem|&lt;br /&gt;
: Let &amp;lt;math&amp;gt;f\in\mathbb{F}[x_1,x_2,\ldots,x_n]&amp;lt;/math&amp;gt; be a multivariate polynomial of degree &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; over a field &amp;lt;math&amp;gt;\mathbb{F}&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;, then for any finite set &amp;lt;math&amp;gt;S\subset\mathbb{F}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;r_1,r_2\ldots,r_n\in S&amp;lt;/math&amp;gt; chosen uniformly and independently at random, &lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[f(r_1,r_2,\ldots,r_n)=0]\le\frac{d}{|S|}.&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
The Schwartz-Zippel Theorem states that for any nonzero &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, the number of roots in any cube &amp;lt;math&amp;gt;S^n&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;d\cdot |S|^{n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Dana Moshkovitz gave a surprisingly simply and elegant [http://eccc.hpi-web.de/report/2010/096/ proof] of Schwartz-Zippel Theorem, using some advanced ideas. Now we introduce the standard proof by induction.&lt;br /&gt;
&lt;br /&gt;
{{Proof| &lt;br /&gt;
The theorem is proved by induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
;Induction basis&#039;&#039;&#039;:&#039;&#039;&#039;&lt;br /&gt;
For &amp;lt;math&amp;gt;n=1&amp;lt;/math&amp;gt;, this is the univariate case. Assume that &amp;lt;math&amp;gt;f\not\equiv 0&amp;lt;/math&amp;gt;. Due to the fundamental theorem of algebra, any polynomial &amp;lt;math&amp;gt;f(x)&amp;lt;/math&amp;gt; of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; must have at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; roots, thus &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[f(r)=0]\le\frac{d}{|S|}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
;Induction hypothesis&#039;&#039;&#039;:&#039;&#039;&#039; &lt;br /&gt;
Assume the theorem holds for any &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-variate polynomials for all &amp;lt;math&amp;gt;m&amp;lt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
;Induction step&#039;&#039;&#039;:&#039;&#039;&#039;&lt;br /&gt;
For any &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-variate polynomial &amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt;  of degree at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, we write &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; as&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)=\sum_{i=0}^kx_n^{i}f_i(x_1,x_2,\ldots,x_{n-1})&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; is the highest degree of &amp;lt;math&amp;gt;x_n&amp;lt;/math&amp;gt;, which means the degree of &amp;lt;math&amp;gt;f_k&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;d-k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_k\not\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In particular, we write &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; as a sum of two parts:&lt;br /&gt;
:&amp;lt;math&amp;gt;f(x_1,x_2,\ldots,x_n)=x_n^k f_k(x_1,x_2,\ldots,x_{n-1})+\bar{f}(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt;,&lt;br /&gt;
where both &amp;lt;math&amp;gt;f_k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\bar{f}&amp;lt;/math&amp;gt; are polynomials, such that &lt;br /&gt;
* &amp;lt;math&amp;gt;f_k\not\equiv 0&amp;lt;/math&amp;gt; is as above, whose degree is  at most &amp;lt;math&amp;gt;d-k&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\bar{f}(x_1,x_2,\ldots,x_n)=\sum_{i=0}^{k-1}x_n^i f_i(x_1,x_2,\ldots,x_{n-1})&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;\bar{f}(x_1,x_2,\ldots,x_n)&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;x_n^{k}&amp;lt;/math&amp;gt; factor in any term.&lt;br /&gt;
&lt;br /&gt;
By the law of total probability, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\Pr[f(r_1,r_2,\ldots,r_n)=0]\\&lt;br /&gt;
=&lt;br /&gt;
&amp;amp;\Pr[f(\vec{r})=0\mid f_k(r_1,r_2,\ldots,r_{n-1})=0]\cdot\Pr[f_k(r_1,r_2,\ldots,r_{n-1})=0]\\&lt;br /&gt;
&amp;amp;+\Pr[f(\vec{r})=0\mid f_k(r_1,r_2,\ldots,r_{n-1})\neq0]\cdot\Pr[f_k(r_1,r_2,\ldots,r_{n-1})\neq0].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that &amp;lt;math&amp;gt;f_k(r_1,r_2,\ldots,r_{n-1})&amp;lt;/math&amp;gt; is a polynomial on &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; variables of degree &amp;lt;math&amp;gt;d-k&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;f_k\not\equiv 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the induction hypothesis, we have &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(*)&lt;br /&gt;
&amp;amp;\qquad&lt;br /&gt;
&amp;amp;\Pr[f_k(r_1,r_2,\ldots,r_{n-1})=0]\le\frac{d-k}{|S|}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Now we look at the case conditioning on &amp;lt;math&amp;gt;f_k(r_1,r_2,\ldots,r_{n-1})\neq0&amp;lt;/math&amp;gt;. Recall that &amp;lt;math&amp;gt;\bar{f}(x_1,\ldots,x_n)&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;x_n^k&amp;lt;/math&amp;gt; factor in any term, thus the condition &amp;lt;math&amp;gt;f_k(r_1,r_2,\ldots,r_{n-1})\neq0&amp;lt;/math&amp;gt; guarantees that &lt;br /&gt;
:&amp;lt;math&amp;gt;f(r_1,\ldots,r_{n-1},x_n)=x_n^k f_k(r_1,r_2,\ldots,r_{n-1})+\bar{f}(r_1,r_2,\ldots,r_{n-1},x_n)=g_{r_1,\ldots,r_{n-1}}(x_n)&amp;lt;/math&amp;gt;&lt;br /&gt;
is a nonzero univariate polynomial of &amp;lt;math&amp;gt;x_n&amp;lt;/math&amp;gt; such that the degree of &amp;lt;math&amp;gt;g_{r_1,\ldots,r_{n-1}}(x_n)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g_{r_1,\ldots,r_{n-1}}\not\equiv 0&amp;lt;/math&amp;gt;, for which we already known that the probability &amp;lt;math&amp;gt;g_{r_1,\ldots,r_{n-1}}(r_n)=0&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;\frac{k}{|S|}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(**)&lt;br /&gt;
&amp;amp;\qquad&lt;br /&gt;
&amp;amp;\Pr[f(\vec{r})=0\mid f_k(r_1,r_2,\ldots,r_{n-1})\neq0]=\Pr[g_{r_1,\ldots,r_{n-1}}(r_n)=0\mid f_k(r_1,r_2,\ldots,r_{n-1})\neq0]\le\frac{k}{|S|}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
Substituting both &amp;lt;math&amp;gt;(*)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(**)&amp;lt;/math&amp;gt; back in the total probability, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[f(r_1,r_2,\ldots,r_n)=0]&lt;br /&gt;
\le\frac{d-k}{|S|}+\frac{k}{|S|}=\frac{d}{|S|},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which proves the theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Detecting perfect matching ==&lt;br /&gt;
&lt;br /&gt;
TBA&lt;br /&gt;
&lt;br /&gt;
;Edmonds matrix&lt;br /&gt;
&lt;br /&gt;
= Fingerprinting =&lt;br /&gt;
The polynomial identity testing algorithm in the Schwartz-Zippel theorem can be abstracted as the following framework:&lt;br /&gt;
Suppose we want to compare two objects &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt;. Instead of comparing them directly, we compute random &#039;&#039;&#039;fingerprints&#039;&#039;&#039; &amp;lt;math&amp;gt;\mathrm{FING}(X)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathrm{FING}(Y)&amp;lt;/math&amp;gt; of them and compare the fingerprints. &lt;br /&gt;
&lt;br /&gt;
The fingerprints has the following properties:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; is a function, meaning that if &amp;lt;math&amp;gt;X= Y&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;\mathrm{FING}(X)=\mathrm{FING}(Y)&amp;lt;/math&amp;gt;.&lt;br /&gt;
* It is much easier to compute and compare the fingerprints.&lt;br /&gt;
* Ideally, the domain of fingerprints is much smaller than the domain of original objects, so storing and comparing fingerprints are easy. This means the fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; cannot be an injection (one-to-one mapping), so it&#039;s possible that different &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; are mapped to the same fingerprint. We resolve this by making fingerprint function &#039;&#039;randomized&#039;&#039;, and for &amp;lt;math&amp;gt;X\neq Y&amp;lt;/math&amp;gt;, we want the probability &amp;lt;math&amp;gt;\Pr[\mathrm{FING}(X)=\mathrm{FING}(Y)]&amp;lt;/math&amp;gt; to be small.&lt;br /&gt;
&lt;br /&gt;
In Schwartz-Zippel theorem, the objects to compare are polynomials from &amp;lt;math&amp;gt;\mathbb{F}[x_1,\ldots,x_n]&amp;lt;/math&amp;gt;. Given a polynomial &amp;lt;math&amp;gt;f\in \mathbb{F}[x_1,\ldots,x_n]&amp;lt;/math&amp;gt;, its fingerprint is computed as &amp;lt;math&amp;gt;\mathrm{FING}(f)=f(r_1,\ldots,r_n)&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;r_i&amp;lt;/math&amp;gt; chosen independently and uniformly at random from some fixed set &amp;lt;math&amp;gt;S\subseteq\mathbb{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
With this generic framework, for various identity testing problems, we may design different fingerprints &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Communication protocols for Equality ==&lt;br /&gt;
Now consider again the communication model where the two players Alice with a private input &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt; and Bob with a private input &amp;lt;math&amp;gt;y\in\{0,1\}^n&amp;lt;/math&amp;gt; together compute a function &amp;lt;math&amp;gt;f(x,y)&amp;lt;/math&amp;gt; by running a communication protocol. &lt;br /&gt;
&lt;br /&gt;
We still consider the communication protocols for the equality function EQ&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathrm{EQ}(x,y)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1&amp;amp; \mbox{if } x=y,\\&lt;br /&gt;
0&amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
With the language of fingerprinting, this communication problem can be solved by the following generic scheme:&lt;br /&gt;
{{Theorem|Communication protocol for EQ by fingerprinting|&lt;br /&gt;
&#039;&#039;&#039;Bob does&#039;&#039;&#039;:&lt;br /&gt;
:* choose a random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; and compute the fingerprint of her input &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt;;&lt;br /&gt;
:* sends both the description of &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; and the value of &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt; to Alice;&lt;br /&gt;
&#039;&#039;&#039;Upon receiving&#039;&#039;&#039; the description of &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; and the value of &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt;, &#039;&#039;&#039;Alice does&#039;&#039;&#039;:&lt;br /&gt;
:* computes &amp;lt;math&amp;gt;\mathrm{FING}(x)&amp;lt;/math&amp;gt; and check whether &amp;lt;math&amp;gt;\mathrm{FING}(x)=\mathrm{FING}(y)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
In this way we have a randomized communication protocol for the equality function EQ with false positive. The communication cost as well as the error probability are reduced to the question of how to design this random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; to guarantee:&lt;br /&gt;
# A random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; can be described succinctly.&lt;br /&gt;
# The range of &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; is small, so the fingerprints are succinct.&lt;br /&gt;
# If &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;, the probability &amp;lt;math&amp;gt;\Pr[\mathrm{FING}(x)=\mathrm{FING}(y)]&amp;lt;/math&amp;gt; is small.&lt;br /&gt;
&lt;br /&gt;
=== Fingerprinting by PIT===&lt;br /&gt;
As before, we can define the fingerprint function as: for any bit-string &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt;,  its random fingerprint is &amp;lt;math&amp;gt;\mathrm{FING}(x)=\sum_{i=1}^n x_i r^{i}&amp;lt;/math&amp;gt;, where the additions and multiplications are defined over a finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen uniformly at random from &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; is some suitable prime which can be represented in &amp;lt;math&amp;gt;\Theta(\log n)&amp;lt;/math&amp;gt; bits. More specifically, we can choose &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; to be any prime from the interval &amp;lt;math&amp;gt;[n^2, 2n^2]&amp;lt;/math&amp;gt;. Due to Chebyshev&#039;s theorem, such prime must exist. &lt;br /&gt;
&lt;br /&gt;
As we have shown before, it takes &amp;lt;math&amp;gt;O(\log p)=O(\log n)&amp;lt;/math&amp;gt; bits to represent &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt; and to describe the random function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; (since it a random function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; from this family is uniquely identified by a random &amp;lt;math&amp;gt;r\in\mathbb{Z}_p&amp;lt;/math&amp;gt;, which can be represented within &amp;lt;math&amp;gt;\log p=O(\log n)&amp;lt;/math&amp;gt; bits). And it follows easily from the fundamental theorem of algebra that for any distinct &amp;lt;math&amp;gt;x, y\in\{0,1\}^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\mathrm{FING}(x)=\mathrm{FING}(y)] \le \frac{n-1}{p}\le \frac{1}{n}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Fingerprinting by randomized checksum===&lt;br /&gt;
Now we consider a new fingerprint function: We treat each input string &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt; as the binary representation of a number, and let &amp;lt;math&amp;gt;\mathrm{FING}(x)=x\bmod p&amp;lt;/math&amp;gt; for some random prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; chosen from &amp;lt;math&amp;gt;[k]=\{0,1,\ldots,k-1\}&amp;lt;/math&amp;gt;, for some &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; to be specified later. &lt;br /&gt;
&lt;br /&gt;
Now a random fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; can be uniquely identified by this random prime &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. The new communication protocol for EQ with this fingerprint is as follows:&lt;br /&gt;
{{Theorem|Communication protocol for EQ by random checksum|&lt;br /&gt;
&#039;&#039;&#039;Bob does&#039;&#039;&#039;:&lt;br /&gt;
:for some parameter &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; (to be specified), &lt;br /&gt;
:* choose a prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt; uniformly at random;&lt;br /&gt;
:* send &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x\bmod p&amp;lt;/math&amp;gt; to Alice;&lt;br /&gt;
&#039;&#039;&#039;Upon receiving&#039;&#039;&#039; &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x\bmod p&amp;lt;/math&amp;gt;, &#039;&#039;&#039;Alice does&#039;&#039;&#039;:&lt;br /&gt;
:* check whether &amp;lt;math&amp;gt;x\bmod p=y\bmod p&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The number of bits to be communicated is obviously &amp;lt;math&amp;gt;O(\log k)&amp;lt;/math&amp;gt;. When &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;, we want to upper bound the error probability &amp;lt;math&amp;gt;\Pr[x\bmod p=y\bmod p]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose without loss of generality &amp;lt;math&amp;gt;x&amp;gt;y&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;z=x-y&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;z&amp;lt;2^n&amp;lt;/math&amp;gt; since &amp;lt;math&amp;gt;x,y\in[2^n]&amp;lt;/math&amp;gt;, and &lt;br /&gt;
&amp;lt;math&amp;gt;z\neq 0&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;.  It holds that &amp;lt;math&amp;gt;x\equiv y\pmod p&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;p\mid z&amp;lt;/math&amp;gt;. Therefore, we only need to upper bound the probability&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[z\bmod p=0]&amp;lt;/math&amp;gt;&lt;br /&gt;
for an arbitrarily fixed &amp;lt;math&amp;gt;0&amp;lt;z&amp;lt;2^n&amp;lt;/math&amp;gt;, and a uniform random prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The probability &amp;lt;math&amp;gt;\Pr[z\bmod p=0]&amp;lt;/math&amp;gt; is computed directly as&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[z\bmod p=0]\le\frac{\mbox{the number of prime divisors of }z}{\mbox{the number of primes in }[k]}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For the numerator, any positive &amp;lt;math&amp;gt;z&amp;lt;2^n&amp;lt;/math&amp;gt; has at most &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; prime factors. To see this, by contradiction assume that &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; has more than &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; prime factors. Note that any prime number is at least 2. Then &amp;lt;math&amp;gt;z&amp;lt;/math&amp;gt; must be greater than &amp;lt;math&amp;gt;2^n&amp;lt;/math&amp;gt;, contradicting the fact that &amp;lt;math&amp;gt;z&amp;lt;2^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For the denominator, we need to lower bound the number of primes in &amp;lt;math&amp;gt;[k]&amp;lt;/math&amp;gt;. This is given by the celebrated [http://en.wikipedia.org/wiki/Prime_number_theorem &#039;&#039;&#039;Prime Number Theorem (PNT)&#039;&#039;&#039;].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Prime Number Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\pi(k)&amp;lt;/math&amp;gt; denote the number of primes less than &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\pi(k)\sim\frac{k}{\ln k}&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;k\rightarrow\infty&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Therefore, by choosing &amp;lt;math&amp;gt;k=2n^2\ln n&amp;lt;/math&amp;gt;, we have that for a &amp;lt;math&amp;gt;0&amp;lt;z&amp;lt;2^n&amp;lt;/math&amp;gt;, and a random prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[z\bmod p=0]\le\frac{n}{\pi(k)}\sim\frac{1}{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
which means the for any inputs &amp;lt;math&amp;gt;x,y\in\{0,1\}^n&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;x\neq y&amp;lt;/math&amp;gt;, then the false positive is bounded as&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathrm{FING}(x)=\mathrm{FING}(y)]\le\Pr[|x-y|\bmod p=0]\le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Moreover, by this choice of parameter &amp;lt;math&amp;gt;k=2n^2\ln n&amp;lt;/math&amp;gt;, the communication complexity of the protocol is bounded by &amp;lt;math&amp;gt;O(\log k)=O(\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Pattern matching ==&lt;br /&gt;
Consider the following problem of pattern matching, which has nothing to do with communication complexity.&lt;br /&gt;
&lt;br /&gt;
*Input: a string &amp;lt;math&amp;gt;x\in\{0,1\}^n&amp;lt;/math&amp;gt; and a &amp;quot;pattern&amp;quot; &amp;lt;math&amp;gt;y\in\{0,1\}^m&amp;lt;/math&amp;gt;.&lt;br /&gt;
*Determine whether the pattern &amp;lt;math&amp;gt;y&amp;lt;/math&amp;gt; is a contiguous  substring of &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;. Usually, we are also asked to find the location of the substring.&lt;br /&gt;
&lt;br /&gt;
A naive algorithm trying every possible match runs in &amp;lt;math&amp;gt;O(nm)&amp;lt;/math&amp;gt; time. The more sophisticated KMP algorithm inspired by automaton theory runs in &amp;lt;math&amp;gt;O(n+m)&amp;lt;/math&amp;gt; time.&lt;br /&gt;
&lt;br /&gt;
A simple randomized algorithm, due to Karp and Rabin, uses the idea of fingerprinting and also runs in &amp;lt;math&amp;gt;O(n + m)&amp;lt;/math&amp;gt; time.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X(j)=x_jx_{j+1}\cdots x_{j+m-1}&amp;lt;/math&amp;gt; denote the substring of &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; of length &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; starting at position &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Algorithm (Karp-Rabin)|&lt;br /&gt;
:pick a random prime &amp;lt;math&amp;gt;p\in[k]&amp;lt;/math&amp;gt;;&lt;br /&gt;
:&#039;&#039;&#039;for&#039;&#039;&#039; &amp;lt;math&amp;gt;j = 1&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;n -m + 1&amp;lt;/math&amp;gt; &#039;&#039;&#039;do&#039;&#039;&#039;&lt;br /&gt;
::&#039;&#039;&#039;if&#039;&#039;&#039; &amp;lt;math&amp;gt;X(j)\bmod p = y \bmod p&amp;lt;/math&amp;gt; then report a match;&lt;br /&gt;
:&#039;&#039;&#039;return&#039;&#039;&#039; &amp;quot;no match&amp;quot;;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
So the algorithm just compares the &amp;lt;math&amp;gt;\mathrm{FING}(X(j))&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathrm{FING}(y)&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;, with the same definition of fingerprint function &amp;lt;math&amp;gt;\mathrm{FING}(\cdot)&amp;lt;/math&amp;gt; as in the communication protocol for EQ.&lt;br /&gt;
&lt;br /&gt;
By the same analysis, by choosing &amp;lt;math&amp;gt;k=n^2m\ln (n^2m)&amp;lt;/math&amp;gt;, the probability of a single false match is &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X(j)\bmod p=y\bmod p\mid X(j)\neq y ]=O\left(\frac{1}{n^2}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the union bound, the probability that a false match occurs is &amp;lt;math&amp;gt;O\left(\frac{1}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The algorithm runs in linear time if we assume that we can compute &amp;lt;math&amp;gt;X(j)\bmod p &amp;lt;/math&amp;gt; for each &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; in constant time. This outrageous assumption can be made realistic by the following observation.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathrm{FING}(a)=a\bmod p&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;\mathrm{FING}(X(j+1))\equiv2(\mathrm{FING}(X(j))-2^{m-1}x_j)+x_{j+m}\pmod p\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| It holds that &lt;br /&gt;
:&amp;lt;math&amp;gt;X(j+1)=2(X(j)-2^{m-1}x_j)+x_{j+m}\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
So the equation holds on the finite field modulo &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Due to this lemma, each fingerprint &amp;lt;math&amp;gt;\mathrm{FING}(X(j))&amp;lt;/math&amp;gt; can be computed in an incremental way, each in constant time. The running time of the algorithm is &amp;lt;math&amp;gt;O(n+m)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Checking distinctness ==&lt;br /&gt;
Consider the following problem of &#039;&#039;&#039;checking distinctness&#039;&#039;&#039;:&lt;br /&gt;
*Given a sequence &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_n\in\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, check whether every element of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; appears &#039;&#039;&#039;exactly&#039;&#039;&#039; once.&lt;br /&gt;
&lt;br /&gt;
Obviously this problem can be solved in linear time and linear space (in addition to the space for storing the input) by maintaining a &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-bit vector that indicates which numbers among &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; have appeared.&lt;br /&gt;
&lt;br /&gt;
When this &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is enormously large, &amp;lt;math&amp;gt;\Omega(n)&amp;lt;/math&amp;gt; space cost is too expensive. We wonder whether we could solve this problem with a space cost (in addition to the space for storing the input) much less than &amp;lt;math&amp;gt;O(n)&amp;lt;/math&amp;gt;. This can be done by fingerprinting if we tolerate a certain degree of inaccuracy.&lt;br /&gt;
&lt;br /&gt;
We consider the following more generalized problem, &#039;&#039;&#039;checking identity of multisets&#039;&#039;&#039;:&lt;br /&gt;
* &#039;&#039;&#039;Input:&#039;&#039;&#039; two multisets &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots, a_n\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B=\{b_1,b_2,\ldots, b_n\}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;a_1,a_2,\ldots,b_1,b_2,\ldots,b_n\in \{1,2,\ldots,n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Determine whether &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt; (multiset equivalence).&lt;br /&gt;
&lt;br /&gt;
Here for a &#039;&#039;&#039;multiset&#039;&#039;&#039; &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots, a_n\}&amp;lt;/math&amp;gt;, its elements &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; are not necessarily distinct. The &#039;&#039;&#039;multiplicity&#039;&#039;&#039; of an element &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; in a multiset &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is the number of times &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; appears in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. Two multisets &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are equivalent if they contain the same set of elements and the multiplicities of every element in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are equal. &lt;br /&gt;
&lt;br /&gt;
Obviously the above problem of checking distinctness can be treated as a special case of checking identity of multisets: by checking the identity of the multiset &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and set &amp;lt;math&amp;gt;\{1,2,\ldots, n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The following fingerprinting function for multisets was introduced by Lipton for solving multiset identity testing.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fingerprint for multiset|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; be a uniform random prime chosen from the interval &amp;lt;math&amp;gt;[(n\log n)^2,2(n\log n)^2]&amp;lt;/math&amp;gt;. By Chebyshev&#039;s theorem, such prime must exist. And consider the the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p=[p]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
:Given a multiset &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots,a_n\}&amp;lt;/math&amp;gt;, we define a univariate polynomial &amp;lt;math&amp;gt;f_A\in\mathbb{Z}_p[x]&amp;lt;/math&amp;gt; over the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
::&amp;lt;math&amp;gt;f_A(x)=\prod_{i=1}^n(x-a_i)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; are defined over the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
:We then define the random fingerprinting function as:&lt;br /&gt;
::&amp;lt;math&amp;gt;\mathrm{FING}(A)=f_A(r)=\prod_{i=1}^n(r-a_i)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen uniformly at random from &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Since all computations of &amp;lt;math&amp;gt;\mathrm{FING}(A)=\prod_{i=1}^n(r-a_i)&amp;lt;/math&amp;gt; are over the finite field &amp;lt;math&amp;gt;\mathbb{Z}_p&amp;lt;/math&amp;gt;, the space cost for computing the fingerprint &amp;lt;math&amp;gt;\mathrm{FING}(A)&amp;lt;/math&amp;gt; is only &amp;lt;math&amp;gt;O(\log p)=O(\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Moreover, the fingerprinting function &amp;lt;math&amp;gt;\mathrm{FING}(A)&amp;lt;/math&amp;gt; is invariant under permutation of elements of the multiset &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots,a_n\}&amp;lt;/math&amp;gt;, thus it is indeed a function of multisets (meaning every multiset has only one fingerprint). Therefore, if &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;\mathrm{FING}(A)=\mathrm{FING}(B)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For two distinct multisets &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;, it is possible that &amp;lt;math&amp;gt;\mathrm{FING}(A)=\mathrm{FING}(B)&amp;lt;/math&amp;gt;, but the following theorem due to Lipton bounds this error probability of false positive.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Lipton 1989)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A=\{a_1,a_2,\ldots,a_n\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B=\{b_1,b_2,\ldots,b_n\}&amp;lt;/math&amp;gt; be two multisets whose elements are from &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[\mathrm{FING}(A)= \mathrm{FING}(B)]=O\left(\frac{1}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| &lt;br /&gt;
Let &amp;lt;math&amp;gt;\tilde{f}_A(x)=\prod_{i=1}^n(x-a_i)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\tilde{f}_B(x)=\prod_{i=1}^n(x-b_i)&amp;lt;/math&amp;gt; be two univariate polynomials defined over reals &amp;lt;math&amp;gt;\mathbb{R}&amp;lt;/math&amp;gt;. Note that in contrast to &amp;lt;math&amp;gt;f_A(x)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_B(x)&amp;lt;/math&amp;gt;,  the &amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\cdot&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;\tilde{f}_A(x), \tilde{f}_B(x)&amp;lt;/math&amp;gt; do not modulo &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. It is easy to verify that the polynomials &amp;lt;math&amp;gt;\tilde{f}_A(x), \tilde{f}_B(x)&amp;lt;/math&amp;gt; have the following properties:&lt;br /&gt;
*&amp;lt;math&amp;gt;\tilde{f}_A\equiv \tilde{f}_B&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt;. Here &amp;lt;math&amp;gt;A=B&amp;lt;/math&amp;gt; means the multiset equivalence.&lt;br /&gt;
*By the properties of finite field, for any value &amp;lt;math&amp;gt;r\in\mathbb{Z}_p&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;f_A(r)=\tilde{f}_A(r)\bmod p&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_B(r)=\tilde{f}_B(r)\bmod p&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, assuming that &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;, we must have &amp;lt;math&amp;gt;\tilde{f}_A(x)\not\equiv \tilde{f}_B(x)&amp;lt;/math&amp;gt;. Then by the law of total probability:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathrm{FING}(A)= \mathrm{FING}(B)]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[f_A(r)=f_B(r)\mid f_A\not\equiv f_B\right]\Pr[f_A\not\equiv f_B]\\&lt;br /&gt;
&amp;amp;\quad\,\,+\Pr\left[f_A(r)=f_B(r)\mid f_A\equiv f_B\right]\Pr[f_A\equiv f_B]\\&lt;br /&gt;
&amp;amp;\le \Pr\left[f_A(r)=f_B(r)\mid f_A\not\equiv f_B\right]+\Pr[f_A\equiv f_B].&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that the degrees of &amp;lt;math&amp;gt;f_A,f_B&amp;lt;/math&amp;gt; are at most &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; is chosen uniformly from &amp;lt;math&amp;gt;[p]&amp;lt;/math&amp;gt;. By the Schwartz-Zippel theorem for univariate polynomials, the first probability&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[f_A(r)=f_B(r)\mid f_A\not\equiv f_B\right]\le \frac{n}{p}=o\left(\frac{1}{n}\right),&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
since &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; is chosen from the interval &amp;lt;math&amp;gt;[(n\log n)^2,2(n\log n)^2]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For the second probability &amp;lt;math&amp;gt;\Pr[f_A\equiv f_B]&amp;lt;/math&amp;gt;, recall that &amp;lt;math&amp;gt;\tilde{f}_A\not\equiv \tilde{f}_B&amp;lt;/math&amp;gt;, therefore there is at least a non-zero coefficient &amp;lt;math&amp;gt;c\le n^n&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;\tilde{f}_A-\tilde{f}_B&amp;lt;/math&amp;gt;. The event &amp;lt;math&amp;gt;f_A\equiv f_B&amp;lt;/math&amp;gt; occurs only if &amp;lt;math&amp;gt;c\bmod p=0&amp;lt;/math&amp;gt;, which means&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[f_A\equiv f_B]&lt;br /&gt;
&amp;amp;\le \Pr[c\bmod p=0]\\&lt;br /&gt;
&amp;amp;=\frac{\text{number of prime factors of }c}{\text{number of primes in }[(n\log n)^2,2(n\log n)^2]}\\&lt;br /&gt;
&amp;amp;\le \frac{n\log_2n}{\pi(2(n\log n)^2)-\pi((n\log n)^2)}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By the prime number theorem, &amp;lt;math&amp;gt;\pi(N)\rightarrow \frac{N}{\ln N}&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;N\to\infty&amp;lt;/math&amp;gt;. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[f_A\equiv f_B]=O\left(\frac{n\log n}{n^2\log n}\right)=O\left(\frac{1}{n}\right).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Combining everything together, we have&lt;br /&gt;
&amp;lt;math&amp;gt;\Pr[\mathrm{FING}(A)= \mathrm{FING}(B)]=O\left(\frac{1}{n}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13890</id>
		<title>高级算法 (Fall 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E9%AB%98%E7%BA%A7%E7%AE%97%E6%B3%95_(Fall_2026)&amp;diff=13890"/>
		<updated>2026-09-02T09:18:27Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;高级算法 &lt;br /&gt;
&amp;lt;br&amp;gt;Advanced Algorithms&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;栗师&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = shili@nju.edu.cn &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7= office&lt;br /&gt;
|data7= 计算机系 605&lt;br /&gt;
|header8 = &lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header9 = &lt;br /&gt;
|label9  = Email&lt;br /&gt;
|data9   = liu@nju.edu.cn &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10= office&lt;br /&gt;
|data10= 计算机系 516&lt;br /&gt;
|header11 = Class&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = &lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = Class meetings&lt;br /&gt;
|data12   = 周一 5-6节 (单) 仙Ⅰ-319&lt;br /&gt;
周三 5-6节 仙Ⅰ-319&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = Place&lt;br /&gt;
|data13   = &lt;br /&gt;
|header14 =&lt;br /&gt;
|label14  = Office hours&lt;br /&gt;
|data14   = 周一4-5pm（尹一通 804）&amp;lt;br/&amp;gt;&lt;br /&gt;
周四4-5pm（刘景铖 516）&amp;lt;br/&amp;gt;&lt;br /&gt;
By appointment&lt;br /&gt;
|header15 = Textbooks&lt;br /&gt;
|label15  = &lt;br /&gt;
|data15   = &lt;br /&gt;
|header16 =&lt;br /&gt;
|label16  = &lt;br /&gt;
|data16   = [[File:MR-randomized-algorithms.png|border|100px]]&lt;br /&gt;
|header17 =&lt;br /&gt;
|label17  = &lt;br /&gt;
|data17   = Motwani and Raghavan. &amp;lt;br&amp;gt;&#039;&#039;Randomized Algorithms&#039;&#039;.&amp;lt;br&amp;gt; Cambridge Univ Press, 1995.&lt;br /&gt;
|header18 =&lt;br /&gt;
|label18  = &lt;br /&gt;
|data18   = [[File:Approximation_Algorithms.jpg|border|100px]]&lt;br /&gt;
|header19 =&lt;br /&gt;
|label19  = &lt;br /&gt;
|data19   =  Vazirani. &amp;lt;br&amp;gt;&#039;&#039;Approximation Algorithms&#039;&#039;. &amp;lt;br&amp;gt; Springer-Verlag, 2001.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Advanced Algorithms&#039;&#039; class of fall 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:*[https://tcs.nju.edu.cn/shili/ 栗师]：[mailto:shili@nju.edu.cn &amp;lt;shili@nju.edu.cn&amp;gt;]，计算机系 605&lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching Assistant&#039;&#039;&#039;: &lt;br /&gt;
** 于逸潇：[mailto:yixiaoyu@smail.nju.edu.cn &amp;lt;yixiaoyu@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
** 张弈垚：[mailto:zhangyiyao@smail.nju.edu.cn &amp;lt;zhangyiyao@smail.nju.edu.cn&amp;gt;]&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: &lt;br /&gt;
** 周一 5-6节 1-17周(单) 仙Ⅰ-319&lt;br /&gt;
** 周三 5-6节 1-18周 仙Ⅰ-319&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
** 周一4-5pm（尹一通 804）&lt;br /&gt;
** 周四4-5pm（刘景铖 516）&lt;br /&gt;
** By appointment&lt;br /&gt;
* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1098567018&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
随着计算机算法理论的不断发展，现代计算机算法的设计与分析大量地使用非初等的数学工具以及非传统的算法思想。“高级算法”这门课程就是面向计算机算法的这一发展趋势而设立的。课程将针对传统算法课程未系统涉及、却在计算机科学各领域的科研和实践中扮演重要角色的高等算法设计思想和算法分析工具进行系统讲授。&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 必须：离散数学，概率论，线性代数。&lt;br /&gt;
* 推荐：算法设计与分析。&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[高级算法 (Fall 2026) / Course materials|&amp;lt;font size=3&amp;gt;教材和参考书&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
Late policy: In general, we will accomodate late submission requests ONLY IF you made such requests ahead of time. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[高级算法 (Fall 2025)/Min Cut, Max Cut, and Spectral Cut|Min Cut, Max Cut, and Spectral Cut]] ([[http://tcs.nju.edu.cn/slides/aa2026/Cut.pdf slides]])&lt;br /&gt;
#*  [[高级算法 (Fall 2025)/Probability Basics|Probability basics]]&lt;br /&gt;
#  [[高级算法 (Fall 2026)/Fingerprinting| Fingerprinting]] ([http://tcs.nju.edu.cn/slides/aa2026/Fingerprinting.pdf slides]) &lt;br /&gt;
#*  [[高级算法 (Fall 2026)/Finite Field Basics|Finite field basics]]&lt;br /&gt;
&lt;br /&gt;
= Related Online Courses=&lt;br /&gt;
* [https://www.cs.cmu.edu/~15850/ Advanced Algorithms] by Anupam Gupta at CMU.&lt;br /&gt;
* [http://people.csail.mit.edu/moitra/854.html Advanced Algorithms] by Ankur Moitra at MIT.&lt;br /&gt;
* [http://courses.csail.mit.edu/6.854/current/ Advanced Algorithms] by David Karger and Aleksander Mądry at MIT.&lt;br /&gt;
* [http://web.stanford.edu/class/cs168/index.html The Modern Algorithmic Toolbox] by Tim Roughgarden and Gregory Valiant at Stanford.&lt;br /&gt;
* [https://www.cs.princeton.edu/courses/archive/fall18/cos521/ Advanced Algorithm Design] by Pravesh Kothari and Christopher Musco at Princeton.&lt;br /&gt;
* [http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15859-f11/www/ Linear and Semidefinite Programming (Advanced Algorithms)] by Anupam Gupta and Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://www.cs.cmu.edu/~odonnell/papers/cs-theory-toolkit-lecture-notes.pdf CS Theory Toolkit] by Ryan O&#039;Donnell at CMU.&lt;br /&gt;
* [https://cs.uwaterloo.ca/~lapchi/cs860/index.html Eigenvalues and Polynomials] by Lap Chi Lau at University of Waterloo.&lt;br /&gt;
* The [https://www.cs.cornell.edu/jeh/book.pdf &amp;quot;Foundations of Data Science&amp;quot; book] by Avrim Blum, John Hopcroft, and Ravindran Kannan.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Matching_theory&amp;diff=13833</id>
		<title>组合数学 (Fall 2026)/Matching theory</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Matching_theory&amp;diff=13833"/>
		<updated>2026-06-17T04:05:18Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== Systems of Distinct Representatives (SDR)== A &amp;#039;&amp;#039;&amp;#039;system of distinct representatives (SDR)&amp;#039;&amp;#039;&amp;#039; (also called a &amp;#039;&amp;#039;&amp;#039;transversal&amp;#039;&amp;#039;&amp;#039;) for a sequence of (not necessarily distinct) sets &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt; is a sequence of &amp;lt;font color=red&amp;gt;&amp;#039;&amp;#039;distinct&amp;#039;&amp;#039;&amp;lt;/font&amp;gt; elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_m&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in S_i&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,2,\ldots,m&amp;lt;/math&amp;gt;.  === Hall&amp;#039;s marriage theorem === If the sets &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt; have a system of dist...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Systems of Distinct Representatives (SDR)==&lt;br /&gt;
A &#039;&#039;&#039;system of distinct representatives (SDR)&#039;&#039;&#039; (also called a &#039;&#039;&#039;transversal&#039;&#039;&#039;) for a sequence of (not necessarily distinct) sets &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt; is a sequence of &amp;lt;font color=red&amp;gt;&#039;&#039;distinct&#039;&#039;&amp;lt;/font&amp;gt; elements &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_m&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in S_i&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,2,\ldots,m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Hall&#039;s marriage theorem ===&lt;br /&gt;
If the sets &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt; have a system of distinct representatives &amp;lt;math&amp;gt;x_1\in S_1,x_2\in S_2,\ldots,x_m\in S_m&amp;lt;/math&amp;gt;, then it is obvious that &amp;lt;math&amp;gt;\left|S_1\cup S_2\cup\cdots\cup S_m\right|\ge |\{x_1,x_2,\ldots,x_m\}|=m&amp;lt;/math&amp;gt;. Moreover, for any subset &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,m\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|\ge |\{x_i\mid i\in I\}|=|I|&amp;lt;/math&amp;gt;&lt;br /&gt;
because &amp;lt;math&amp;gt;x_1,x_2,\ldots,x_m&amp;lt;/math&amp;gt; are distinct.&lt;br /&gt;
&lt;br /&gt;
Surprisingly, this obvious necessary condition for the existence of SDR is also sufficient, which is stated by the Hall&#039;s theorem, also called the &#039;&#039;&#039;mariage theorem&#039;&#039;&#039;.&lt;br /&gt;
{{Theorem|Hall&#039;s Theorem|&lt;br /&gt;
:The sets &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt; have a system of distinct representatives (SDR) if and only if&lt;br /&gt;
::&amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|\ge |I|&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,m\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The condition that &amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|\ge |I|&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,m\}&amp;lt;/math&amp;gt; is also called the &#039;&#039;&#039;Hall&#039;s condition&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{{Proof|&lt;br /&gt;
We only need to prove the sufficiency of Hall&#039;s condition for the existence of SDR. We do it by induction on &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;. When &amp;lt;math&amp;gt;m=1&amp;lt;/math&amp;gt;, the theorem trivially holds. Now assume the theorem hold for any integer smaller than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
A subcollection of sets &amp;lt;math&amp;gt;\{S_i\mid i\in I\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;|I|&amp;lt;m\,&amp;lt;/math&amp;gt;, is called a &#039;&#039;&#039;critical family&#039;&#039;&#039; if &amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|=|I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Case.1:&#039;&#039;&#039; There is no critical family, i.e. for each &amp;lt;math&amp;gt;I\subset\{1,2,\ldots,m\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|&amp;gt;|I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;S_m&amp;lt;/math&amp;gt; and choose an arbitrary &amp;lt;math&amp;gt;x\in S_m&amp;lt;/math&amp;gt; as its representative. Remove &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; from all other sets by letting &amp;lt;math&amp;gt;S&#039;_i=S_i\setminus \{x\}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;1\le i\le m-1&amp;lt;/math&amp;gt;. Then for all &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,m-1\}&amp;lt;/math&amp;gt;, &lt;br /&gt;
:&amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i&#039;\right|\ge \left|\bigcup_{i\in I}S_i\right|-1\ge |I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the induction hypothesis, &amp;lt;math&amp;gt;S_1,\ldots,S_m&amp;lt;/math&amp;gt; have an SDR, say &amp;lt;math&amp;gt;x_1\in S_1,\ldots,x_{m-1}\in S_{m-1}&amp;lt;/math&amp;gt;. It is obvious that none of them equals &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; because &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; is removed. Thus, &amp;lt;math&amp;gt;x_1,\ldots,x_{m-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x\,&amp;lt;/math&amp;gt; form an SDR for &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Case.2:&#039;&#039;&#039; There is a critical family, i.e. &amp;lt;math&amp;gt;\exists I\subset\{1,2,\ldots,m\}, |I|&amp;lt;m&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|=|I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose &amp;lt;math&amp;gt;S_{m-k+1},\ldots, S_m&amp;lt;/math&amp;gt; are such a collection of &amp;lt;math&amp;gt;k&amp;lt;m&amp;lt;/math&amp;gt; sets. Hall&#039;s condition certainly holds for these sets. Since &amp;lt;math&amp;gt;k&amp;lt;m&amp;lt;/math&amp;gt;, due to the induction hypothesis, there is an SDR for the sets, say &amp;lt;math&amp;gt;x_{m-k+1}\in S_{m-k+1},\ldots,x_m\in S_m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Again, remove the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; elements from the remaining sets by letting &amp;lt;math&amp;gt;S&#039;_i=S_i\setminus\{x_{m-k+1},\ldots,x_m\}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;1\le i\le m-k&amp;lt;/math&amp;gt;. By the Hall&#039;s condition, &lt;br /&gt;
for any &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,m-k\}&amp;lt;/math&amp;gt;, writing that &amp;lt;math&amp;gt;S=\bigcup_{i\in I}S_i\cup\bigcup_{i=m-k+1}^m S_i&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\left|S\right|\ge |I|+k&amp;lt;/math&amp;gt;,&lt;br /&gt;
thus &lt;br /&gt;
:&amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i&#039;\right|\ge |S|-\left|\bigcup_{i=m-k+1}^m S_i\right|\ge |I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the induction hypothesis, there is an SDR for &amp;lt;math&amp;gt;S_1,\ldots, S_{m-k}&amp;lt;/math&amp;gt;, say &amp;lt;math&amp;gt;x_1\in S_1,\ldots, x_{m-k}\in S_{m-k}&amp;lt;/math&amp;gt;. Combining it with the SDR &amp;lt;math&amp;gt;x_{m-k+1},\ldots, x_{m}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;S_{m-k+1},\ldots, S_{m}&amp;lt;/math&amp;gt;, we have an SDR for &amp;lt;math&amp;gt;S_{1},\ldots, S_{m}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Hall&#039;s theorem is usually stated as a theorem for the existence of matching in a bipartite graph.&lt;br /&gt;
&lt;br /&gt;
In a graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt;, a &#039;&#039;&#039;matching&#039;&#039;&#039; &amp;lt;math&amp;gt;M\subseteq E&amp;lt;/math&amp;gt; is an independent set for edges, that is, for any &amp;lt;math&amp;gt;e_1,e_2\in M&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;e_1\neq e_2&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;e_1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;e_2&amp;lt;/math&amp;gt; are not adjacent to the same vertex.&lt;br /&gt;
&lt;br /&gt;
In a bipartite graph &amp;lt;math&amp;gt;G(U,V,E)&amp;lt;/math&amp;gt;, we say &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; is a matching of &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; (or a matching of &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;), if every vertex in &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; (or &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;) is adjacent to some edge in &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt;, i.e., all vertices in &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; (or &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;) are matched.&lt;br /&gt;
&lt;br /&gt;
In a graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt;, for any vertex &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;N(v)&amp;lt;/math&amp;gt; denote the set of neighbors of &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;; and for any vertex set &amp;lt;math&amp;gt;S\subseteq V&amp;lt;/math&amp;gt;, we override the notation as &amp;lt;math&amp;gt;N(S)=\bigcup_{v\in S}N(v)&amp;lt;/math&amp;gt;, i.e. the set of vertices that are adjacent to one of the vertices in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Hall&#039;s Theorem (graph theory form)|&lt;br /&gt;
:A bipartite graph &amp;lt;math&amp;gt;G(U,V,E)&amp;lt;/math&amp;gt; has a matching of &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; if and only if&lt;br /&gt;
::&amp;lt;math&amp;gt;\left|N(S)\right|\ge |S|&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;S\subseteq U&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Consider the collection of sets &amp;lt;math&amp;gt;N(u), u\in U&amp;lt;/math&amp;gt;. A matching of &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; is an SDR for these sets. Then clearly the theorem is equivalent to Hall&#039;s theorem.&lt;br /&gt;
&lt;br /&gt;
=== Min-max theorems ===&lt;br /&gt;
In combinatorics (and also in other branches of mathematics), there is a family of theorems which relate the minimum of one thing to the maximum of something else. The following are some examples.&lt;br /&gt;
*&#039;&#039;&#039;König-Egerváry theorem&#039;&#039;&#039; (König 1931; Egerváry 1931): in a bipartite graph, the maximum number of edges in a matching equals the minimum number of vertices in a vertex cover.&lt;br /&gt;
*&#039;&#039;&#039;Menger&#039;s theorem&#039;&#039;&#039; (Menger 1927): the minimum number of vertices separating two given vertices in a graph equals the maximum number of vertex-disjoint paths between the two vertices.&lt;br /&gt;
*&#039;&#039;&#039;Dilworth&#039;s theorem&#039;&#039;&#039; (Dilworth 1950): the minimum number of chains which cover a partially ordered set equals the maximum number of elements in an antichain.&lt;br /&gt;
&lt;br /&gt;
== König-Egerváry theorem==&lt;br /&gt;
A [http://en.wikipedia.org/wiki/Matching_(graph_theory) &#039;&#039;&#039;matching&#039;&#039;&#039;] in a graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is a set &amp;lt;math&amp;gt;M\subseteq E&amp;lt;/math&amp;gt; of edges such that no two edges &amp;lt;math&amp;gt;e_1,e_2\in M&amp;lt;/math&amp;gt; shares a vertex. In other words, a matching is just an edge version of independent set.&lt;br /&gt;
&lt;br /&gt;
A [http://en.wikipedia.org/wiki/Vertex_cover &#039;&#039;&#039;vertex cover&#039;&#039;&#039;] in a graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is a vertex set &amp;lt;math&amp;gt;C\subseteq V&amp;lt;/math&amp;gt; such that every edge &amp;lt;math&amp;gt;e\in E&amp;lt;/math&amp;gt; is adjacent to some &amp;lt;math&amp;gt;u\in C&amp;lt;/math&amp;gt;, that is, all edges in the graph are &amp;quot;covered&amp;quot; by some vertex in the vertex cover &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;König-Egerváry theorem&#039;&#039;&#039; (also called the &#039;&#039;&#039;König&#039;s theorem&#039;&#039;&#039;) states the equality of the sizes of maximum matching and minimum vertex cover.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|König-Egerváry Theorem (graph theory form)|&lt;br /&gt;
:In any bipartite graph, the size of a &#039;&#039;maximum&#039;&#039; matching equals the size of a &#039;&#039;minimum&#039;&#039; vertex cover.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The König-Egerváry theorem can be reformulated in its matrix form. A bipartite graph &amp;lt;math&amp;gt;G(U,V,E)&amp;lt;/math&amp;gt; can be represented as a &amp;lt;math&amp;gt;|U|\times |V|&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; with 0-1 entries. For any &amp;lt;math&amp;gt;u\in U&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;A(u,v)=1&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;. (Note that this definition is different from adjacency matrix for graphs.)&lt;br /&gt;
&lt;br /&gt;
Then, a matching in &amp;lt;math&amp;gt;G(U,V,E)&amp;lt;/math&amp;gt; corresponds to a set of &#039;&#039;&#039;independent 1&#039;s&#039;&#039;&#039; in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;: a set of 1&#039;s that do not share rows or columns. A vertex cover corresponds to a set of rows and columns &#039;&#039;&#039;covering&#039;&#039;&#039; all 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;: a set of rows and columns that every 1-entry in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; belongs to at least one of these rows or columns.&lt;br /&gt;
&lt;br /&gt;
It is easy to see the König-Egerváry theorem for bipartite graphs can be equivalently described as follows:&lt;br /&gt;
&lt;br /&gt;
{{Theorem|König-Egerváry Theorem (matrix form)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;m\times n&amp;lt;/math&amp;gt; 0-1 matrix. The &#039;&#039;maximum&#039;&#039; number of independent 1&#039;s is equal to the &#039;&#039;minimum&#039;&#039; number of rows and columns required to cover all the 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We give a proof by the Hall&#039;s theorem.&lt;br /&gt;
&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; denote the maximum number of independent 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; be the minimum number of rows and columns to cover all 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;r\le s&amp;lt;/math&amp;gt;, since any set of &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; independent 1&#039;s requires together &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; rows and columns to cover. &lt;br /&gt;
&lt;br /&gt;
We now prove &amp;lt;math&amp;gt;r\ge s&amp;lt;/math&amp;gt;. Assume that some &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; rows and &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; columns cover all the 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;s=a+b&amp;lt;/math&amp;gt;, i.e. the covering is minimum. Because permuting the rows or the columns change neither &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; nor &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; (as reordering the vertices on either side in a bipartite graph changes nothing to the size of matchings and vertex covers), we may assume that the first &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; rows and the first &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; columns cover the 1&#039;s. Write &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; in the form&lt;br /&gt;
:&amp;lt;math&amp;gt;A=\begin{bmatrix}&lt;br /&gt;
B_{a\times b} &amp;amp;C_{a\times (n-b)}\\&lt;br /&gt;
D_{(m-a)\times b} &amp;amp;E_{(m-a)\times (n-b)}&lt;br /&gt;
\end{bmatrix}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where the submatrix &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; has only zero entries. We will show that there are &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; independent 1&#039;s in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; independent 1&#039;s in &amp;lt;math&amp;gt;D&amp;lt;/math&amp;gt;, thus together &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;a+b=s&amp;lt;/math&amp;gt; independent 1&#039;s, which will imply that &amp;lt;math&amp;gt;r\ge s&amp;lt;/math&amp;gt;, as desired.&lt;br /&gt;
&lt;br /&gt;
Define&lt;br /&gt;
:&amp;lt;math&amp;gt;S_i=\{j\mid c_{ij}=1\}&amp;lt;/math&amp;gt;&lt;br /&gt;
It is obvious that &amp;lt;math&amp;gt;S_i\subseteq\{1,2,\ldots,n-b\}&amp;lt;/math&amp;gt;. We claim that the sequence &amp;lt;math&amp;gt;S_1,S_2,\ldots, S_a&amp;lt;/math&amp;gt; has a system of distinct representatives, i.e., we can choose a 1 from each row, no two in the same column. Otherwise, Hall&#039;s theorem tells us that there exists some &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,a\}&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|&amp;lt;|I|&amp;lt;/math&amp;gt;, that is, the 1&#039;s in the rows contained by &amp;lt;math&amp;gt;I&amp;lt;/math&amp;gt; can be covered by less than &amp;lt;math&amp;gt;|I|&amp;lt;/math&amp;gt; columns. Thus, the 1&#039;s in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; can be covered by &amp;lt;math&amp;gt;a-|I|&amp;lt;/math&amp;gt; and less than &amp;lt;math&amp;gt;|I|&amp;lt;/math&amp;gt; columns, altogether less than &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; rows and columns. Therefore, the 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; can be covered by less than &amp;lt;math&amp;gt;a+b&amp;lt;/math&amp;gt; rows and columns, which contradicts the assumption that the size of the minimum covering of all 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a+b&amp;lt;/math&amp;gt;. Therefore, we show that &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;a&amp;lt;/math&amp;gt; independent 1&#039;s.&lt;br /&gt;
&lt;br /&gt;
By the same argument, we can show that &amp;lt;math&amp;gt;D&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; independent 1&#039;s. Since &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;D&amp;lt;/math&amp;gt; share no rows or columns, the number of independent 1&#039;s in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;r\ge a+b=s&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Dilworth&#039;s theorem ==&lt;br /&gt;
Recall that a [http://en.wikipedia.org/wiki/Partially_ordered_set &#039;&#039;&#039;partially ordered set&#039;&#039;&#039;] (or &#039;&#039;&#039;poset&#039;&#039;&#039;) consists of a set &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; and a binary relation &amp;lt;math&amp;gt;\le&amp;lt;/math&amp;gt; defined on &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;, satisfying&lt;br /&gt;
*&#039;&#039;&#039;reflexivity&#039;&#039;&#039;: &amp;lt;math&amp;gt;x\le x&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&#039;&#039;&#039;antisymmetry&#039;&#039;&#039;: &amp;lt;math&amp;gt;x\le y&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;y\le x&amp;lt;/math&amp;gt; only if &amp;lt;math&amp;gt;x=y&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&#039;&#039;&#039;transitivity&#039;&#039;&#039;: if &amp;lt;math&amp;gt;x\le y&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;y\le z&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;x\le z&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Two elements &amp;lt;math&amp;gt;x,y\in P&amp;lt;/math&amp;gt; are said to be &#039;&#039;&#039;comparable&#039;&#039;&#039; if &amp;lt;math&amp;gt;x\le y&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;y\le x&amp;lt;/math&amp;gt;; and &amp;lt;math&amp;gt;x,y\in P&amp;lt;/math&amp;gt; are &#039;&#039;&#039;incomparable&#039;&#039;&#039; if otherwise.&lt;br /&gt;
&lt;br /&gt;
A poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; is a &#039;&#039;&#039;totally ordered set&#039;&#039;&#039;, or a &#039;&#039;&#039;chain&#039;&#039;&#039;, if all pairs of elements in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; are comparable. A poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; is an &#039;&#039;&#039;antichain&#039;&#039;&#039; if all pairs of elements in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; are incomparable.&lt;br /&gt;
&lt;br /&gt;
Given a poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;, we can partition it into chains. What is the minimum number of chains that we can break &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; into? Dilworth&#039;s theorem tells us that it is equal to the size of the maximum antichain.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Dilworth&#039;s Theorem|&lt;br /&gt;
:Suppose that the largest antichain in the poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; has size &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; can be partitioned into &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; disjoint chains.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Proof|&lt;br /&gt;
Suppose &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; has an antichain &amp;lt;math&amp;gt;|A|&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; can be partitioned into disjoint chains &amp;lt;math&amp;gt;C_1,C_2,\ldots,C_s&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;|A|\le s&amp;lt;/math&amp;gt;, since every chain can pass though an antichain at most once, that is, &amp;lt;math&amp;gt;|C_i\cap A|\le 1&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,2,\ldots,s&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, we only need to prove that there exist an antichain &amp;lt;math&amp;gt;A\subseteq P&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;, and a partition of &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; into at most &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; chains.&lt;br /&gt;
&lt;br /&gt;
Define a bipartite graph &amp;lt;math&amp;gt;G(U,V,E)&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;U=V=P&amp;lt;/math&amp;gt;, and for any &amp;lt;math&amp;gt;u\in U&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v\in v&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;u&amp;lt;v&amp;lt;/math&amp;gt; in the poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;. By König-Egerváry theorem, there is a matching &amp;lt;math&amp;gt;M\subseteq E&amp;lt;/math&amp;gt; and a vertex set &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; such that every edge in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; is adjacent to at least a vertex in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;|M|=|C|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote &amp;lt;math&amp;gt;|M|=|C|=m&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the set of &#039;&#039;&#039;uncovered&#039;&#039;&#039; elements in poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;, i.e., the elements of &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; that do not correspond to any vertex in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;|A|\ge n-m&amp;lt;/math&amp;gt;. We claim that &amp;lt;math&amp;gt;A\subseteq P&amp;lt;/math&amp;gt; is an antichain. By contradiction, assume there exists &amp;lt;math&amp;gt;x,y\in A&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x&amp;lt;y&amp;lt;/math&amp;gt;. Then, by the definition of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, there exist &amp;lt;math&amp;gt;u_x\in U,v_x\in V&amp;lt;/math&amp;gt; which corresponds to &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;, and  &amp;lt;math&amp;gt;u_y\in U,v_y\in V&amp;lt;/math&amp;gt; which corresponds to &amp;lt;math&amp;gt;y&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;u_xv_y\in E&amp;lt;/math&amp;gt;. But since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; includes only those elements whose corresponding vertices are not in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, none of &amp;lt;math&amp;gt;u_x,v_x,u_y,v_y&amp;lt;/math&amp;gt; can be in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, which contradicts the fact that &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is a vertex cover of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; that every edges in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; are adjacent to at least a vertex in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; be a family of chains formed by including &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; in the same chain whenever &amp;lt;math&amp;gt;uv\in M&amp;lt;/math&amp;gt;. A moment thought would tell us that the number of chains in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; is equal to the &#039;&#039;&#039;unmatched&#039;&#039;&#039; vertices in &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; (or &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;). Thus, &amp;lt;math&amp;gt;|B|=n-m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Altogether, we construct an antichain of size &amp;lt;math&amp;gt;|A|\ge n-m&amp;lt;/math&amp;gt; and partition the poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;|B|=n-m&amp;lt;/math&amp;gt; disjoint chains. The theorem is proved.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Application: Erdős-Szekeres Theorem ===&lt;br /&gt;
Let &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; be a sequence of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct real numbers.A &#039;&#039;&#039;subsequence&#039;&#039;&#039; of &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;(a_{i_1},a_{i_2},\ldots,a_{i_k})&amp;lt;/math&amp;gt;, with &amp;lt;math&amp;gt;i_1&amp;lt;i_2&amp;lt;\cdots&amp;lt;i_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A sequence &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; is &#039;&#039;&#039;increasing&#039;&#039;&#039; if &amp;lt;math&amp;gt;a_1&amp;lt;a_2&amp;lt;\cdots&amp;lt;a_n&amp;lt;/math&amp;gt;, and &#039;&#039;&#039;decreasing&#039;&#039;&#039; if &amp;lt;math&amp;gt;a_1&amp;gt;a_2&amp;gt;\cdots&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Recall that the Erdős-Szekeres theorem states the existence of long increasing subsequence or decreasing subsequence. Last time we prove this by the pigeonhole principle. Now we use the Dilworth&#039;s theorem to prove it, which is also the original proof due to Erdős-Szekeres.&lt;br /&gt;
{{Theorem|Erdős-Szekeres Theorem|&lt;br /&gt;
:A sequence of more than &amp;lt;math&amp;gt;mn&amp;lt;/math&amp;gt; different real numbers must contain either an increasing subsequence of length &amp;lt;math&amp;gt;m+1&amp;lt;/math&amp;gt;, or a decreasing subsequence of length &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Proof by Dilworth&#039;s theorem|(Original proof of Erdős-Szekeres)&lt;br /&gt;
Let &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_N)&amp;lt;/math&amp;gt; be the sequence of &amp;lt;math&amp;gt;N&amp;gt;mn&amp;lt;/math&amp;gt; distinct real numbers. Define the poset as&lt;br /&gt;
:&amp;lt;math&amp;gt;P=\{(i,a_i)\mid i=1,2,\ldots,N\}&amp;lt;/math&amp;gt;&lt;br /&gt;
and &amp;lt;math&amp;gt;(i,a_i)\le (j,a_j)&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;i\le j&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;a_i\le a_j&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A chain &amp;lt;math&amp;gt;(i_1,a_{i_1})&amp;lt;(i_2,a_{i_2})&amp;lt;\cdots&amp;lt;(i_k,a_{i_k})&amp;lt;/math&amp;gt; must have &amp;lt;math&amp;gt;i_1&amp;lt;i_2&amp;lt;\cdots&amp;lt;i_k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;a_{i_1}&amp;lt;a_{i_2}&amp;lt;\cdots&amp;lt;a_{i_k}&amp;lt;/math&amp;gt;. Thus, each chain correspond to an increasing subsequence.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;(j_1,a_{j_1}),(j_2,a_{j_2}),\cdots,(j_k,a_{j_k})&amp;lt;/math&amp;gt; be an antichain. Without loss of generality, we can assume that &amp;lt;math&amp;gt;j_1&amp;lt;j_2&amp;lt;\cdots&amp;lt;j_k&amp;lt;/math&amp;gt;. The only case that these elements are non-comparable is that &amp;lt;math&amp;gt;a_{j_1}&amp;gt;a_{j_2}&amp;gt;\cdots&amp;gt;a_{j_k}&amp;lt;/math&amp;gt;, otherwise if &amp;lt;math&amp;gt;a_{j_s}&amp;lt; a_{j_t}&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;s&amp;lt;t&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;(j_s,a_{j_s})&amp;lt;(j_t,a_{j_t})&amp;lt;/math&amp;gt;, which contradicts the fact that it is an antichain. Thus, each antichain corresponds to a decreasing subsequence.&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; has an antichain of size &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_N)&amp;lt;/math&amp;gt; has a decreasing subsequence of size &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt;, and we are done.&lt;br /&gt;
&lt;br /&gt;
Alternatively, if the largest antichain in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; is of size at most &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, then by Dilworth&#039;s theorem, &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; can be partitioned into no more than &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; disjoint chains, due to pigeonhole principle, one of which must be of length &amp;lt;math&amp;gt;m+1&amp;lt;/math&amp;gt;, which means  &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_N)&amp;lt;/math&amp;gt; has an increasing subsequence of size &amp;lt;math&amp;gt;m+1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Application: Hall&#039;s Theorem ===&lt;br /&gt;
To recognize the power of Dilworth&#039;s theorem, we show that it contains Hall&#039;s theorem as a special case!&lt;br /&gt;
{{Theorem|Hall&#039;s Theorem |&lt;br /&gt;
:The sets &amp;lt;math&amp;gt;S_1,S_2,\ldots,S_m&amp;lt;/math&amp;gt; have a system of distinct representatives (SDR) if and only if&lt;br /&gt;
::&amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|\ge |I|&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,m\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Proof by Dilworth&#039;s theorem|&lt;br /&gt;
As we discussed before, the necessity of Hall&#039;s condition for the existence of SDR is easy. We prove its sufficiency by Dilworth&#039;s theorem.&lt;br /&gt;
&lt;br /&gt;
Denote &amp;lt;math&amp;gt;X=\bigcup_{i=1}^mS_i&amp;lt;/math&amp;gt;. Construct a poset &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; by letting &amp;lt;math&amp;gt;P=X\cup\{S_1,S_2,\ldots,S_m\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x&amp;lt;S_{i}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;x\in S_{i}&amp;lt;/math&amp;gt;. There are no other comparabilities.&lt;br /&gt;
&lt;br /&gt;
It is obvious that &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is an antichain. We claim it is also the largest one. To prove this, let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be an arbitrary antichain, and let &amp;lt;math&amp;gt;I=\{i\mid S_i\in A\}&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; contains no elements of &amp;lt;math&amp;gt;\bigcup_{i\in I}S_i&amp;lt;/math&amp;gt;, since if &amp;lt;math&amp;gt;x\in A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x\in S_i\in A&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;x&amp;lt;S_i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; cannot be an antichain. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;|A|\le |I|+|X|-\left|\bigcup_{i\in I}S_i\right|&amp;lt;/math&amp;gt;&lt;br /&gt;
and by Hall&#039;s condition &amp;lt;math&amp;gt;\left|\bigcup_{i\in I}S_i\right|\ge|I|&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;|A|\le |X|&amp;lt;/math&amp;gt;, as claimed.&lt;br /&gt;
&lt;br /&gt;
Now, Dilworth&#039;s theorem implies that &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; can be partitioned into &amp;lt;math&amp;gt;|X|&amp;lt;/math&amp;gt; chains. Since &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is an antichain and each chain can pass though an antichain on at most one element, each of the &amp;lt;math&amp;gt;|X|&amp;lt;/math&amp;gt; chains contain precisely one &amp;lt;math&amp;gt;x\in X&amp;lt;/math&amp;gt;. And since &amp;lt;math&amp;gt;\{S_1,\ldots,S_m\}&amp;lt;/math&amp;gt; is also an antichain, each of these &amp;lt;math&amp;gt;|X|&amp;lt;/math&amp;gt; chains contain at most one &amp;lt;math&amp;gt;S_i&amp;lt;/math&amp;gt;. Altogether, the &amp;lt;math&amp;gt;|X|&amp;lt;/math&amp;gt; chains are in the form:&lt;br /&gt;
:&amp;lt;math&amp;gt;\{x_1,S_1\},\{x_2,S_2\},\ldots,\{x_m,S_m\},\{x_{m+1}\}\ldots,\{x_{|X|}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since the only comparabilities in our posets are &amp;lt;math&amp;gt;x\in S_i&amp;lt;/math&amp;gt; and the above chains are disjoint, we have &amp;lt;math&amp;gt;x_1\in S_1, x_2\in S_2,\ldots,x_m\in S_m&amp;lt;/math&amp;gt; as an SDR.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Flow and Cut==&lt;br /&gt;
&lt;br /&gt;
=== Flows ===&lt;br /&gt;
An instance of the maximum flow problem consists of:&lt;br /&gt;
* a directed graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt;;&lt;br /&gt;
* two distinguished vertices &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; (the &#039;&#039;&#039;source&#039;&#039;&#039;) and &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; (the &#039;&#039;&#039;sink&#039;&#039;&#039;), where the in-degree of &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and the out-degree of &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; are both 0;&lt;br /&gt;
* the &#039;&#039;&#039;capacity function&#039;&#039;&#039;  &amp;lt;math&amp;gt;c:E\rightarrow\mathbb{R}^+&amp;lt;/math&amp;gt; which associates each directed edge &amp;lt;math&amp;gt;(u,v)\in E&amp;lt;/math&amp;gt; a nonnegative real number &amp;lt;math&amp;gt;c_{uv}&amp;lt;/math&amp;gt; called the &#039;&#039;&#039;capacity&#039;&#039;&#039; of the edge.&lt;br /&gt;
&lt;br /&gt;
The quadruple &amp;lt;math&amp;gt;(G,c,s,t)&amp;lt;/math&amp;gt; is called a &#039;&#039;&#039;flow network&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
A function &amp;lt;math&amp;gt;f:E\rightarrow\mathbb{R}^+&amp;lt;/math&amp;gt; is called a &#039;&#039;&#039;flow&#039;&#039;&#039; (to be specific an &#039;&#039;&#039;&amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flow&#039;&#039;&#039;) in the network &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; if it satisfies:&lt;br /&gt;
* &#039;&#039;&#039;Capacity constraint:&#039;&#039;&#039; &amp;lt;math&amp;gt;f_{uv}\le c_{uv}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;(u,v)\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
* &#039;&#039;&#039;Conservation constraint:&#039;&#039;&#039; &amp;lt;math&amp;gt;\sum_{u:(u,v)\in E}f_{uv}=\sum_{w:(v,w)\in E}f_{vw}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V\setminus\{s,t\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;value&#039;&#039;&#039; of the flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;\sum_{v:(s,v)\in E}f_{sv}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Given a flow network, the maximum flow problem asks to find the flow of the maximum value.&lt;br /&gt;
&lt;br /&gt;
The maximum flow problem can be described as the following linear program.&lt;br /&gt;
:&amp;lt;math&amp;gt;\text{maximize} \quad \sum_{v:(s,v)\in E}f_{sv}&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\text{s.t.} &lt;br /&gt;
&amp;amp;&amp;amp;f_{uv} &amp;amp;\le c_{uv} &amp;amp;\quad&amp;amp; \forall (u,v)\in E\\&lt;br /&gt;
&amp;amp;&amp;amp;\sum_{u:(u,v)\in E}f_{uv}-\sum_{w:(v,w)\in E}f_{vw} &amp;amp;=0 &amp;amp;\quad&amp;amp; \forall v\in V\setminus\{s,t\}\\&lt;br /&gt;
&amp;amp;&amp;amp;f_{uv} &amp;amp;\ge 0 &amp;amp;\quad&amp;amp; \forall (u,v)\in E&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Cuts ===&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;(G(V,E),c,s,t)&amp;lt;/math&amp;gt; be a flow network. Let &amp;lt;math&amp;gt;S\subset V&amp;lt;/math&amp;gt;. We call &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; an &#039;&#039;&#039;&amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut&#039;&#039;&#039; if &amp;lt;math&amp;gt;s\in S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;t\not\in S&amp;lt;/math&amp;gt;.&lt;br /&gt;
:The &#039;&#039;&#039;value&#039;&#039;&#039; of  the cut (also called the &#039;&#039;&#039;capacity&#039;&#039;&#039; of the cut) is defined as &amp;lt;math&amp;gt;\sum_{u\in S,v\not\in S\atop (u,v)\in E}c_{uv}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
A fundamental fact in flow theory is that cuts always upper bound flows.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;(G(V,E),c,s,t)&amp;lt;/math&amp;gt; be a flow network. Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be an arbitrary flow in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; be an arbitrary &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\sum_{v:(s,v)}f_{sv}\le \sum_{u\in S,v\not\in S\atop (u,v)\in E}c_{uv}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:that is, the value of any flow is no greater than the value of any cut.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|By the definition of &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut, &amp;lt;math&amp;gt;s\in S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;t\not\in S&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the conservation of flow, &lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{u\in S}\left(\sum_{v:(u,v)\in E}f_{uv}-\sum_{v:(v,u)\in E}f_{vu}\right)=\sum_{v:(s,v)\in E}f_{sv}+\sum_{u\in S\setminus\{s\}}\left(\sum_{v:(u,v)\in E}f_{uv}-\sum_{v:(v,u)\in E}f_{vu}\right)=\sum_{v:(s,v)\in E}f_{sv}\,.&amp;lt;/math&amp;gt;&lt;br /&gt;
On the other hand, summing flow over edges,&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v\in S}\left(\sum_{u:(u,v)\in E}f_{uv}-\sum_{u:(v,u)\in E}f_{vu}\right)=\sum_{u\in S,v\in S\atop (u,v)\in E}\left(f_{uv}-f_{uv}\right)+\sum_{u\in S,v\not\in S\atop (u,v)\in E}f_{uv}-\sum_{u\in S,v\not\in S\atop (v,u)\in E}f_{vu}=\sum_{u\in S,v\not\in S\atop (u,v)\in E}f_{uv}-\sum_{u\in S,v\not\in S\atop (v,u)\in E}f_{vu}\,.&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v:(s,v)\in E}f_{sv}=\sum_{u\in S,v\not\in S\atop (u,v)\in E}f_{uv}-\sum_{u\in S,v\not\in S\atop (v,u)\in E}f_{vu}\le\sum_{u\in S,v\not\in S\atop (u,v)\in E}f_{uv}\le  \sum_{u\in S,v\not\in S\atop (u,v)\in E}c_{uv}\,,&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Augmenting paths ===&lt;br /&gt;
{{Theorem|Definition (Augmenting path)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a flow in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. An &#039;&#039;&#039;augmenting path to &amp;lt;math&amp;gt;u_k&amp;lt;/math&amp;gt;&#039;&#039;&#039; is a sequence of distinct vertices &amp;lt;math&amp;gt;P=(u_0,u_1,\cdots, u_k)&amp;lt;/math&amp;gt;, such that &lt;br /&gt;
:* &amp;lt;math&amp;gt;u_0=s\,&amp;lt;/math&amp;gt;;&lt;br /&gt;
:and each pair of consecutive vertices &amp;lt;math&amp;gt;u_{i}u_{i+1}\,&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; corresponds to either a &#039;&#039;&#039;forward edge&#039;&#039;&#039; &amp;lt;math&amp;gt;(u_{i},u_{i+1})\in E&amp;lt;/math&amp;gt; or a &#039;&#039;&#039;reverse edge&#039;&#039;&#039; &amp;lt;math&amp;gt;(u_{i+1},u_{i})\in E&amp;lt;/math&amp;gt;, and &lt;br /&gt;
:* &amp;lt;math&amp;gt;f(u_i,u_{i+1})&amp;lt;c(u_i,u_{i+1})\,&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;u_{i}u_{i+1}\,&amp;lt;/math&amp;gt; corresponds to a forward edge &amp;lt;math&amp;gt;(u_{i},u_{i+1})\in E&amp;lt;/math&amp;gt;, and &lt;br /&gt;
:* &amp;lt;math&amp;gt;f(u_{i+1},u_i)&amp;gt;0\,&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;u_{i}u_{i+1}\,&amp;lt;/math&amp;gt; corresponds to a reverse edge &amp;lt;math&amp;gt;(u_{i+1},u_{i})\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
:If &amp;lt;math&amp;gt;u_k=t\,&amp;lt;/math&amp;gt;, we simply call &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; an &#039;&#039;&#039;augmenting path&#039;&#039;&#039;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a flow in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Suppose there is an augmenting path &amp;lt;math&amp;gt;P=u_0u_1\cdots u_k&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;u_0=s&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u_k=t&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt; be a positive constant satisfying &lt;br /&gt;
*&amp;lt;math&amp;gt;\epsilon \le c(u_{i},u_{i+1})-f(u_i,u_{i+1})&amp;lt;/math&amp;gt; for all forward edges &amp;lt;math&amp;gt;(u_{i},u_{i+1})\in E&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&amp;lt;math&amp;gt;\epsilon \le f(u_{i+1},u_i)&amp;lt;/math&amp;gt; for all reverse edges &amp;lt;math&amp;gt;(u_{i+1},u_i)\in E&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the definition of augmenting path, we can always find such a positive &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Increase &amp;lt;math&amp;gt;f(u_i,u_{i+1})&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; for all forward edges &amp;lt;math&amp;gt;(u_{i},u_{i+1})\in E&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt; and decrease &amp;lt;math&amp;gt;f(u_{i+1},u_i)&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; for all reverse edges &amp;lt;math&amp;gt;(u_{i+1},u_i)\in E&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;P&amp;lt;/math&amp;gt;. Denote the modified flow by &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt;. It can be verified that &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt; satisfies the capacity constraint and conservation constraint thus is still a valid flow. On the other hand, the value of the new flow &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v:(s,v)\in E}f_{sv}&#039;=\epsilon+\sum_{v:(s,v)\in E}f_{sv}&amp;gt;\sum_{v:(s,v)\in E}f_{sv}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, the value of the flow can be &amp;quot;augmented&amp;quot; by adjusting the flow on the augmenting path. This immediately implies that if a flow is maximum, then there is no augmenting path. Surprisingly, the converse is also true, thus maximum flows are &amp;quot;characterized&amp;quot; by augmenting paths.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:A flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is maximum if and only if there are no augmenting paths.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|We have already proved the &amp;quot;only if&amp;quot; direction above. Now we prove the &amp;quot;if&amp;quot; direction.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{u\in V\mid \exists\text{an augmenting path to }u\}&amp;lt;/math&amp;gt;. Clearly &amp;lt;math&amp;gt;s\in S&amp;lt;/math&amp;gt;, and since there is no augmenting path &amp;lt;math&amp;gt;t\not\in S&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; defines an &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut. &lt;br /&gt;
&lt;br /&gt;
We claim that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v:(s,v)}f_{sv}= \sum_{u\in S,v\not\in S\atop (u,v)\in E}c_{uv}&amp;lt;/math&amp;gt;,&lt;br /&gt;
that is, the value of flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; approach the value of the cut &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; defined above. By the above lemma, this will imply that the current flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is maximum.&lt;br /&gt;
&lt;br /&gt;
To prove this claim, we first observe that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v:(s,v)}f_{sv}= \sum_{u\in S,v\not\in S\atop (u,v)\in E}f_{uv}-\sum_{u\in S,v\not\in S\atop (v,u)\in E}f_{vu}&amp;lt;/math&amp;gt;.&lt;br /&gt;
This identity is implied by the flow conservation constraint, and holds for any &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We then claim that &lt;br /&gt;
*&amp;lt;math&amp;gt;f_{uv}=c_{uv}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;u\in S,v\not\in S, (u,v)\in E&amp;lt;/math&amp;gt;; and &lt;br /&gt;
*&amp;lt;math&amp;gt;f_{vu}=0&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;u\in S,v\not\in S, (v,u)\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
If otherwise, then the augmenting path to &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; apending &amp;lt;math&amp;gt;uv&amp;lt;/math&amp;gt; becomes a new augmenting path to &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, which contradicts that &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; includes all vertices to which there exist augmenting paths.&lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v:(s,v)}f_{sv}= \sum_{u\in S,v\not\in S\atop (u,v)\in E}f_{uv}-\sum_{u\in S,v\not\in S\atop (v,u)\in E}f_{vu} = \sum_{u\in S,v\not\in S\atop (u,v)\in E}c_{uv}&amp;lt;/math&amp;gt;.&lt;br /&gt;
As discussed above, this proves the theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Max-Flow Min-Cut ==&lt;br /&gt;
&lt;br /&gt;
=== The max-flow min-cut theorem ===&lt;br /&gt;
{{Theorem|Max-Flow Min-Cut Theorem|&lt;br /&gt;
:In a flow network, the maximum value of any &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flow equals the minimum value of any &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be a flow with maximum value, so there is no augmenting path.&lt;br /&gt;
&lt;br /&gt;
Again, let &lt;br /&gt;
:&amp;lt;math&amp;gt;S=\{u\in V\mid \exists\text{an augmenting path to }u\}&amp;lt;/math&amp;gt;. &lt;br /&gt;
As proved above, &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; forms an &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut, and&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v:(s,v)}f_{sv}= \sum_{u\in S,v\not\in S\atop (u,v)\in E}c_{uv}&amp;lt;/math&amp;gt;,&lt;br /&gt;
that is, the value of flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; equals the value of cut &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Since we know that all &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flows are not greater than any &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut, the value of flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; equals the minimum value of any &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Flow Integrality Theorem ===&lt;br /&gt;
{{Theorem|Flow Integrality Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;(G,c,s,t)&amp;lt;/math&amp;gt; be a flow network with integral capacity &amp;lt;math&amp;gt;c&amp;lt;/math&amp;gt;. There exists an integral flow which is maximum.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; be an integral flow of maximum value. If there is an augmenting path, since both &amp;lt;math&amp;gt;c&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; are integral, a new flow can be constructed of value 1+the value of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;, contradicting that &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is maximum over all integral flows. Therefore, there is no augmenting path, which means that &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is maximum over all flows, integral or not.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Applications: Menger&#039;s theorem ===&lt;br /&gt;
Given an undirected graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; and two distinct vertices &amp;lt;math&amp;gt;s,t\in V&amp;lt;/math&amp;gt;, a set of edges &amp;lt;math&amp;gt;C\subseteq E&amp;lt;/math&amp;gt; is called an &#039;&#039;&#039;&amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut&#039;&#039;&#039;, if deleting edges in &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; disconnects &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A simple path from &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; is called an &#039;&#039;&#039;&amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; path&#039;&#039;&#039;. Two paths are said to be &#039;&#039;&#039;edge-disjoint&#039;&#039;&#039; if they do not share any edge.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Menger 1927)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be an arbitrary undirected graph and &amp;lt;math&amp;gt;s,t\in V&amp;lt;/math&amp;gt; be two distinct vertices. The minimum size of any &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut equals the maximum number of edge-disjoint &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; paths.&lt;br /&gt;
}}&lt;br /&gt;
{{proof|&lt;br /&gt;
Construct a directed graph &amp;lt;math&amp;gt;G&#039;(V,E&#039;)&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; as follows: replace every undirected edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;s,t\not\in\{u,v\}&amp;lt;/math&amp;gt; by two directed edges &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(v,u)&amp;lt;/math&amp;gt;; replace every undirected edge &amp;lt;math&amp;gt;sv\in E&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;(s,v)&amp;lt;/math&amp;gt;, and very undirected edge &amp;lt;math&amp;gt;vt\in E&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;(v,t)&amp;lt;/math&amp;gt;. Then assign every directed edge with capacity 1.&lt;br /&gt;
&lt;br /&gt;
It is easy to verify that in the flow network constructed as above, the followings hold:&lt;br /&gt;
*An integral &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flow corresponds to a set of &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; paths in the undirected graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, where the value of the flow is the number of &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; paths.&lt;br /&gt;
*An &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut in the flow network corresponds to an &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut in the undirected graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with the same value.&lt;br /&gt;
The Menger&#039;s theorem follows as a direct consequence of the max-flow min-cut theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Applications: König-Egerváry theorem ===&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph. An edge set &amp;lt;math&amp;gt;M\subseteq E&amp;lt;/math&amp;gt; is called a &#039;&#039;&#039;matching&#039;&#039;&#039; if no edge in &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; shares any vertex. A vertex set &amp;lt;math&amp;gt;C\subseteq V&amp;lt;/math&amp;gt; is called a &#039;&#039;&#039;vertex cover&#039;&#039;&#039; if for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, either &amp;lt;math&amp;gt;u\in C&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;v\in C&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (König 1936)|&lt;br /&gt;
:In any bipartite graph &amp;lt;math&amp;gt;G(V_1,V_2,E)&amp;lt;/math&amp;gt;, the size of a &#039;&#039;maximum&#039;&#039; matching equals the size of a &#039;&#039;minimum&#039;&#039; vertex cover.&lt;br /&gt;
}}&lt;br /&gt;
We now show how a reduction of bipartite matchings to flows. &lt;br /&gt;
&lt;br /&gt;
Construct a flow network &amp;lt;math&amp;gt;(G&#039;(V,E&#039;),c,s,t)&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
* &amp;lt;math&amp;gt;V=V_1\cup V_2\cup\{s,t\}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; are two new vertices.&lt;br /&gt;
* For ever &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, add &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;E&#039;&amp;lt;/math&amp;gt;; for every &amp;lt;math&amp;gt;u\in V_1&amp;lt;/math&amp;gt;, add &amp;lt;math&amp;gt;(s,u)&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;E&#039;&amp;lt;/math&amp;gt;; and for every &amp;lt;math&amp;gt;v\in V_2&amp;lt;/math&amp;gt;, add &amp;lt;math&amp;gt;(v,t)&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;E&#039;&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Let &amp;lt;math&amp;gt;c_{su}=1&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;u\in V_1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;c_{vt}=1&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;v\in V_2&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;c_{uv}=\infty&amp;lt;/math&amp;gt; for every bipartite edges &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:The size of a maximum matching in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is equal to the value of a maximum &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flow in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{proof|&lt;br /&gt;
Given an integral &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;M=\{uv\in E\mid f_{uv}=1\}&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; must be a matching since for every &amp;lt;math&amp;gt;u\in V_1&amp;lt;/math&amp;gt;. To see this, observe that there is at most one &amp;lt;math&amp;gt;v\in V_2&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;f_{uv}=1&amp;lt;/math&amp;gt;, because of that &amp;lt;math&amp;gt;f_{su}\le c_{su}=1&amp;lt;/math&amp;gt; and conservation of flows. The same holds for vertices in &amp;lt;math&amp;gt;V_2&amp;lt;/math&amp;gt; by the same argument. Therefore, each flow corresponds to a matching.&lt;br /&gt;
&lt;br /&gt;
Given a matching &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; in bipartite graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, define an integral flow &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; as such: for &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;f_{uv}=1&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;uv\in M&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_{uv}=0&amp;lt;/math&amp;gt; if otherwise; for &amp;lt;math&amp;gt;u\in V_1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;f_{su}=1&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;uv\in M&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_{su}=0&amp;lt;/math&amp;gt; if otherwise; for &amp;lt;math&amp;gt;v\in V_2&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;f_{vt}=1&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;uv\in M&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;f_{vt}=0&amp;lt;/math&amp;gt; if otherwise.&lt;br /&gt;
&lt;br /&gt;
It is easy to check that &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is valid &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; flow in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt;. Therefore, there is an one-one correspondence between flows in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; and matchings in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. The lemma follows naturally.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We then establish a correspondence between &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cuts in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; and vertex covers in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:The size of a minimum vertex cover in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is equal to the value of a minimum &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut of minimum capacity in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\sum_{u\in S, v\not\in S\atop (u,v)\in E&#039;}c_{uv}&amp;lt;/math&amp;gt; must be finite since &amp;lt;math&amp;gt;S=\{s\}&amp;lt;/math&amp;gt; gives us an &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut whose capacity is &amp;lt;math&amp;gt;|V_1|&amp;lt;/math&amp;gt; which is finite. Therefore, no edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;u\in V_1\cap S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v\in V_2\setminus S&amp;lt;/math&amp;gt;, i.e., for all &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, either &amp;lt;math&amp;gt;u\in V_1\setminus S&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;v\in V_2\cap S&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;(V_1\setminus S)\cup(V_2\cap S)&amp;lt;/math&amp;gt; is a vertex cover in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, whose size is&lt;br /&gt;
:&amp;lt;math&amp;gt;|(V_1\setminus S)\cup(V_2\cap S)|=|V_1\setminus S|+|V_2\cap S|=\sum_{u\in V_1\setminus S}c_{su}+\sum_{v\in V_2\cap S}c_{ut}=\sum_{u\in S,v\not\in S\atop (u,v)\in E&#039;}c_{uv}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The last term is the capacity of the minimum &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;-&amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; cut &amp;lt;math&amp;gt;(S,\bar{S})&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The König-Egerváry theorem then holds as a consequence of the max-flow min-cut theorem.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13832</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13832"/>
		<updated>2026-06-17T04:04:50Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/04/21)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第二次作业已发布&amp;lt;/font&amp;gt;，请在 2026/05/13 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A2.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/05/22)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第三次作业已发布&amp;lt;/font&amp;gt;，请在 2026/06/03 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A3.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 2|Problem Set 2]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 3|Problem Set 3]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal graph theory|Extremal graph theory | 极值图论]] ([http://tcs.nju.edu.cn/slides/comb2026/ExtremalGraphs.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal set theory|Extremal set theory | 极值集合论]]（[http://tcs.nju.edu.cn/slides/comb2026/ExtremalSets.pdf slides]）&lt;br /&gt;
#* [https://mathweb.ucsd.edu/~ronspubs/90_03_erdos_ko_rado.pdf Old and new proofs of the Erdős–Ko–Rado theorem] by Frankl and Graham&lt;br /&gt;
#* An [http://tcs.nju.edu.cn/slides/comb2026/sunflower-note.pdf LLM-generated lecture note] on Alweiss-Lovet-Wu-Zhang&#039;s improvement over the sunflower lemma, with simplified proofs by Rao-Tao&lt;br /&gt;
# [[组合数学 (Fall 2026)/Ramsey theory|Ramsey theory | Ramsey理论]]（[http://tcs.nju.edu.cn/slides/comb2026/Ramsey.pdf slides]）&lt;br /&gt;
# [[组合数学 (Fall 2026)/Matching theory|Matching theory | 匹配论]]（[http://tcs.nju.edu.cn/slides/comb2026/Matchings.pdf slides]）&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Ramsey_theory&amp;diff=13754</id>
		<title>组合数学 (Fall 2026)/Ramsey theory</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Ramsey_theory&amp;diff=13754"/>
		<updated>2026-05-20T13:33:09Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== Ramsey&amp;#039;s Theorem == === Ramsey&amp;#039;s theorem for graph === {{Theorem|Ramsey&amp;#039;s Theorem| :Let &amp;lt;math&amp;gt;k,\ell&amp;lt;/math&amp;gt; be positive integers. Then there exists an integer &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; satisfying: :If &amp;lt;math&amp;gt;n\ge R(k,\ell)&amp;lt;/math&amp;gt;, for any coloring of edges of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; with two colors red and blue, there exists a red &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; or a blue &amp;lt;math&amp;gt;K_\ell&amp;lt;/math&amp;gt;. }} {{Proof| We show that &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is finite by induction on &amp;lt;math&amp;gt;k+\ell&amp;lt;/math&amp;gt;. For the...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Ramsey&#039;s Theorem ==&lt;br /&gt;
=== Ramsey&#039;s theorem for graph ===&lt;br /&gt;
{{Theorem|Ramsey&#039;s Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;k,\ell&amp;lt;/math&amp;gt; be positive integers. Then there exists an integer &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; satisfying:&lt;br /&gt;
:If &amp;lt;math&amp;gt;n\ge R(k,\ell)&amp;lt;/math&amp;gt;, for any coloring of edges of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; with two colors red and blue, there exists a red &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; or a blue &amp;lt;math&amp;gt;K_\ell&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We show that &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is finite by induction on &amp;lt;math&amp;gt;k+\ell&amp;lt;/math&amp;gt;. For the base case, it is easy to verify that&lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,1)=R(1,\ell)=1&amp;lt;/math&amp;gt;.&lt;br /&gt;
For general &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;, we will show that &lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,\ell)\le R(k,\ell-1)+R(k-1,\ell)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose we have a two coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;n=R(k,\ell-1)+R(k-1,\ell)&amp;lt;/math&amp;gt;. Take an arbitrary vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, and split &amp;lt;math&amp;gt;V\setminus\{v\}&amp;lt;/math&amp;gt; into two subsets &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt;, where&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
S&amp;amp;=\{u\in V\setminus\{v\}\mid uv \text{ is blue }\}\\&lt;br /&gt;
T&amp;amp;=\{u\in V\setminus\{v\}\mid uv \text{ is red }\}&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
Since &lt;br /&gt;
:&amp;lt;math&amp;gt;|S|+|T|+1=n=R(k,\ell-1)+R(k-1,\ell)&amp;lt;/math&amp;gt;,&lt;br /&gt;
we have either &amp;lt;math&amp;gt;|S|\ge R(k,\ell-1)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;|T|\ge R(k-1,\ell)&amp;lt;/math&amp;gt;. By symmetry, suppose &amp;lt;math&amp;gt;|S|\ge R(k,\ell-1)&amp;lt;/math&amp;gt;. By induction hypothesis, the complete subgraph defined on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; has either a red &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;, in which case we are done; or a blue &amp;lt;math&amp;gt;K_{\ell-1}&amp;lt;/math&amp;gt;, in which case the complete subgraph defined on &amp;lt;math&amp;gt;S\cup{v}&amp;lt;/math&amp;gt; must have a blue &amp;lt;math&amp;gt;K_\ell&amp;lt;/math&amp;gt; since all edges from &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; to vertices in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; are blue.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Ramsey&#039;s Theorem (graph, multicolor)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;r, k_1,k_2,\ldots,k_r&amp;lt;/math&amp;gt; be positive integers. Then there exists an integer &amp;lt;math&amp;gt;R(r;k_1,k_2,\ldots,k_r)&amp;lt;/math&amp;gt; satisfying:&lt;br /&gt;
:For any &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-coloring of a complete graph of &amp;lt;math&amp;gt;n\ge R(r;k_1,k_2,\ldots,k_r)&amp;lt;/math&amp;gt; vertices, there exists a monochromatic &amp;lt;math&amp;gt;k_i&amp;lt;/math&amp;gt;-clique with the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th color for some &amp;lt;math&amp;gt;i\in\{1,2,\ldots,r\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma (the &amp;quot;mixing color&amp;quot; trick)|&lt;br /&gt;
:&amp;lt;math&amp;gt;R(r;k_1,k_2,\ldots,k_r)\le R(r-1;k_1,k_2,\ldots,k_{r-2},R(2;k_{r-1},k_r))&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We transfer the &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-coloring to &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;-coloring by identifying the &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;th and the &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;th colors. &lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;n\ge R(r-1;k_1,k_2,\ldots,k_{r-2},R(2;k_{r-1},k_r))&amp;lt;/math&amp;gt;, then for any &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt;, there either exist an &amp;lt;math&amp;gt;i\in\{1,2,\ldots,r-2\}&amp;lt;/math&amp;gt; and a &amp;lt;math&amp;gt;k_i&amp;lt;/math&amp;gt;-clique which is monochromatically colored with the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th color; or exists clique of &amp;lt;math&amp;gt;R(2;k_{r-1},k_r)&amp;lt;/math&amp;gt; vertices which is monochromatically colored with the mixed color of the original &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;th and &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;th colors, which again implies that there exists either a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-clique which is monochromatically colored with the original &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;th color, or a &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;-clique which is monochromatically colored with the original &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;th color. This implies the recursion.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Ramsey number ===&lt;br /&gt;
The smallest number &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; satisfying the condition in the Ramsey theory is called the &#039;&#039;&#039;Ramsey number&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Alternatively, we can define &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; as the smallest &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; such that if &amp;lt;math&amp;gt;n\ge N&amp;lt;/math&amp;gt;, for any 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; in red and blue, there is either a red &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; or a blue &amp;lt;math&amp;gt;K_\ell&amp;lt;/math&amp;gt;. The Ramsey theorem is stated as:&lt;br /&gt;
:&amp;quot;&#039;&#039;&amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is finite for any positive integers &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;.&#039;&#039;&amp;quot;&lt;br /&gt;
&lt;br /&gt;
The core of the inductive proof of the Ramsey theorem is the following recursion&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
R(k,1) &amp;amp;=R(1,\ell)=1\\&lt;br /&gt;
R(k,\ell) &amp;amp;\le R(k,\ell-1)+R(k-1,\ell).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
From this recursion, we can deduce an upper bound for the Ramsey number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,\ell)\le{k+\ell-2\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|It is easy to verify the bound by induction.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
------&lt;br /&gt;
&lt;br /&gt;
The following theorem is due to Spencer in 1975, which is the best known lower bound for diagonal Ramsey number.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Spencer 1975)|&lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,k)\ge Ck2^{k/2}&amp;lt;/math&amp;gt; for some constant &amp;lt;math&amp;gt;C&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Its proof uses the Lovász local lemma in the probabilistic method.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lovász Local Lemma (symmetric case)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; be a set of events, and assume that the following hold:&lt;br /&gt;
:#for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\Pr[A_i]\le p&amp;lt;/math&amp;gt;;&lt;br /&gt;
:# each event &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; is independent of all but at most &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; other events, and&lt;br /&gt;
:::&amp;lt;math&amp;gt;ep(d+1)\le 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We can use the local lemma to prove a lower bound for the diagonal Ramsey number.&lt;br /&gt;
{{Proof|&lt;br /&gt;
To prove a lower bound &amp;lt;math&amp;gt;R(k,k)&amp;gt;n&amp;lt;/math&amp;gt;, it is sufficient to show that there exists a 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; without a monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;. We prove this by the probabilistic method.&lt;br /&gt;
&lt;br /&gt;
Pick a random 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; by coloring each edge uniformly and independently with one of the two colors. For any set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vertices, let &amp;lt;math&amp;gt;A_S&amp;lt;/math&amp;gt; denote the event that &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; forms a monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;. It is easy to see that &amp;lt;math&amp;gt;\Pr[A_s]=2^{1-{k\choose 2}}=p&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subset &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; of vertices, &amp;lt;math&amp;gt;A_S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;A_T&amp;lt;/math&amp;gt; are dependent if and only if &amp;lt;math&amp;gt;|S\cap T|\ge 2&amp;lt;/math&amp;gt;. For each &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, the number of &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|S\cap T|\ge 2&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;{k\choose 2}{n\choose k-2}&amp;lt;/math&amp;gt;, so the max degree of the dependency graph is &amp;lt;math&amp;gt;d\le{k\choose 2}{n\choose k-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;n=Ck2^{k/2}&amp;lt;/math&amp;gt; for some appropriate constant &amp;lt;math&amp;gt;C&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathrm{e}p(d+1)&lt;br /&gt;
&amp;amp;\le \mathrm{e}2^{1-{k\choose 2}}\left({k\choose 2}{n\choose k-2}+1\right)\\&lt;br /&gt;
&amp;amp;\le 2^{3-{k\choose 2}}{k\choose 2}{n\choose k-2}\\&lt;br /&gt;
&amp;amp;\le 1&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying the local lemma, the probability that there is no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{S\in{[n]\choose k}}\overline{A_S}\right]&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, there exists a 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; which has no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;, which means&lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,k)&amp;gt;n=Ck2^{k/2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;\Omega\left(k2^{k/2}\right)\le R(k,k)\le{2k-2\choose k-1}=O\left(k^{-1/2}4^{k}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
! &#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;&#039;&#039;,&#039;&#039;&amp;lt;math&amp;gt;l&amp;lt;/math&amp;gt;&#039;&#039;&lt;br /&gt;
! 1&lt;br /&gt;
! 2&lt;br /&gt;
! 3&lt;br /&gt;
! 4&lt;br /&gt;
! 5&lt;br /&gt;
! 6&lt;br /&gt;
! 7&lt;br /&gt;
! 8&lt;br /&gt;
! 9&lt;br /&gt;
! 10&lt;br /&gt;
|-&lt;br /&gt;
! 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
| 1&lt;br /&gt;
|-&lt;br /&gt;
! 2&lt;br /&gt;
| 1&lt;br /&gt;
| 2&lt;br /&gt;
| 3&lt;br /&gt;
| 4&lt;br /&gt;
| 5&lt;br /&gt;
| 6&lt;br /&gt;
| 7&lt;br /&gt;
| 8&lt;br /&gt;
| 9&lt;br /&gt;
| 10&lt;br /&gt;
|-&lt;br /&gt;
! 3&lt;br /&gt;
| 1&lt;br /&gt;
| 3&lt;br /&gt;
| 6&lt;br /&gt;
| 9&lt;br /&gt;
| 14&lt;br /&gt;
| 18&lt;br /&gt;
| 23&lt;br /&gt;
| 28&lt;br /&gt;
| 36&lt;br /&gt;
| 40–43&lt;br /&gt;
|-&lt;br /&gt;
! 4&lt;br /&gt;
| 1&lt;br /&gt;
| 4&lt;br /&gt;
| 9&lt;br /&gt;
| 18&lt;br /&gt;
| 25&lt;br /&gt;
| 35–41&lt;br /&gt;
| 49–61&lt;br /&gt;
| 56–84&lt;br /&gt;
| 73–115&lt;br /&gt;
| 92–149&lt;br /&gt;
|-&lt;br /&gt;
! 5&lt;br /&gt;
| 1&lt;br /&gt;
| 5&lt;br /&gt;
| 14&lt;br /&gt;
| 25&lt;br /&gt;
| 43–49&lt;br /&gt;
| 58–87&lt;br /&gt;
| 80–143&lt;br /&gt;
| 101–216&lt;br /&gt;
| 125–316&lt;br /&gt;
| 143–442&lt;br /&gt;
|-&lt;br /&gt;
! 6&lt;br /&gt;
| 1&lt;br /&gt;
| 6&lt;br /&gt;
| 18&lt;br /&gt;
| 35–41&lt;br /&gt;
| 58–87&lt;br /&gt;
| 102–165&lt;br /&gt;
| 113–298&lt;br /&gt;
| 127–495&lt;br /&gt;
| 169–780&lt;br /&gt;
| 179–1171&lt;br /&gt;
|-&lt;br /&gt;
! 7&lt;br /&gt;
| 1&lt;br /&gt;
| 7&lt;br /&gt;
| 23&lt;br /&gt;
| 49–61&lt;br /&gt;
| 80–143&lt;br /&gt;
| 113–298&lt;br /&gt;
| 205–540&lt;br /&gt;
| 216–1031&lt;br /&gt;
| 233–1713&lt;br /&gt;
| 289–2826&lt;br /&gt;
|-&lt;br /&gt;
! 8&lt;br /&gt;
| 1&lt;br /&gt;
| 8&lt;br /&gt;
| 28&lt;br /&gt;
| 56–84&lt;br /&gt;
| 101–216&lt;br /&gt;
| 127–495&lt;br /&gt;
| 216–1031&lt;br /&gt;
| 282–1870&lt;br /&gt;
| 317–3583&lt;br /&gt;
| 317-6090&lt;br /&gt;
|-&lt;br /&gt;
! 9&lt;br /&gt;
| 1&lt;br /&gt;
| 9&lt;br /&gt;
| 36&lt;br /&gt;
| 73–115&lt;br /&gt;
| 125–316&lt;br /&gt;
| 169–780&lt;br /&gt;
| 233–1713&lt;br /&gt;
| 317–3583&lt;br /&gt;
| 565–6588&lt;br /&gt;
| 580–12677&lt;br /&gt;
|-&lt;br /&gt;
! 10&lt;br /&gt;
| 1&lt;br /&gt;
| 10&lt;br /&gt;
| 40–43&lt;br /&gt;
| 92–149&lt;br /&gt;
| 143–442&lt;br /&gt;
| 179–1171&lt;br /&gt;
| 289–2826&lt;br /&gt;
| 317-6090&lt;br /&gt;
| 580–12677&lt;br /&gt;
| 798–23556&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Ramsey&#039;s theorem for hypergraph ===&lt;br /&gt;
{{Theorem|Ramsey&#039;s Theorem (hypergraph, multicolor)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;r, t, k_1,k_2,\ldots,k_r&amp;lt;/math&amp;gt; be positive integers. Then there exists an integer &amp;lt;math&amp;gt;R_t(r;k_1,k_2,\ldots,k_r)&amp;lt;/math&amp;gt; satisfying:&lt;br /&gt;
:For any &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-coloring of &amp;lt;math&amp;gt;{[n]\choose t}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n\ge R_t(r;k_1,k_2,\ldots,k_r)&amp;lt;/math&amp;gt;,  there exist an &amp;lt;math&amp;gt;i\in\{1,2,\ldots,r\}&amp;lt;/math&amp;gt; and  a subset &amp;lt;math&amp;gt;X\subseteq [n]&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|X|\ge k_i&amp;lt;/math&amp;gt; such that all members of &amp;lt;math&amp;gt;{X\choose t}&amp;lt;/math&amp;gt; are colored with the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th color.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;n\rightarrow(k_1,k_2,\ldots,k_r)^t&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma (the &amp;quot;mixing color&amp;quot; trick)|&lt;br /&gt;
:&amp;lt;math&amp;gt;R_t(r;k_1,k_2,\ldots,k_r)\le R_t(r-1;k_1,k_2,\ldots,k_{r-2},R_t(2;k_{r-1},k_r))&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is then sufficient to prove the Ramsey&#039;s theorem for the two-coloring of a hypergraph, that is, to prove &amp;lt;math&amp;gt;R_t(k,\ell)=R_t(2;k,\ell)&amp;lt;/math&amp;gt; is finite.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:&amp;lt;math&amp;gt;R_t(k,\ell)\le R_{t-1}(R_t(k-1,\ell),R_t(k,\ell-1))+1&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;n=R_{t-1}(R_t(k-1,\ell),R_t(k,\ell-1))+1&amp;lt;/math&amp;gt;. Denote &amp;lt;math&amp;gt;[n]=\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;f:{[n]\choose t}\rightarrow\{{\color{red}\text{red}},{\color{blue}\text{blue}}\}&amp;lt;/math&amp;gt; be an arbitrary 2-coloring of &amp;lt;math&amp;gt;{[n]\choose t}&amp;lt;/math&amp;gt;. It is then sufficient to show that there either exists an &amp;lt;math&amp;gt;X\subseteq[n]&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|X|=k&amp;lt;/math&amp;gt; such that all members of &amp;lt;math&amp;gt;{X\choose t}&amp;lt;/math&amp;gt; are colored red by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;; or exists an &amp;lt;math&amp;gt;X\subseteq[n]&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|X|=\ell&amp;lt;/math&amp;gt; such that all members of &amp;lt;math&amp;gt;{X\choose t}&amp;lt;/math&amp;gt; are colored blue by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We remove &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt; and define a new coloring &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;{[n-1]\choose t-1}&amp;lt;/math&amp;gt; by&lt;br /&gt;
:&amp;lt;math&amp;gt;f&#039;(A)=f(A\cup\{n\})&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;A\in{[n-1]\choose t-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the choice of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; and by symmetry, there exists a subset &amp;lt;math&amp;gt;S\subseteq[n-1]&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|X|=R_t(k-1,\ell)&amp;lt;/math&amp;gt; such that all members of &amp;lt;math&amp;gt;{S\choose t-1}&amp;lt;/math&amp;gt; are colored with red by &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt;. Then there either exists an &amp;lt;math&amp;gt;X\subseteq S&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|X|=\ell&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;{X\choose t}&amp;lt;/math&amp;gt; is colored all blue by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;, in which case we are done; or exists an &amp;lt;math&amp;gt;X\subseteq S&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|X|=k-1&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;{X\choose t}&amp;lt;/math&amp;gt; is colored all red by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;. Next we prove that in the later case &amp;lt;math&amp;gt;{X\cup{n}\choose t}&amp;lt;/math&amp;gt; is all red, which will close our proof. Since all &amp;lt;math&amp;gt;A\in{S\choose t-1}&amp;lt;/math&amp;gt; are colored with red by &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt;, then by our definition of &amp;lt;math&amp;gt;f&#039;&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;f(A\cup\{n\})={\color{red}\text{red}}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;A\in {X\choose t-1}\subseteq{S\choose t-1}&amp;lt;/math&amp;gt;. Recalling that &amp;lt;math&amp;gt;{X\choose t}&amp;lt;/math&amp;gt; is colored all red by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;{X\cup\{n\}\choose t}&amp;lt;/math&amp;gt; is colored all red by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; and we are done.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
==  Applications of Ramsey Theorem ==&lt;br /&gt;
=== The &amp;quot;Happy Ending&amp;quot; problem ===&lt;br /&gt;
{{Theorem|The happy ending problem|&lt;br /&gt;
:Any set of 5 points in the plane, no three on a line, has a subset of 4 points that form the vertices of a convex quadrilateral.&lt;br /&gt;
}}&lt;br /&gt;
See the article&lt;br /&gt;
[http://www.maa.org/mathland/mathtrek_10_3_00.html] for the proof.&lt;br /&gt;
&lt;br /&gt;
We say a set of points in the plane in [http://en.wikipedia.org/wiki/General_position &#039;&#039;&#039;general positions&#039;&#039;&#039;] if no three of the points are on the same line.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Erdős-Szekeres 1935)|&lt;br /&gt;
:For any positive integer &amp;lt;math&amp;gt;m\ge 3&amp;lt;/math&amp;gt;, there is an &amp;lt;math&amp;gt;N(m)&amp;lt;/math&amp;gt; such that any set of at least &amp;lt;math&amp;gt;N(m)&amp;lt;/math&amp;gt; points in general position in the plane (i.e., no three of the points are on a line) contains &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; points that are the vertices of a convex &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-gon.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;N(m)=R_3(m,m)&amp;lt;/math&amp;gt;. For &amp;lt;math&amp;gt;n\ge N(m)&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be an arbitrary set of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; points in the plane, no three of which are on a line. Define a 2-coloring of the 3-subsets of points &amp;lt;math&amp;gt;f:{X\choose 3}\rightarrow\{0,1\}&amp;lt;/math&amp;gt; as follows: for any &amp;lt;math&amp;gt;\{a,b,c\}\in{X\choose 3}&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;\triangle_{abc}\subset X&amp;lt;/math&amp;gt; be the set of points covered by the triangle &amp;lt;math&amp;gt;abc&amp;lt;/math&amp;gt;; and &amp;lt;math&amp;gt;f(\{a,b,c\})=|\triangle_{abc}|\bmod 2&amp;lt;/math&amp;gt;, that is, &amp;lt;math&amp;gt;f(\{a,b,c\})&amp;lt;/math&amp;gt; indicates the oddness of the number of points covered by the triangle &amp;lt;math&amp;gt;abc&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Since &amp;lt;math&amp;gt;|X|\ge R_3(m,m)&amp;lt;/math&amp;gt;, there exists a &amp;lt;math&amp;gt;Y\subseteq X&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;|Y|=m&amp;lt;/math&amp;gt; and all members of &amp;lt;math&amp;gt;{Y\choose 3}&amp;lt;/math&amp;gt; are colored with the same value by &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
We claim that the &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; points in &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; are the vertices of a convex &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-gon. If otherwise, by the definition of convexity, there exist &amp;lt;math&amp;gt;\{a,b,c,d\}\subseteq Y&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;d\in\triangle_{abc}&amp;lt;/math&amp;gt;. Since no three points are in the same line, &lt;br /&gt;
:&amp;lt;math&amp;gt;\triangle_{abc}=\triangle_{abd}\cup\triangle_{acd}\cup\triangle_{bcd}\cup\{d\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where all unions are disjoint. Then &amp;lt;math&amp;gt;|\triangle_{abc}|=|\triangle_{abd}|+|\triangle_{acd}|+|\triangle_{bcd}|+1&amp;lt;/math&amp;gt;, which implies that &amp;lt;math&amp;gt;f(\{a,b,c\}), f(\{a,b,d\}), f(\{a,c,d\}), f(\{b,c,d\})\,&amp;lt;/math&amp;gt; cannot be equal, contradicting that all members of &amp;lt;math&amp;gt;{Y\choose 3}&amp;lt;/math&amp;gt; have the same color.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Yao&#039;s lower bound for implicit data structures ===&lt;br /&gt;
We consider the following fundamental problem of &#039;&#039;&#039;membership query&#039;&#039;&#039;.&lt;br /&gt;
{{Theorem|Membership Query|&lt;br /&gt;
:&#039;&#039;&#039;Input&#039;&#039;&#039;: A data set &amp;lt;math&amp;gt;S\subset U&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; is a data universe of size &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&#039;&#039;&#039;Query&#039;&#039;&#039;: a data item (also called a &#039;&#039;&#039;key&#039;&#039;&#039;) &amp;lt;math&amp;gt;x\in U&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&#039;&#039;&#039;Answer&#039;&#039;&#039;: Whether &amp;lt;math&amp;gt;x\in S&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
This is a basic problem for data structures. People want to design efficient data structures to store the data set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; so that the query &amp;quot;Is &amp;lt;math&amp;gt;x\in S&amp;lt;/math&amp;gt;?&amp;quot; can be efficiently answered by accessing the data structure as little as possible in the worst case.&lt;br /&gt;
&lt;br /&gt;
A &#039;&#039;&#039;sorted table&#039;&#039;&#039; for a data set &amp;lt;math&amp;gt;S\subset [N]&amp;lt;/math&amp;gt; is a natural data structure in which the elements of &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;x_1&amp;lt;x_2&amp;lt;\cdots&amp;lt;x_n&amp;lt;/math&amp;gt;, are stored in an array, one element &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt; in each entry, in the increasing order.&lt;br /&gt;
For a sorted table, the membership query problem can be solved via &#039;&#039;&#039;binary search&#039;&#039;&#039; within &amp;lt;math&amp;gt;\Omega(\log_2 n)&amp;lt;/math&amp;gt; memory accesses in the worst case. The following [https://dl.acm.org/doi/pdf/10.1145/322261.322274 fundamental result of Andrew Chi-Chih Yao (姚期智)] shows that this is the best possible for sorted tables. The proof is an elegant application of the &#039;&#039;&#039;adversarial argument&#039;&#039;&#039;(对手论证).&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma (Yao 1981)|&lt;br /&gt;
:Suppose that &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;N\ge 2n-1&amp;lt;/math&amp;gt;, the data universe is &amp;lt;math&amp;gt;U=[N]&amp;lt;/math&amp;gt;, and the size of the data set is &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:If the data structure is a &#039;&#039;&#039;sorted table&#039;&#039;&#039;, any search algorithm requires at least &amp;lt;math&amp;gt;\lceil\log_2 (n+1)\rceil&amp;lt;/math&amp;gt; accesses to the data structure in the worst case.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We will show by an adversarial argument that &amp;lt;math&amp;gt;\lceil\log_2 (n+1)\rceil&amp;lt;/math&amp;gt; accesses are required to search for the key value &amp;lt;math&amp;gt;x=n&amp;lt;/math&amp;gt; in the universe &amp;lt;math&amp;gt;[N]=\{1,2,\ldots,N\}&amp;lt;/math&amp;gt;. The construction of the adversarial data set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is by induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;n=2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;N\ge 2n-1=3&amp;lt;/math&amp;gt; it is easy to see that 2 memory accesses are required to make sure whether the key value &amp;lt;math&amp;gt;n=2&amp;lt;/math&amp;gt; presents in a sorted table containing 2 keys out of a data universe of size 3, in the worst case.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;n&amp;gt;2&amp;lt;/math&amp;gt;. Assume the induction hypothesis for all smaller &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. We will prove it for the size of data set &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, size of universe &amp;lt;math&amp;gt;N\ge 2n-1&amp;lt;/math&amp;gt; and the search key &amp;lt;math&amp;gt;x=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose that the first access position is &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. The adversary chooses the table content &amp;lt;math&amp;gt;T[k]&amp;lt;/math&amp;gt;. The adversary&#039;s strategy is:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
T[k]=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
k &amp;amp; k\le \frac{n}{2},\\&lt;br /&gt;
N-(n-k) &amp;amp; k&amp;gt; \frac{n}{2}.&lt;br /&gt;
\end{cases}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By symmetry, suppose it is the first case that &amp;lt;math&amp;gt;k\le \frac{n}{2}&amp;lt;/math&amp;gt;.  Then the key &amp;lt;math&amp;gt;x=n&amp;lt;/math&amp;gt; may be in any position &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;\frac{n}{2}+1\le i\le n&amp;lt;/math&amp;gt;. In fact, &amp;lt;math&amp;gt;T\left[ \left\lceil \frac{n}{2}\right\rceil +1\right]&amp;lt;/math&amp;gt; through &amp;lt;math&amp;gt;T[n]&amp;lt;/math&amp;gt; is a sorted table of size &amp;lt;math&amp;gt;n&#039;=\left\lfloor \frac{n}{2}\right\rfloor&amp;lt;/math&amp;gt; which may contain any &amp;lt;math&amp;gt;n&#039;&amp;lt;/math&amp;gt;-subset of &amp;lt;math&amp;gt;\left\{\left\lceil \frac{n}{2}\right\rceil+1, \left\lceil \frac{n}{2}\right\rceil+2,\ldots,N\right\}&amp;lt;/math&amp;gt;, and hence, in particular, any &amp;lt;math&amp;gt;n&#039;&amp;lt;/math&amp;gt;-subset of the new universe&lt;br /&gt;
:&amp;lt;math&amp;gt;U&#039;=\left\{\left\lceil \frac{n}{2}\right\rceil+1, \left\lceil \frac{n}{2}\right\rceil+2,\ldots,N-\left\lceil \frac{n}{2}\right\rceil\right\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The size &amp;lt;math&amp;gt;N&#039;&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;U&#039;&amp;lt;/math&amp;gt; satisfies&lt;br /&gt;
:&amp;lt;math&amp;gt;N&#039;=N-2\left\lceil \frac{n}{2}\right\rceil\ge 2(n-1)-2\left\lceil \frac{n}{2}\right\rceil \ge 2\left\lfloor \frac{n}{2}\right\rfloor-1= 2n&#039;-1&amp;lt;/math&amp;gt;,&lt;br /&gt;
and the desired key &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; has the relative value &amp;lt;math&amp;gt;x&#039;=n- \left\lceil \frac{n}{2}\right\rceil=\left\lfloor \frac{n}{2}\right\rfloor=n&#039;&amp;lt;/math&amp;gt; in the new universe &amp;lt;math&amp;gt;U&#039;&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
By the induction hypothesis, &amp;lt;math&amp;gt;\lceil\log_2 (n&#039;+1)\rceil&amp;lt;/math&amp;gt; more memory accesses will be required. Hence the total number of memory accesses is at least &lt;br /&gt;
:&amp;lt;math&amp;gt;1+\lceil\log_2 (n&#039;+1)\rceil=1+\left\lceil\log_2 \left(\left\lfloor \frac{n}{2}\right\rfloor+1\right)\right\rceil\ge \lceil\log_2 (n+1)\rceil&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If the first access is &amp;lt;math&amp;gt;k&amp;gt; \frac{n}{2}&amp;lt;/math&amp;gt;, we symmetrically get that &amp;lt;math&amp;gt;T[1]&amp;lt;/math&amp;gt; through &amp;lt;math&amp;gt;T\left[\left\lfloor \frac{n}{2}\right\rfloor\right]&amp;lt;/math&amp;gt; is a sorted table of size &amp;lt;math&amp;gt;n&#039;=\left\lfloor \frac{n}{2}\right\rfloor&amp;lt;/math&amp;gt; which may contain any &amp;lt;math&amp;gt;n&#039;&amp;lt;/math&amp;gt;-subset of the universe&lt;br /&gt;
:&amp;lt;math&amp;gt;U&#039;=\left\{\left\lceil \frac{n}{2}\right\rceil+1, \left\lceil \frac{n}{2}\right\rceil+2,\ldots,N-\left\lceil \frac{n}{2}\right\rceil\right\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The rest is the same as before. This completes the induction.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We have seen that on a sorted table, there is no search algorithm outperforming the binary search in the worst case.&lt;br /&gt;
Our question is:&lt;br /&gt;
:&#039;&#039;Is there any other order than the increasing order, on which there is a better search algorithm?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
An &#039;&#039;&#039;implicit data structure&#039;&#039;&#039; use no extra space in addition to the original data set, thus a data structure can only be represented &#039;&#039;implicitly&#039;&#039; by the order of the data items in the table. That is, each data set is stored as a permutation of the set. Formally, an implicit data structure is described by a function&lt;br /&gt;
:&amp;lt;math&amp;gt;f:{U\choose n}\rightarrow[n!]&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;\pi\in[n!]&amp;lt;/math&amp;gt; specify a permutation of the sorted table, and a data set &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;x_1&amp;lt;x_2&amp;lt;\cdots&amp;lt;x_n&amp;lt;/math&amp;gt; is stored as an array &amp;lt;math&amp;gt;(x_{\pi(1)},x_{\pi(2)},\ldots,x_{\pi(n)}\}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;\pi=f(S)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Thus, the sorted table is the simplest implicit data structure, in which &amp;lt;math&amp;gt;f(S)&amp;lt;/math&amp;gt; always gives the identity permutation for all &amp;lt;math&amp;gt;S\in{U\choose n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We observe that if &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; maps all data sets &amp;lt;math&amp;gt;S\in{U\choose n}&amp;lt;/math&amp;gt; to the same permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, then the data structure is equivalent to the sorted table, under the bijection that the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th entry in the sorted table corresponds to the &amp;lt;math&amp;gt;\pi(i)&amp;lt;/math&amp;gt;th entry of the actual array, where the same &amp;lt;math&amp;gt;\Omega(\log_2 n)&amp;lt;/math&amp;gt; lower bound applies.&lt;br /&gt;
&lt;br /&gt;
This observation can be generalized and made formal as follows.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Observation|&lt;br /&gt;
:If there is a sub-universe &amp;lt;math&amp;gt;X\subseteq U&amp;lt;/math&amp;gt; such that for every data set &amp;lt;math&amp;gt;S\in {X\choose n}&amp;lt;/math&amp;gt;, the implicit data structure &amp;lt;math&amp;gt;f(S)=\pi&amp;lt;/math&amp;gt; stores the data set using the the same permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, i.e.&lt;br /&gt;
::&amp;lt;math&amp;gt;f\left({X\choose n}\right)=\{\pi\}&amp;lt;/math&amp;gt;&lt;br /&gt;
:then this implicit data structure is equivalent to the sorted table for all data sets from the new universe &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, under the bijection that the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th entry in the sorted table corresponds to the &amp;lt;math&amp;gt;\pi(i)&amp;lt;/math&amp;gt;th entry of the array.&lt;br /&gt;
:Therefore, if &amp;lt;math&amp;gt;|X|\ge 2n&amp;lt;/math&amp;gt;, then the same &amp;lt;math&amp;gt;\Omega(\log_2 n)&amp;lt;/math&amp;gt; lower bound for searching in a sorted table applies.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Due to Ramsey theorem, for sufficiently large &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; which satisfies &amp;lt;math&amp;gt;N\ge R_{n}(n!;2n)&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;f:{U\choose n}\rightarrow[n!]&amp;lt;/math&amp;gt;, there is an &amp;lt;math&amp;gt;X\subseteq U&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;|X|\ge 2n&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;\left|f\left({X\choose n}\right)\right|=1&amp;lt;/math&amp;gt;, which guarantees the existence of the sub-universe &amp;lt;math&amp;gt;X\subseteq U&amp;lt;/math&amp;gt; required in the above observation for (wildly) large universe sizes &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt;, which implies the following lower bound for implicit data structures.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Yao 1981)|&lt;br /&gt;
:Suppose that &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;N\ge 2n&amp;lt;/math&amp;gt;, the data universe is &amp;lt;math&amp;gt;U=[N]&amp;lt;/math&amp;gt;, and the size of the data set is &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:For any &#039;&#039;&#039;implicit data structure&#039;&#039;&#039;, if &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; is sufficiently large, then any search algorithm requires at least &amp;lt;math&amp;gt;\lfloor\log_2 n\rfloor&amp;lt;/math&amp;gt; accesses to the data structure in the worst case.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Linial&#039;s lower bound for local computation===&lt;br /&gt;
In the studies of &#039;&#039;&#039;local computation&#039;&#039;&#039; (initiated by [https://www.cs.huji.ac.il/~nati/PAPERS/locality_dist_graph_algs.pdf Linial] and [https://www.wisdom.weizmann.ac.il/~naor/PAPERS/lcl.pdf Naor and Stockmeyer]), people wants to answer questions like:&lt;br /&gt;
::&#039;&#039;Can locally defined problems be computed locally?&#039;&#039;&lt;br /&gt;
In general, the answer is no to the above question. A famous example is Linial&#039;s lower bound for &#039;&#039;&#039;maximal independent set&#039;&#039;&#039; (&#039;&#039;&#039;MIS&#039;&#039;&#039;) in a ring.&lt;br /&gt;
&lt;br /&gt;
Consider a very simple distributed network, a ring that contains &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; nodes, where each node is assigned a unique ID from &amp;lt;math&amp;gt;[n]=\{1,2,\ldots, n\}&amp;lt;/math&amp;gt;. The labeled network is then described by a tuple &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; of IDs, which is a permutation of &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;a_i\in [n]&amp;lt;/math&amp;gt; gives the ID of the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th node in the ring.&lt;br /&gt;
&lt;br /&gt;
In a distributed algorithm, in each round, every node communicates with its 2 neighbors in the ring, and when the algorithm terminates, each node returns its local output. For example, in the MIS problem, the goal of the algorithm is to construct a maximal independent set: upon termination, each node returns a bit to indicate whether the node is in the constructed independent set. And the output gives a correct MIS as long as it satisfies both the followings: &lt;br /&gt;
* there are no two consecutive nodes in the ring both outputting 1;&lt;br /&gt;
* there are no three consecutive nodes in the ring all outputting 0.&lt;br /&gt;
This is clearly a locally defined problem. In fact, it is a constraint satisfaction problem (CSP) where each constraint only involves 1-local or 2-local neighborhood.&lt;br /&gt;
&lt;br /&gt;
We are interested in the distributed algorithms that can always produce the correct answer, and want to prove a lower bound for the number of rounds required in the worst case by such distributed algorithms.&lt;br /&gt;
&lt;br /&gt;
As a local distributed algorithm, each node &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; initially does not know anything beyond its local information, which is just its own ID &amp;lt;math&amp;gt;a_i\in [n]&amp;lt;/math&amp;gt;. &lt;br /&gt;
And after &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; rounds, information-theoretically, each node &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; can at best know all information within its &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;-local neighborhood, which is represented by the &amp;lt;math&amp;gt;(2t-1)&amp;lt;/math&amp;gt;-tuple &amp;lt;math&amp;gt;(a_{i-t},\ldots,a_{i-1},a_i,a_{i+1},\ldots a_{i+t})&amp;lt;/math&amp;gt;, with the addition/subtraction modulo &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; along the ring.&lt;br /&gt;
&lt;br /&gt;
This suggests us to define such a computational model for local distributed algorithms: any &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;-round local algorithm is described by a function&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{L}:[n]^{2t+1}\to\{0,1\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th node in the ring, its output is given by &amp;lt;math&amp;gt;\mathcal{L}(a_{i-t},\ldots,a_{i-1},a_i,a_{i+1},\ldots a_{i+t})&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt; represents the ID of the &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;th node in the ring, with the addition/subtraction modulo &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
As a correct algorithm for constructing MIS, it must hold that any three consecutive nodes can never output the same value. We then have the following lower bound for &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; for such algorithms.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Linial 1992)|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;-round local algorithm for maximal independent set (MIS) on a ring of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; nodes, it holds that&lt;br /&gt;
:: &amp;lt;math&amp;gt;t=\Omega(\log^*n)&amp;lt;/math&amp;gt;&lt;br /&gt;
:where &amp;lt;math&amp;gt;\log^*n&amp;lt;/math&amp;gt; represents the [https://en.wikipedia.org/wiki/Iterated_logarithm iterated logarithm], which is the number of times the logarithm function must be iteratively applied to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; before the result is less than or equal to 1.&lt;br /&gt;
}}&lt;br /&gt;
This lower bound shows that even on very simple network like ring, some very basic locally defined problem (MIS) cannot be computed locally (within constant locality).&lt;br /&gt;
&lt;br /&gt;
The original proof of Linial relies on chromatic number of so-called neighborhood graphs. Here we give an alternative proof based on Ramsey theorem found by Baruch Awerbuch.&lt;br /&gt;
{{Proof|&lt;br /&gt;
As we discussed earlier, any &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;-round local algorithm can be represented by a mapping&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{L}:[n]^{2t+1}\to\{0,1\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
It can naturally defines a 2-coloring:&lt;br /&gt;
:&amp;lt;math&amp;gt;f:{[n]\choose {2t+1}}\to\{0,1\}&amp;lt;/math&amp;gt;&lt;br /&gt;
by the following construction: for any &amp;lt;math&amp;gt;\{a_1,a_2,\ldots,a_{2t+1}\}\subseteq [n]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;a_1&amp;lt;a_2&amp;lt;\cdots&amp;lt;a_{2t+1}&amp;lt;/math&amp;gt;, we define&lt;br /&gt;
:&amp;lt;math&amp;gt;f(\{a_1,a_2,\ldots,a_{2t+1}\})=\mathcal{L}(a_1,a_2,\ldots,a_{2t+1})&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Ramsey theorem, for &amp;lt;math&amp;gt;n\ge R_{2t+1}(2;2t+3,2t+3)&amp;lt;/math&amp;gt;, there exists a subset &amp;lt;math&amp;gt;\{a_1,a_2,\ldots,a_{2t+3}\}\subseteq [n]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;a_1&amp;lt;a_2&amp;lt;\cdots&amp;lt;a_{2t+3}&amp;lt;/math&amp;gt;, such that &lt;br /&gt;
:&amp;lt;math&amp;gt;\left|f\left({S\choose {2t+1}}\right)\right|=1&amp;lt;/math&amp;gt;.&lt;br /&gt;
By our construction of &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;, this means&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{L}(a_1,a_2,\ldots,a_{2t+1})=\mathcal{L}(a_2,a_3,\ldots,a_{2t+2})=\mathcal{L}(a_3,a_4,\ldots,a_{2t+3})&amp;lt;/math&amp;gt;,&lt;br /&gt;
which contradicts to that the output of &amp;lt;math&amp;gt;\mathcal{L}&amp;lt;/math&amp;gt; should indicate an MIS, on any ring with &amp;lt;math&amp;gt;2t+3&amp;lt;/math&amp;gt; consecutive nodes labeled as &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_{2t+3})&amp;lt;/math&amp;gt;, because on such rings, there would be 3 consecutive nodes with the same output bit.&lt;br /&gt;
&lt;br /&gt;
Therefore, any &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;-round local algorithm that can always correctly produce an MIS on a ring, must satisfies that&lt;br /&gt;
:&amp;lt;math&amp;gt;n&amp;lt;R_{2t+1}(2;2t+3,2t+3)\le \underbrace{2^{2^{\unicode{x22F0}^{2}}}}_{ct}&amp;lt;/math&amp;gt;,&lt;br /&gt;
for some constant &amp;lt;math&amp;gt;c&amp;gt;0&amp;lt;/math&amp;gt;, whose inverse function gives the lower bound&lt;br /&gt;
:&amp;lt;math&amp;gt;t=\Omega(\log^*n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13753</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13753"/>
		<updated>2026-05-20T13:32:42Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/04/21)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第二次作业已发布&amp;lt;/font&amp;gt;，请在 2026/05/13 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A2.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 2|Problem Set 2]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal graph theory|Extremal graph theory | 极值图论]] ([http://tcs.nju.edu.cn/slides/comb2026/ExtremalGraphs.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal set theory|Extremal set theory | 极值集合论]]（[http://tcs.nju.edu.cn/slides/comb2026/ExtremalSets.pdf slides]）&lt;br /&gt;
#* [https://mathweb.ucsd.edu/~ronspubs/90_03_erdos_ko_rado.pdf Old and new proofs of the Erdős–Ko–Rado theorem] by Frankl and Graham&lt;br /&gt;
#* An [http://tcs.nju.edu.cn/slides/comb2026/sunflower-note.pdf LLM-generated lecture note] on Alweiss-Lovet-Wu-Zhang&#039;s improvement over the sunflower lemma, with simplified proofs by Rao-Tao&lt;br /&gt;
# [[组合数学 (Fall 2026)/Ramsey theory|Ramsey theory | Ramsey理论]]（[http://tcs.nju.edu.cn/slides/comb2026/Ramsey.pdf slides]）&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Extremal_set_theory&amp;diff=13736</id>
		<title>组合数学 (Fall 2026)/Extremal set theory</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Extremal_set_theory&amp;diff=13736"/>
		<updated>2026-05-13T08:49:22Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== Sunflowers == An set system is a &amp;#039;&amp;#039;&amp;#039;sunflower&amp;#039;&amp;#039;&amp;#039; if all its member sets intersect at the same set of elements. {{Theorem|Definition (sunflower)| : A set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; is a &amp;#039;&amp;#039;&amp;#039;sunflower&amp;#039;&amp;#039;&amp;#039; of size &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; with a &amp;#039;&amp;#039;&amp;#039;core&amp;#039;&amp;#039;&amp;#039; &amp;lt;math&amp;gt;C\subseteq X&amp;lt;/math&amp;gt; if  ::&amp;lt;math&amp;gt;\forall S,T\in\mathcal{F}&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;S\neq T&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;S\cap T=C&amp;lt;/math&amp;gt;. }} Note that we do not require the core to be nonempty, thus a family of disjoint sets is...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Sunflowers ==&lt;br /&gt;
An set system is a &#039;&#039;&#039;sunflower&#039;&#039;&#039; if all its member sets intersect at the same set of elements.&lt;br /&gt;
{{Theorem|Definition (sunflower)|&lt;br /&gt;
: A set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; is a &#039;&#039;&#039;sunflower&#039;&#039;&#039; of size &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; with a &#039;&#039;&#039;core&#039;&#039;&#039; &amp;lt;math&amp;gt;C\subseteq X&amp;lt;/math&amp;gt; if &lt;br /&gt;
::&amp;lt;math&amp;gt;\forall S,T\in\mathcal{F}&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;S\neq T&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;S\cap T=C&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
Note that we do not require the core to be nonempty, thus a family of disjoint sets is also a sunflower (with the core &amp;lt;math&amp;gt;\emptyset&amp;lt;/math&amp;gt;).&lt;br /&gt;
&lt;br /&gt;
The next result due to Erdős and Rado, called the sunflower lemma, is a famous result in extremal set theory, and has some important applications in Boolean circuit complexity.&lt;br /&gt;
{{Theorem|Sunflower Lemma (Erdős-Rado)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;|\mathcal{F}|&amp;gt;k!(r-1)^k&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; contains a sunflower of size  &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We proceed by induction on &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. For &amp;lt;math&amp;gt;k=1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathcal{F}\subseteq{X\choose 1}&amp;lt;/math&amp;gt;, thus all sets in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; are disjoint. And since &amp;lt;math&amp;gt;|\mathcal{F}|&amp;gt;r-1&amp;lt;/math&amp;gt;, we can choose &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; of these sets and form a sunflower.&lt;br /&gt;
&lt;br /&gt;
Now let &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt; and assume the lemma holds for all smaller &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. Take a maximal family &amp;lt;math&amp;gt;\mathcal{G}\subseteq \mathcal{F}&amp;lt;/math&amp;gt; whose members are disjoint, i.e. for any &amp;lt;math&amp;gt;S,T\in \mathcal{G}&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;S\neq T&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;S\cap T=\emptyset&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;|\mathcal{G}|\ge r&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;\mathcal{G}&amp;lt;/math&amp;gt; is a sunflower of size at least &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and we are done.&lt;br /&gt;
&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;|\mathcal{G}|\le r-1&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;Y=\bigcup_{S\in\mathcal{G}}S&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;|Y|=k|\mathcal{G}|\le k(r-1)&amp;lt;/math&amp;gt; (since all members of &amp;lt;math&amp;gt;\mathcal{G}&amp;lt;/math&amp;gt;) are disjoint). We claim that &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; intersets all members of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;, since if otherwise, there exists an &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;S\cap Y=\emptyset&amp;lt;/math&amp;gt;, then we can enlarge &amp;lt;math&amp;gt;\mathcal{G}&amp;lt;/math&amp;gt; by adding &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;\mathcal{G}&amp;lt;/math&amp;gt; and still have all members of &amp;lt;math&amp;gt;\mathcal{G}&amp;lt;/math&amp;gt; disjoint, which contradicts the assumption that &amp;lt;math&amp;gt;\mathcal{G}&amp;lt;/math&amp;gt; is the maximum of such families.&lt;br /&gt;
&lt;br /&gt;
By the pigeonhole principle, some elements &amp;lt;math&amp;gt;y\in Y&amp;lt;/math&amp;gt; must contained in at least&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{|\mathcal{F}|}{|Y|}&amp;gt;\frac{k!(r-1)^k}{k(r-1)}=(k-1)!(r-1)^{k-1}&amp;lt;/math&amp;gt;&lt;br /&gt;
members of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;. We delete this &amp;lt;math&amp;gt;y&amp;lt;/math&amp;gt; from these sets and consider the family &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{H}=\{S\setminus\{y\}\mid S\in\mathcal{F}\wedge y\in S\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have &amp;lt;math&amp;gt;\mathcal{H}\subseteq {X\choose k-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|\mathcal{H}|&amp;gt;(k-1)!(r-1)^{k-1}&amp;lt;/math&amp;gt;, thus by the induction hypothesis, &amp;lt;math&amp;gt;\mathcal{H}&amp;lt;/math&amp;gt;contains a sunflower of size &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;. Adding &amp;lt;math&amp;gt;y&amp;lt;/math&amp;gt; to the members of this sunflower, we get the desired sunflower in the original family &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
==The Erdős–Ko–Rado Theorem ==&lt;br /&gt;
A set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; is called &#039;&#039;&#039;intersecting&#039;&#039;&#039;, if for any &amp;lt;math&amp;gt;S,T\in\mathcal{F}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;S\cap T\neq\emptyset&amp;lt;/math&amp;gt;. A natural question of extremal favor is: &amp;quot;how large can an intersecting family be?&amp;quot;&lt;br /&gt;
&lt;br /&gt;
Assume &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt;. When &amp;lt;math&amp;gt;n&amp;lt;2k&amp;lt;/math&amp;gt;, every pair of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subsets of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; intersects. So the non-trivial case is when &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;. The famous Erdős–Ko–Rado theorem gives the largest possible cardinality of a nontrivially intersecting family. &lt;br /&gt;
&lt;br /&gt;
According to Erdős, the theorem itself was proved in 1938, but was not published until 23 years later.&lt;br /&gt;
{{Theorem|Erdős–Ko–Rado theorem (proved in 1938, published in 1961)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|\mathcal{F}|\le{n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Katona&#039;s proof ===&lt;br /&gt;
We first introduce a proof discovered by Katona in 1972. The proof uses double counting.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; be a &#039;&#039;&#039;cyclic permutation&#039;&#039;&#039; of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, that is, we think of assigning &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; in a circle and ignore the rotations of the circle. It is easy to see that there are &amp;lt;math&amp;gt;(n-1)!&amp;lt;/math&amp;gt; cyclic permutations of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set (each cyclic permutation corresponds to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; permutations).&lt;br /&gt;
Let &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{G}_\pi=\{\{\pi_{(i+j)\bmod n}\mid j\in[k]\}\mid i\in [n]\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The next lemma states the following observation: in a circle of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; points, supposed &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;, there can be at most &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; arcs, each consisting of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; points, such that every pair of arcs share at least one point.&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting, then for any cyclic permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;|\mathcal{G}_\pi\cap\mathcal{F}|\le k&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Fix a cyclic permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;A_i=\{\pi_{(i+j+n)\bmod n}\mid j\in[k]\}&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\mathcal{G}_\pi&amp;lt;/math&amp;gt; can be written as &amp;lt;math&amp;gt;\mathcal{G}_\pi=\{A_i\mid i\in [n]\}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Suppose that &amp;lt;math&amp;gt;A_t\in\mathcal{F}&amp;lt;/math&amp;gt;. Since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting, the only sets &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; that can be in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; other than &amp;lt;math&amp;gt;A_t&amp;lt;/math&amp;gt; itself are the &amp;lt;math&amp;gt;2k-2&amp;lt;/math&amp;gt; sets &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;t-(k-1)\le i\le t+k-1, i\neq t&amp;lt;/math&amp;gt;. We partition these sets into &amp;lt;math&amp;gt;k-1&amp;lt;/math&amp;gt; pairs &amp;lt;math&amp;gt;\{A_i,A_{i+k}\}&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;t-(k-1)\le i\le t-1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Note that for &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;A_i\cap C_{i+k}=\emptyset&amp;lt;/math&amp;gt;. Since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting, &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; can contain at most one set of each such pair. The lemma follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The Katona&#039;s proof of Erdős–Ko–Rado theorem is done by counting in two ways the pairs of member &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; and cyclic permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which contain &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; as a continuous path on the circle (i.e., an arc).&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Katona&#039;s proof of Erdős–Ko–Rado theorem|(double counting)&lt;br /&gt;
Let &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{R}=\{(S,\pi)\mid \pi \text{ is a cyclic permutation of }X, \text{and }S\in\mathcal{F}\cap\mathcal{G}_\pi\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We count &amp;lt;math&amp;gt;\mathcal{R}&amp;lt;/math&amp;gt; in two ways.&lt;br /&gt;
&lt;br /&gt;
First, due to the lemma, &amp;lt;math&amp;gt;|\mathcal{F}\cap\mathcal{G}_\pi|\le k&amp;lt;/math&amp;gt; for any cyclic permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;. There are &amp;lt;math&amp;gt;(n-1)!&amp;lt;/math&amp;gt; cyclic permutations in total. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{R}|=\sum_{\text{cyclic }\pi}|\mathcal{F}\cap\mathcal{G}_\pi|\le k(n-1)!&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Next, for each &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt;, the number of cyclic permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; in which &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is continuous is &amp;lt;math&amp;gt;|S|!(n-|S|)!=k!(n-k)!&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{R}|=\sum_{S\in\mathcal{F}}k!(n-k)!=|\mathcal{F}|k!(n-k)!&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Altogether, we have &lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|\le\frac{k(n-1)!}{k!(n-k)!}=\frac{(n-1)!}{(k-1)!(n-k)!}={n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős&#039; shifting technique ===&lt;br /&gt;
We now introduce the original proof of the Erdős–Ko–Rado theorem, which uses a technique called &#039;&#039;&#039;shifting&#039;&#039;&#039; (originally called &#039;&#039;&#039;compression&#039;&#039;&#039;).&lt;br /&gt;
&lt;br /&gt;
Without loss of generality, we assume &amp;lt;math&amp;gt;X=[n]&amp;lt;/math&amp;gt;, and restate the Erdős–Ko–Rado theorem as follows.&lt;br /&gt;
{{Theorem|Erdős–Ko–Rado theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {[n]\choose k}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting, then &amp;lt;math&amp;gt;|\mathcal{F}|\le{n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We define a &#039;&#039;&#039;shift operator&#039;&#039;&#039; for the set family.&lt;br /&gt;
{{Theorem|Definition (shift operator)|&lt;br /&gt;
: Assume &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^{[n]}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;0\le i&amp;lt;j\le n-1&amp;lt;/math&amp;gt;. Define the &#039;&#039;&#039;&amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shift&#039;&#039;&#039; &amp;lt;math&amp;gt;S_{ij}&amp;lt;/math&amp;gt; as an operator on &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
:*for each &amp;lt;math&amp;gt;T\in\mathcal{F}&amp;lt;/math&amp;gt;, write &amp;lt;math&amp;gt;T_{ij}=(T\setminus\{j\})\cup\{i\} &amp;lt;/math&amp;gt;, and let&lt;br /&gt;
::&amp;lt;math&amp;gt;S_{ij}(T)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
T_{ij} &amp;amp; \mbox{if }j\in T, i\not\in T, \mbox{ and }T_{ij} \not\in\mathcal{F},\\&lt;br /&gt;
T &amp;amp; \mbox{otherwise;}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
:* let &amp;lt;math&amp;gt;S_{ij}(\mathcal{F})=\{S_{ij}(T)\mid T\in \mathcal{F}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the following propositions of shifts.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
# &amp;lt;math&amp;gt;|S_{ij}(T)|=|T|\,&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|S_{ij}(\mathcal{F})|=\mathcal{F}&amp;lt;/math&amp;gt;;&lt;br /&gt;
# if &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting, then so is &amp;lt;math&amp;gt;S_{ij}(\mathcal{F})&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
(1) is quite obvious. Now we prove (2).&lt;br /&gt;
&lt;br /&gt;
Consider any &amp;lt;math&amp;gt;A,B\in\mathcal{F}&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;A\cap B&amp;lt;/math&amp;gt; includes any element other than &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are still intersecting after &amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shift. Thus without loss of generality, we may consider the only unsafe case where &amp;lt;math&amp;gt;A\cap B=\{j\}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; is successfully shifted to &amp;lt;math&amp;gt;B_{ij}=(B\setminus\{j\})\cup\{i\}&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; fails to shift to &amp;lt;math&amp;gt;A_{ij}=(A\setminus\{j\})\cup\{i\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Since  &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; is successfully shifted to &amp;lt;math&amp;gt;B_{ij}&amp;lt;/math&amp;gt;, we know that it must hold that &amp;lt;math&amp;gt;i\not\in B&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B_{ij}\not\in\mathcal{F}&amp;lt;/math&amp;gt;. And the only two reasons for which &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; may fail to shift to &amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt; are: (1) &amp;lt;math&amp;gt;i\in A&amp;lt;/math&amp;gt; and (2) &amp;lt;math&amp;gt;i\not\in A&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;A_{ij}\in\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Case.1: &amp;lt;math&amp;gt;i\in A&amp;lt;/math&amp;gt;. In this case, since &amp;lt;math&amp;gt;i\in B_{ij}&amp;lt;/math&amp;gt;, we have &amp;lt;math&amp;gt;S_{ij}(A)\cap S_{ij}(B)=A\cap B_{ij}=\{i\}&amp;lt;/math&amp;gt;, i.e. &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; is still intersecting with &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; after shifting.&lt;br /&gt;
&lt;br /&gt;
Case.2: &amp;lt;math&amp;gt;i\not\in A&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;A_{ij}\in\mathcal{F}&amp;lt;/math&amp;gt;. In this case, it is easy to verify that &amp;lt;math&amp;gt;A_{ij}\cap B=(A\cap B)\setminus\{j\}=\emptyset&amp;lt;/math&amp;gt;. Recall that we assume &amp;lt;math&amp;gt;A_{ij}\in\mathcal{F}&amp;lt;/math&amp;gt;. This contradicts to that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting.&lt;br /&gt;
&lt;br /&gt;
In conclusion, in all cases, &amp;lt;math&amp;gt;S_{ij}(\mathcal{F})&amp;lt;/math&amp;gt; remains intersecting.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Repeatedly applying &amp;lt;math&amp;gt;S_{ij}(\mathcal{F})&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;0\le i&amp;lt;j\le n-1&amp;lt;/math&amp;gt;, since we only replace elements by smaller elements, eventually &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; will stop changing, that is, &amp;lt;math&amp;gt;S_{ij}(\mathcal{F})=\mathcal{F}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;0\le i&amp;lt;j\le n-1&amp;lt;/math&amp;gt;. We call such an &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; &#039;&#039;&#039;shifted&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
The idea behind the shifting technique is very natural: by applying shifting, all intersecting families are transformed to some &#039;&#039;special forms&#039;&#039;, and we only need to prove the theorem for these special form of intersecting families.&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Proof of Erdős-Ko-Rado theorem| (The original proof of Erdős-Ko-Rado by shifting)&lt;br /&gt;
By the above lemma, it is sufficient to prove the Erdős-Ko-Rado theorem holds for shifted &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;. We assume that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is shifted.&lt;br /&gt;
&lt;br /&gt;
First, it is trivial to see that the theorem holds for &amp;lt;math&amp;gt;k=1&amp;lt;/math&amp;gt; (no matter whether shifted).&lt;br /&gt;
&lt;br /&gt;
Next, we show that the theorem holds when &amp;lt;math&amp;gt;n=2k&amp;lt;/math&amp;gt;  (no matter whether shifted). For any &amp;lt;math&amp;gt;S\in{X\choose k}&amp;lt;/math&amp;gt;, both &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;X\setminus S&amp;lt;/math&amp;gt; are in &amp;lt;math&amp;gt;{X\choose k}&amp;lt;/math&amp;gt;, but at most one of them can be in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|\le\frac{1}{2}{n\choose k}=\frac{n!}{2k!(n-k)!}=\frac{(n-1)!}{(k-1)!(n-k)!}={n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We then apply the induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. For &amp;lt;math&amp;gt;n&amp;gt; 2k&amp;lt;/math&amp;gt;, the induction hypothesis is stated as:&lt;br /&gt;
* the Erdős-Ko-Rado theorem holds for any smaller &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Define&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{F}_0=\{S\in\mathcal{F}\mid n\not\in S\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}_1=\{S\in\mathcal{F}\mid n\in S\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Clearly, &amp;lt;math&amp;gt;\mathcal{F}_0\subseteq{[n-1]\choose k}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}_0&amp;lt;/math&amp;gt; is intersecting. Due to the induction hypothesis, &amp;lt;math&amp;gt;|\mathcal{F}_0|\le{n-2\choose k-1}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
In order to apply the induction, we let&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{F}_1&#039;=\{S\setminus\{n\}\mid S\in\mathcal{F}_1\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Clearly, &amp;lt;math&amp;gt;\mathcal{F}_1&#039;\subseteq{[n-1]\choose k-1}&amp;lt;/math&amp;gt;. If only it is also intersecting, we can apply the induction hypothesis, and indeed it is. To see this, by contradiction we assume that &amp;lt;math&amp;gt;\mathcal{F}_1&#039;&amp;lt;/math&amp;gt; is not intersecting. Then there must exist &amp;lt;math&amp;gt;A,B\in\mathcal{F}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;A\cap B=\{n\}&amp;lt;/math&amp;gt;, which means that &amp;lt;math&amp;gt;|A\cup B|\le 2k-1&amp;lt;n-1&amp;lt;/math&amp;gt;. Thus, there is some &amp;lt;math&amp;gt;0\le i\le n-1&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;i\not\in A\cup B&amp;lt;/math&amp;gt;. Since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is shifted, &amp;lt;math&amp;gt;A_{in}=A\setminus\{n\}\cup\{i\}\in\mathcal{F}&amp;lt;/math&amp;gt;. On the other hand it can be verified that &amp;lt;math&amp;gt;A_{in}\cap B=\emptyset&amp;lt;/math&amp;gt;, which contradicts that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is intersecting. &lt;br /&gt;
&lt;br /&gt;
Thus, &amp;lt;math&amp;gt;\mathcal{F}_1&#039;\subseteq{[n-1]\choose k-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}_1&#039;&amp;lt;/math&amp;gt; is intersecting. Due to the induction hypothesis, &amp;lt;math&amp;gt;|\mathcal{F}_1&#039;|\le{n-2\choose k-2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Combining these together,&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|=|\mathcal{F}_0|+|\mathcal{F}_1|=|\mathcal{F}_0|+|\mathcal{F}_1&#039;|\le {n-2\choose k-1}+{n-2\choose k-2}={n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Sperner system ==&lt;br /&gt;
A set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; with the relation &amp;lt;math&amp;gt;\subseteq&amp;lt;/math&amp;gt; define a poset. Thus, a &#039;&#039;&#039;chain&#039;&#039;&#039; is a sequence &amp;lt;math&amp;gt;S_1\subseteq S_2\subseteq\cdots\subseteq S_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; is an &#039;&#039;&#039;antichain&#039;&#039;&#039; (also called a &#039;&#039;&#039;Sperner system&#039;&#039;&#039;) if for all &amp;lt;math&amp;gt;S,T\in\mathcal{F}&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;S\neq T&amp;lt;/math&amp;gt;, we have &amp;lt;math&amp;gt;S\not\subseteq T&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-uniform &amp;lt;math&amp;gt;{X\choose k}&amp;lt;/math&amp;gt; is an antichain. Let &amp;lt;math&amp;gt;n=|X|&amp;lt;/math&amp;gt;. The size of &amp;lt;math&amp;gt;{X\choose k}&amp;lt;/math&amp;gt; is maximized when &amp;lt;math&amp;gt;k=\lfloor n/2\rfloor&amp;lt;/math&amp;gt;. We wonder whether this is also the largest possible size of any antichain &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In 1928, Emanuel Sperner proved a theorem saying that it is indeed the largest possible antichain. This result, called Sperner&#039;s theorem today, initiated the studies of extremal set theory.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Sperner 1928)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|\mathcal{F}|\le{n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== First proof (shadows)===&lt;br /&gt;
We first introduce the original proof by Sperner, which uses concepts called &#039;&#039;&#039;shadows&#039;&#039;&#039; and &#039;&#039;&#039;shades&#039;&#039;&#039; of set systems.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;|X|=n\,&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;k&amp;lt;n\,&amp;lt;/math&amp;gt;. &lt;br /&gt;
:The &#039;&#039;&#039;shade&#039;&#039;&#039; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is defined to be&lt;br /&gt;
::&amp;lt;math&amp;gt;\nabla\mathcal{F}=\left\{T\in {X\choose k+1}\,\,\bigg|\,\, \exists S\in\mathcal{F}\mbox{ such that } S\subset T\right\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Thus the shade &amp;lt;math&amp;gt;\nabla\mathcal{F}&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; consists of all subsets of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; which can be obtained by adding an element to a set in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Similarly, the &#039;&#039;&#039;shadow&#039;&#039;&#039; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is defined to be&lt;br /&gt;
::&amp;lt;math&amp;gt;\Delta\mathcal{F}=\left\{T\in {X\choose k-1}\,\,\bigg|\,\, \exists S\in\mathcal{F}\mbox{ such that } T\subset S\right\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Thus the shadow &amp;lt;math&amp;gt;\Delta\mathcal{F}&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; consists of all subsets of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; which can be obtained by removing an element from a set in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Next lemma bounds the effects of shadows and shades on the sizes of set systems.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma (Sperner)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;|X|=n\,&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;|\nabla\mathcal{F}|\ge\frac{n-k}{k+1}|\mathcal{F}| &amp;amp;\text{ if } k&amp;lt;n\\&lt;br /&gt;
&amp;amp;|\Delta\mathcal{F}|\ge\frac{k}{n-k+1}|\mathcal{F}| &amp;amp;\text{ if } k&amp;gt;0.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Proof|&lt;br /&gt;
The lemma is proved by double counting. We prove the inequality of &amp;lt;math&amp;gt;|\nabla\mathcal{F}|&amp;lt;/math&amp;gt;. Assume that &amp;lt;math&amp;gt;0\le k&amp;lt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Define&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{R}=\{(S,T)\mid S\in\mathcal{F}, T\in\nabla\mathcal{F}, S\subset T\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We estimate &amp;lt;math&amp;gt;|\mathcal{R}|&amp;lt;/math&amp;gt; in two ways. &lt;br /&gt;
&lt;br /&gt;
For each &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; different &amp;lt;math&amp;gt;T\in\nabla\mathcal{F}&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;S\subset T&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{R}|=(n-k)|\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
For each &amp;lt;math&amp;gt;T\in\nabla\mathcal{F}&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;k+1&amp;lt;/math&amp;gt; ways to choose an &amp;lt;math&amp;gt;S\subset T&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|S|=k&amp;lt;/math&amp;gt;, some of which may not be in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{R}|\le (k+1)|\nabla\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Altogether, we show that &amp;lt;math&amp;gt;|\nabla\mathcal{F}|\ge\frac{n-k}{k+1}|\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The inequality of &amp;lt;math&amp;gt;|\Delta\mathcal{F}|&amp;lt;/math&amp;gt; can be proved in the same way.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
An immediate corollary of the previous lemma is as follows.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition 1|&lt;br /&gt;
:If &amp;lt;math&amp;gt;k\le \frac{n-1}{2}&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;|\nabla\mathcal{F}|\ge|\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
:If &amp;lt;math&amp;gt;k\ge \frac{n-1}{2}&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;|\Delta\mathcal{F}|\ge|\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The idea of Sperner&#039;s proof is pretty clear: &lt;br /&gt;
* we &amp;quot;push up&amp;quot; all the sets in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;&amp;lt;\frac{n-1}{2}&amp;lt;/math&amp;gt; replacing them by their shades; &lt;br /&gt;
* and also &amp;quot;push down&amp;quot; all the sets in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;\ge\frac{n+1}{2}&amp;lt;/math&amp;gt; replacing them by their shadows. &lt;br /&gt;
Repeat this process we end up with a set system &amp;lt;math&amp;gt;\mathcal{F}\subseteq{X\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;. We need to show that this process does not decrease the size of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition 2|&lt;br /&gt;
:Suppose that &amp;lt;math&amp;gt;\mathcal{F}\subseteq2^X&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;\mathcal{F}_k=\mathcal{F}\cap{X\choose k}&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;k_\min&amp;lt;/math&amp;gt; be the smallest &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|\mathcal{F}_k|&amp;gt;0&amp;lt;/math&amp;gt;, and let&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{F}&#039;=\begin{cases}&lt;br /&gt;
\mathcal{F}\setminus\mathcal{F}_{k_\min}\cup \nabla\mathcal{F}_{k_\min} &amp;amp; \mbox{if }k_\min&amp;lt;\frac{n-1}{2},\\&lt;br /&gt;
\mathcal{F} &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:Similarly, let &amp;lt;math&amp;gt;k_\max&amp;lt;/math&amp;gt; be the largest &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|\mathcal{F}_k|&amp;gt;0&amp;lt;/math&amp;gt;, and let&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\mathcal{F}&#039;&#039;=\begin{cases}&lt;br /&gt;
\mathcal{F}\setminus\mathcal{F}_{k_\max}\cup \Delta\mathcal{F}_{k_\max} &amp;amp; \mbox{if }k_\max\ge\frac{n+1}{2},\\&lt;br /&gt;
\mathcal{F} &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
:If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain, &amp;lt;math&amp;gt;\mathcal{F}&#039;&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}&#039;&#039;&amp;lt;/math&amp;gt; are antichains, and we have &amp;lt;math&amp;gt;|\mathcal{F}&#039;|\ge|\mathcal{F}|&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|\mathcal{F}&#039;&#039;|\ge|\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We show that &amp;lt;math&amp;gt;\mathcal{F}&#039;&amp;lt;/math&amp;gt; is an antichain and &amp;lt;math&amp;gt;|\mathcal{F}&#039;|\ge|\mathcal{F}|&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
First, observe that &amp;lt;math&amp;gt;\nabla\mathcal{F}_k\cap\mathcal{F}=\emptyset&amp;lt;/math&amp;gt;, otherwise &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; cannot be an antichain, and due to Proposition 1, &amp;lt;math&amp;gt;|\nabla\mathcal{F}_k|\ge|\mathcal{F}_k|&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;k\le \frac{n-1}{2}&amp;lt;/math&amp;gt;, so &amp;lt;math&amp;gt;|\mathcal{F}&#039;|=|\mathcal{F}|-|\mathcal{F}_k|+|\nabla\mathcal{F}_k|\ge |\mathcal{F}|&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Now we prove that &amp;lt;math&amp;gt;\mathcal{F}&#039;&amp;lt;/math&amp;gt; is an antichain . By contradiction, assume that there are &amp;lt;math&amp;gt;S, T\in \mathcal{F}&#039;&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;S\subset T&amp;lt;/math&amp;gt;. One of the &amp;lt;math&amp;gt;S,T&amp;lt;/math&amp;gt; must be in &amp;lt;math&amp;gt;\nabla\mathcal{F}_{k_\min}&amp;lt;/math&amp;gt;, or otherwise &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; cannot be an antichain. Recall that &amp;lt;math&amp;gt;k_\min&amp;lt;/math&amp;gt; is the smallest &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|\mathcal{F}_k|&amp;gt;0&amp;lt;/math&amp;gt;, thus it must be &amp;lt;math&amp;gt;S\in \nabla\mathcal{F}_{k_\min}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;T\in\mathcal{F}&amp;lt;/math&amp;gt;. This implies that there is an &amp;lt;math&amp;gt;R\in \mathcal{F}_{k_\min}\subseteq \mathcal{F}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;R\subset S\subset T&amp;lt;/math&amp;gt;, which contradicts that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain.&lt;br /&gt;
&lt;br /&gt;
The statement for &amp;lt;math&amp;gt;\mathcal{F}&#039;&#039;&amp;lt;/math&amp;gt; can be proved in the same way.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Applying the above process, we prove the Sperner&#039;s theorem.&lt;br /&gt;
{{Prooftitle|Proof of Sperner&#039;s theorem | (original proof of Sperner)&lt;br /&gt;
Let &amp;lt;math&amp;gt;\mathcal{F}_k=\{S\in\mathcal{F}\mid |S|=k\}&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;0\le k\le n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We change &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; as follows: &lt;br /&gt;
* for the smallest &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|\mathcal{F}_k|&amp;gt;0&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;k&amp;lt;\frac{n-1}{2}&amp;lt;/math&amp;gt;, replace &amp;lt;math&amp;gt;\mathcal{F}_k&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;\nabla\mathcal{F}_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Due to Proposition 2, this procedure preserves &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; as an antichain and does not decrease &amp;lt;math&amp;gt;|\mathcal{F}|&amp;lt;/math&amp;gt;. Repeat this procedure, until &amp;lt;math&amp;gt;|\mathcal{F}_k|=0&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;k&amp;lt;\frac{n-1}{2}&amp;lt;/math&amp;gt;, that is, there is no member set of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; has size less than &amp;lt;math&amp;gt;\frac{n-1}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We then define another symmetric procedure:&lt;br /&gt;
* for the largest &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|\mathcal{F}_k|&amp;gt;0&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;k\ge\frac{n+1}{2}&amp;lt;/math&amp;gt;, replace &amp;lt;math&amp;gt;\mathcal{F}_k&amp;lt;/math&amp;gt; by &amp;lt;math&amp;gt;\Delta\mathcal{F}_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
Also due to Proposition 2, this procedure preserves &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; as an antichain and does not decrease &amp;lt;math&amp;gt;|\mathcal{F}|&amp;lt;/math&amp;gt;. After repeatedly applying this procedure, &amp;lt;math&amp;gt;|\mathcal{F}_k|=0&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;k\ge\frac{n+1}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
The resulting &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\mathcal{F}\subseteq{X\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;, and since &amp;lt;math&amp;gt;|\mathcal{F}|&amp;lt;/math&amp;gt; is never decreased, for the original &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|\le {n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Second proof (counting)===&lt;br /&gt;
We now introduce an elegant proof due to Lubell. The proof uses a counting argument, and tells more information than just the size of the set system.&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Proof of Sperner&#039;s theorem | (Lubell 1966)&lt;br /&gt;
Let &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; be a permutation of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. We say that an &amp;lt;math&amp;gt;S\subseteq X&amp;lt;/math&amp;gt; &#039;&#039;&#039;prefixes&#039;&#039;&#039; &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;S=\{\pi_1,\pi_2,\ldots, \pi_{|S|}\}&amp;lt;/math&amp;gt;, that is, &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is precisely the set of the first &amp;lt;math&amp;gt;|S|&amp;lt;/math&amp;gt; elements in the permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Fix an &amp;lt;math&amp;gt;S\subseteq X&amp;lt;/math&amp;gt;. It is easy to see that the number of permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; prefixed by &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;|S|!(n-|S|)!&amp;lt;/math&amp;gt;.  Also, since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain, no permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; can be prefixed by more than one members of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;, otherwise one of the member sets must contain the other, which contradicts that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain. Thus, the number of permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; prefixed by some &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}|S|!(n-|S|)!&amp;lt;/math&amp;gt;,&lt;br /&gt;
which cannot be larger than the total number of permutations, &amp;lt;math&amp;gt;n!&amp;lt;/math&amp;gt;, therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}|S|!(n-|S|)!\le n!&amp;lt;/math&amp;gt;.&lt;br /&gt;
Dividing both sides by &amp;lt;math&amp;gt;n!&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}=\sum_{S\in\mathcal{F}}\frac{|S|!(n-|S|)!}{n!}\le 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{n\choose |S|}\le {n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;, so&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}\ge \frac{|\mathcal{F}|}{{n\choose \lfloor n/2\rfloor}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Combining this with the above inequality, we prove the Sperner&#039;s theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== The LYM inequality ===&lt;br /&gt;
Lubell&#039;s proof proves the following inequality:&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}\le 1&amp;lt;/math&amp;gt;&lt;br /&gt;
which is actually stronger than Sperner&#039;s original statement that &amp;lt;math&amp;gt;|\mathcal{F}|\le{n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
This inequality is independently discovered by Lubell-Yamamoto, Meschalkin, and Bollobás, and is called the LYM inequality today.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Lubell, Yamamoto 1954; Meschalkin 1963)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain, then&lt;br /&gt;
::&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}\le 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
In Lubell&#039;s counting argument proves the LYM inequality, which implies the Sperner&#039;s theorem. Here we give another proof of the LYM inequality by the probabilistic method,  due to Noga Alon.&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof (the probabilistic method)| (Due to Alon.)&lt;br /&gt;
Let &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; be a uniformly random permutation of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. Define a random maximal chain by&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{C}_\pi=\{\{\pi_i\mid 1\le i\le k\}\mid 0\le k\le n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_S&amp;lt;/math&amp;gt; be the 0-1 random variable which indicates whether &amp;lt;math&amp;gt;S\in\mathcal{C}_\pi&amp;lt;/math&amp;gt;, that is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X_S=\begin{cases}&lt;br /&gt;
1 &amp;amp; \mbox{if }S\in\mathcal{C}_\pi,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that for a uniformly random &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathcal{C}_\pi&amp;lt;/math&amp;gt; has exact one member set of size &amp;lt;math&amp;gt;|S|&amp;lt;/math&amp;gt;, uniformly distributed over &amp;lt;math&amp;gt;{X\choose |S|}&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_S]=\Pr[S\in\mathcal{C}_\pi]=\frac{1}{{n\choose |S|}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let &amp;lt;math&amp;gt;X=\sum_{S\in\mathcal{F}}X_S&amp;lt;/math&amp;gt;. Note that &amp;lt;math&amp;gt;X=|\mathcal{F}\cap\mathcal{C}_\pi|&amp;lt;/math&amp;gt;. By the linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X]=\sum_{S\in\mathcal{F}}\mathbf{E}[X_S]=\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
On the other hand, since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is an antichain, it can never intersect a chain at more than one elements, thus we always have &amp;lt;math&amp;gt;X=|\mathcal{F}\cap\mathcal{C}_\pi|\le 1&amp;lt;/math&amp;gt;. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}\le \mathbf{E}[X] \le 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The Sperner&#039;s theorem is an immediate consequence of the LYM inequality.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}\le 1&amp;lt;/math&amp;gt; implies that &amp;lt;math&amp;gt;|\mathcal{F}|\le{n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
It holds that &amp;lt;math&amp;gt;{n\choose k}\le {n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;1\ge \sum_{S\in\mathcal{F}}\frac{1}{{n\choose |S|}}\ge \frac{|\mathcal{F}|}{{n\choose \lfloor n/2\rfloor}}&amp;lt;/math&amp;gt;,&lt;br /&gt;
which implies that &amp;lt;math&amp;gt;|\mathcal{F}|\le {n\choose \lfloor n/2\rfloor}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Sauer&#039;s lemma and VC-dimension ==&lt;br /&gt;
&lt;br /&gt;
=== Shattering and the VC-dimension ===&lt;br /&gt;
{{Theorem|Definition (shatter)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; be set family and let &amp;lt;math&amp;gt;R\subseteq X&amp;lt;/math&amp;gt; be a subset. The &#039;&#039;&#039;trace&#039;&#039;&#039; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; on &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;, denoted &amp;lt;math&amp;gt;\mathcal{F}|_R&amp;lt;/math&amp;gt; is defined as&lt;br /&gt;
::&amp;lt;math&amp;gt;\mathcal{F}|_R=\{S\cap R\mid S\in\mathcal{F}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:We say that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; &#039;&#039;&#039;shatters&#039;&#039;&#039; &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;\mathcal{F}|_R=2^R&amp;lt;/math&amp;gt;, i.e. for all &amp;lt;math&amp;gt;T\subseteq R&amp;lt;/math&amp;gt;, there exists an &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;T=S\cap R&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The [http://en.wikipedia.org/wiki/VC_dimension &#039;&#039;&#039;VC dimension&#039;&#039;&#039;] is defined by the power of a family to shatter a set.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition (VC-dimension)|&lt;br /&gt;
:The &#039;&#039;&#039;Vapnik–Chervonenkis dimension&#039;&#039;&#039; (&#039;&#039;&#039;VC-dimension&#039;&#039;&#039;) of a set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt;, denoted &amp;lt;math&amp;gt;\text{VC-dim}(\mathcal{F})&amp;lt;/math&amp;gt;, is the size of the largest &amp;lt;math&amp;gt;R\subseteq X&amp;lt;/math&amp;gt; shattered by &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is a core concept in [http://en.wikipedia.org/wiki/Computational_learning_theory computational learning theory].&lt;br /&gt;
&lt;br /&gt;
Each subset &amp;lt;math&amp;gt;S\subseteq X&amp;lt;/math&amp;gt; can be equivalently represented by its characteristic function &amp;lt;math&amp;gt;f_S:X\rightarrow\{0,1\}&amp;lt;/math&amp;gt;, such that for each &amp;lt;math&amp;gt;x\in X&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;f_S(x)=\begin{cases}&lt;br /&gt;
1 &amp;amp; x\in S\\&lt;br /&gt;
0 &amp;amp; x\not\in S.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Then a set family &amp;lt;math&amp;gt;\mathcal{F}\subseteq2^X&amp;lt;/math&amp;gt; corresponds to a collection of boolean functions &amp;lt;math&amp;gt;\{f_S\mid S\in\mathcal{F}\}&amp;lt;/math&amp;gt;, which is a subset of all Boolean functions in the form &amp;lt;math&amp;gt;f:X\rightarrow\{0,1\}&amp;lt;/math&amp;gt;. We wonder on how large a subdomain &amp;lt;math&amp;gt;Y\subseteq X&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; includes all the &amp;lt;math&amp;gt;2^{|Y|}&amp;lt;/math&amp;gt; mappings &amp;lt;math&amp;gt;Y\rightarrow\{0,1\}&amp;lt;/math&amp;gt;. The largest size of such subdomain is the VC-dimension. It measures how complicated a collection of boolean functions (or equivalently a set family) is.&lt;br /&gt;
&lt;br /&gt;
=== Sauer&#039;s Lemma ===&lt;br /&gt;
The definition of the VC-dimension involves enumerating all subsets, thus is difficult to analyze in general. The following famous result state a very simple sufficient condition to lower bound the VC-dimension, regarding only the size of the family. The lemma is due to Sauer, and independently due to Shelah and Perles. A slightly weaker version is found by Vapnik and Chervonenkis, who use the framework to develop a theory of classifications.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Sauer&#039;s Lemma (Sauer; Shelah-Perles; Vapnik-Chervonenkis)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;|\mathcal{F}|&amp;gt;\sum_{1\le i&amp;lt;k}{n\choose i}&amp;lt;/math&amp;gt;, then there exists an &amp;lt;math&amp;gt;R\in{X\choose k}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; shatters &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
In other words, for any set family &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|\mathcal{F}|&amp;gt;\sum_{1\le i&amp;lt;k}{n\choose i}&amp;lt;/math&amp;gt;, its VC-dimension &amp;lt;math&amp;gt;\text{VC-dim}(\mathcal{F})\ge k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Hereditary family ===&lt;br /&gt;
We note the Sauer&#039;s lemma is especially easy to prove for a special type of set families, called the &#039;&#039;&#039;hereditary&#039;&#039;&#039; families.&lt;br /&gt;
{{Theorem|Definition (hereditary family)|&lt;br /&gt;
:A set system &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; is said to be &#039;&#039;&#039;hereditary&#039;&#039;&#039; (also called an &#039;&#039;&#039;ideal&#039;&#039;&#039; or an &#039;&#039;&#039;abstract simplicial complex&#039;&#039;&#039;), if&lt;br /&gt;
::&amp;lt;math&amp;gt;S\subseteq T\in\mathcal{F}&amp;lt;/math&amp;gt; implies &amp;lt;math&amp;gt;S\in\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
In other words, for a hereditary family &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;R\in\mathcal{F}&amp;lt;/math&amp;gt;, then all subsets of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; are also in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;. An immediate consequence is the following proposition.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; be a hereditary family. If &amp;lt;math&amp;gt;R\in\mathcal{F}&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; shatters &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Therefore, it is very easy to prove the Sauer&#039;s lemma for hereditary families: &lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:For &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is hereditary and &amp;lt;math&amp;gt;|\mathcal{F}|&amp;gt;\sum_{1\le i&amp;lt;k}{n\choose i}&amp;lt;/math&amp;gt; then there exists an &amp;lt;math&amp;gt;R\in{X\choose k}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; shatters &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is hereditary, we only need to show that there exists an &amp;lt;math&amp;gt;R\in\mathcal{F}&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;|R|\ge k&amp;lt;/math&amp;gt;, which must be true, because if all members of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; are of sizes &amp;lt;math&amp;gt;&amp;lt;k&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;|\mathcal{F}|\le\left|\bigcup_{1\le i&amp;lt;k}{X\choose k}\right|=\sum_{1\le i&amp;lt;k}{n\choose i}&amp;lt;/math&amp;gt;, a contradiction.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
To prove the Sauer&#039;s lemma for general non-hereditary families, we can use some way to reduce arbitrary families to hereditary families. Here we apply the shifting technique to achieve this.&lt;br /&gt;
&lt;br /&gt;
=== Down-shifts ===&lt;br /&gt;
Note that we work on &amp;lt;math&amp;gt;\mathcal{F}\subseteq2^X&amp;lt;/math&amp;gt;, instead of &amp;lt;math&amp;gt;\mathcal{F}\subseteq{X\choose k}&amp;lt;/math&amp;gt; like in the Erdős–Ko–Rado theorem, so we do not need to preserve the size of member sets. Instead, we need to reduce an arbitrary family to a hereditary one, thus we use a shift operator which replaces a member set by a subset of it.&lt;br /&gt;
{{Theorem|Definition (down-shifts)|&lt;br /&gt;
: Assume &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^{[n]}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;. Define the &#039;&#039;&#039;down-shift&#039;&#039;&#039; operator &amp;lt;math&amp;gt;S_{i}&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
:* for each &amp;lt;math&amp;gt;T\in\mathcal{F}&amp;lt;/math&amp;gt;, let&lt;br /&gt;
::&amp;lt;math&amp;gt;S_{i}(T)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
T\setminus\{i\} &amp;amp; \mbox{if }i\in T \mbox{ and }T\setminus\{i\} \not\in\mathcal{F},\\&lt;br /&gt;
T &amp;amp; \mbox{otherwise;}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
:* let &amp;lt;math&amp;gt;S_{i}(\mathcal{F})=\{S_{i}(T)\mid T\in \mathcal{F}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Repeatedly applying &amp;lt;math&amp;gt;S_i&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;, due to the finiteness, eventually &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is not changed by any &amp;lt;math&amp;gt;S_i&amp;lt;/math&amp;gt;. We call such a family &#039;&#039;&#039;down-shifted&#039;&#039;&#039;. A family &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is down-shifted if and only if &amp;lt;math&amp;gt;S_i(\mathcal{F})=\mathcal{F}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;. It is then easy to see that a down-shifted &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; must be hereditary.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:If &amp;lt;math&amp;gt;\mathcal{F}\subseteq2^X&amp;lt;/math&amp;gt; is down-shifted, then &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is hereditary.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
In order to use down-shift to prove the Sauer&#039;s lemma, we need to make sure that down-shift does not decrease &amp;lt;math&amp;gt;|\mathcal{F}|&amp;lt;/math&amp;gt; and does not increase the VC-dimension &amp;lt;math&amp;gt;\text{VC-dim}(\mathcal{F})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
# &amp;lt;math&amp;gt;|S_{i}(\mathcal{F})|=\mathcal{F}&amp;lt;/math&amp;gt;;&lt;br /&gt;
# &amp;lt;math&amp;gt;|S_i(\mathcal{F})|_R|\le |\mathcal{F}|_R|&amp;lt;/math&amp;gt;, thus if &amp;lt;math&amp;gt;S_{i}(\mathcal{F})&amp;lt;/math&amp;gt; shatters an &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;, so does &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
(1) is immediate. (2) is proved by case analysis. We omit the proof.&lt;br /&gt;
&lt;br /&gt;
;Proof of Sauer&#039;s lemma&lt;br /&gt;
Now we can prove the Sauer&#039;s lemma for arbitrary &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^{[n]}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^{[n]}&amp;lt;/math&amp;gt;, repeatedly apply &amp;lt;math&amp;gt;S_i&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i\in&amp;lt;/math&amp;gt; till the family is down-shifted, which is denoted by &amp;lt;math&amp;gt;\mathcal{F}&#039;&amp;lt;/math&amp;gt;. We have proved that &amp;lt;math&amp;gt;|\mathcal{F}&#039;|=|\mathcal{F}|&amp;gt;\sum_{1\le i&amp;lt;k}{n\choose i}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{F}&#039;&amp;lt;/math&amp;gt; is hereditary, thus as argued before, here exists an &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; shattered by &amp;lt;math&amp;gt;\mathcal{F}&#039;&amp;lt;/math&amp;gt;. By the above proposition, &amp;lt;math&amp;gt;|\mathcal{F}&#039;|_R|\le |\mathcal{F}|_R|&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; also shatters &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;. The lemma is proved.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== The Kruskal–Katona Theorem ==&lt;br /&gt;
The &#039;&#039;&#039;shadow&#039;&#039;&#039; of a set system &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;, denoted &amp;lt;math&amp;gt;\Delta\mathcal{F}&amp;lt;/math&amp;gt;, consists of all sets  which can be obtained by removing an element from a set in &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt;. The &#039;&#039;&#039;shadow&#039;&#039;&#039; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is defined to be&lt;br /&gt;
::&amp;lt;math&amp;gt;\Delta\mathcal{F}=\left\{T\in {X\choose k-1}\,\,\bigg|\,\, \exists S\in\mathcal{F}\mbox{ such that } T\subset S\right\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The shadow contains rich information about the set system. An extremal question is: for a system of fixed number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-sets, how small can its shadow be? The Kruskal–Katona theorem gives an answer to this question.&lt;br /&gt;
&lt;br /&gt;
To state the result of the Kruskal–Katona theorem, we need to introduce the concepts of the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation of numbers and the colex order of sets.&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation of a number ===&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:Given positive integers &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;, there exists a unique representation of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; in the form&lt;br /&gt;
::&amp;lt;math&amp;gt;m={m_k\choose k}+{m_{k-1}\choose k-1}+\cdots+{m_t\choose t}=\sum_{\ell=t}^k{m_\ell\choose \ell}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;m_k&amp;gt;m_{k-1}&amp;gt;\cdots&amp;gt;m_t\ge t\ge 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
:This representation of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; is called a &#039;&#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade&#039;&#039;&#039; (or a &#039;&#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-binomial&#039;&#039;&#039;) representation of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;. &lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
In fact, the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation &amp;lt;math&amp;gt;(m_k,m_{k-1},\ldots,m_t)&amp;lt;/math&amp;gt; of an &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; can be found by the following simple greedy algorithm:&lt;br /&gt;
----&lt;br /&gt;
:&amp;lt;math&amp;gt;\ell=k;&amp;lt;/math&amp;gt;&lt;br /&gt;
:while (&amp;lt;math&amp;gt;m&amp;gt;0&amp;lt;/math&amp;gt;) do&lt;br /&gt;
::let &amp;lt;math&amp;gt;m_\ell&amp;lt;/math&amp;gt; be the largest integer for which &amp;lt;math&amp;gt;{m_\ell\choose \ell}\le m;&amp;lt;/math&amp;gt;&lt;br /&gt;
::&amp;lt;math&amp;gt;m=m-m_\ell;&amp;lt;/math&amp;gt;&lt;br /&gt;
::&amp;lt;math&amp;gt;\ell=\ell-1;&amp;lt;/math&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
We then show that the above algorithm constructs a sequence &amp;lt;math&amp;gt;(m_k,m_{k-1},\ldots,m_t)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;m_k&amp;gt;m_{k-1}&amp;gt;\cdots&amp;gt;m_t\ge t\ge 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose the current &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt; and the current &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;. To see that &amp;lt;math&amp;gt;m_{\ell-1}&amp;lt; m_\ell&amp;lt;/math&amp;gt;, we suppose otherwise &amp;lt;math&amp;gt;m_{\ell-1}\ge m_\ell&amp;lt;/math&amp;gt;. Then &lt;br /&gt;
:&amp;lt;math&amp;gt;m\ge {m_\ell\choose \ell}+{m_{\ell-1}\choose \ell-1}\ge{m_\ell\choose \ell}+{m_{\ell}\choose \ell-1}={1+m_{\ell}\choose \ell}&amp;lt;/math&amp;gt;&lt;br /&gt;
contradicting the maximality of &amp;lt;math&amp;gt;m_\ell&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;m_\ell&amp;gt;m_{\ell-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The algorithm continues reducing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; to smaller nonnegative values, and eventually reaches a stage where the choice of &amp;lt;math&amp;gt;m_t&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;t\ge 2&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;{m_t\choose t}&amp;lt;/math&amp;gt; equals the current &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;; or gets right down to choosing &amp;lt;math&amp;gt;m_1&amp;lt;/math&amp;gt; as the integer such that &amp;lt;math&amp;gt;m_1={m_1\choose 1}&amp;lt;/math&amp;gt; equals the current &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore,  &amp;lt;math&amp;gt;m={m_k\choose k}+{m_{k-1}\choose k-1}+\cdots+{m_t\choose t}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;m_k&amp;gt;m_{k-1}&amp;gt;\cdots&amp;gt;m_t\ge t\ge 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The uniqueness of the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation follows by the induction on &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
When &amp;lt;math&amp;gt;k=1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; has a unique &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation &amp;lt;math&amp;gt;m={m\choose 1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For general &amp;lt;math&amp;gt;k&amp;gt;1&amp;lt;/math&amp;gt;, suppose that every nonnegative integer &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; has a unique &amp;lt;math&amp;gt;(k-1)&amp;lt;/math&amp;gt;-cascade representation.&lt;br /&gt;
&lt;br /&gt;
Suppose then that &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; has two &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representations:&lt;br /&gt;
:&amp;lt;math&amp;gt;m={a_k\choose k}+{a_{k-1}\choose k-1}+\cdots+{a_t\choose t}={b_k\choose k}+{b_{k-1}\choose k-1}+\cdots+{b_r\choose r}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We then show that it must hold that &amp;lt;math&amp;gt;a_k=b_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
If &amp;lt;math&amp;gt;a_k\neq b_k&amp;lt;/math&amp;gt;, WLOG, suppose that &amp;lt;math&amp;gt;a_k&amp;lt;b_k&amp;lt;/math&amp;gt;. We obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
m&amp;amp;&lt;br /&gt;
={a_k\choose k}+{a_{k-1}\choose k-1}+\cdots+{a_t\choose t}\\&lt;br /&gt;
&amp;amp;\le {a_k\choose k}+{a_{k}-1\choose k-1}+\cdots+{a_k-(k-t)\choose t}+\cdots+{a_k-k+1\choose 1}\\&lt;br /&gt;
&amp;amp;={a_k+1\choose k}-1,&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is got by repeatedly applying the identity&lt;br /&gt;
:&amp;lt;math&amp;gt;{n\choose k}+{n\choose k-1}={n+1\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We then obtain &lt;br /&gt;
:&amp;lt;math&amp;gt;m&amp;lt;{a_k+1\choose k}\le{b_k\choose k}\le m&amp;lt;/math&amp;gt;, &lt;br /&gt;
which is a contradiction. Therefore, &amp;lt;math&amp;gt;a_k=b_k&amp;lt;/math&amp;gt;, and by the induction hypothesis, the remaining value &amp;lt;math&amp;gt;m-a_k=m-b_k&amp;lt;/math&amp;gt; has a unique &amp;lt;math&amp;gt;(k-1)&amp;lt;/math&amp;gt;-cascade representation, so &amp;lt;math&amp;gt;a_i=b_i&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Co-lexicographic order of subsets ===&lt;br /&gt;
The co-lexicographic order of sets plays a particularly important role in the investigation of the size of the shadow of a system of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-sets.&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
: The &#039;&#039;&#039;co-lexicographic (colex) order&#039;&#039;&#039; (also called the &#039;&#039;&#039;reverse lexicographic order&#039;&#039;&#039;) of sets is defined as follows: for any &amp;lt;math&amp;gt;A,B\subseteq [n]&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;A\neq B&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;A&amp;lt;B&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;\max A\setminus B &amp;lt; \max B\setminus A&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We can sort sets in colex order by first writing each set as a tuple, whose elements are in decreasing order, and then sorting the tuples in the lexicographic order of tuples.&lt;br /&gt;
&lt;br /&gt;
For example, &amp;lt;math&amp;gt;{[5]\choose 3}&amp;lt;/math&amp;gt; in colex order is&lt;br /&gt;
 {3,2,1}&lt;br /&gt;
 {4,2,1}&lt;br /&gt;
 {4,3,1}&lt;br /&gt;
 {4,3,2}&lt;br /&gt;
 {5,2,1}&lt;br /&gt;
 {5,3,1}&lt;br /&gt;
 {5,3,2}&lt;br /&gt;
 {5,4,1}&lt;br /&gt;
 {5,4,2}&lt;br /&gt;
 {5,4,3}&lt;br /&gt;
&lt;br /&gt;
We find that the first &amp;lt;math&amp;gt;{n\choose 3}&amp;lt;/math&amp;gt; sets in this order for &amp;lt;math&amp;gt;n=3,4,5&amp;lt;/math&amp;gt;, form precisely &amp;lt;math&amp;gt;{[n]\choose 3}&amp;lt;/math&amp;gt;. And if we write the colex order of &amp;lt;math&amp;gt;{[6]\choose 3}&amp;lt;/math&amp;gt;, the above colex order of &amp;lt;math&amp;gt;{[5]\choose 3}&amp;lt;/math&amp;gt; appears as a prefix of that order. Elaborating on this, we have:&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{R}(m,k)&amp;lt;/math&amp;gt; be the first &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; sets in the colex order of &amp;lt;math&amp;gt;{\mathbb{N}\choose k}&amp;lt;/math&amp;gt;. Then &lt;br /&gt;
::&amp;lt;math&amp;gt;\mathcal{R}\left({n\choose k},k\right)={[n]\choose k}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:that is, the first &amp;lt;math&amp;gt;{n\choose k}&amp;lt;/math&amp;gt; sets in the colex order of all &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-sets of natural numbers is precisely &amp;lt;math&amp;gt;{[n]\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This proposition says that the sets in &amp;lt;math&amp;gt;\mathcal{R}(m,k)&amp;lt;/math&amp;gt; is highly overlapped, which suggests that &amp;lt;math&amp;gt;\mathcal{R}(m,k)&amp;lt;/math&amp;gt; may have small shadow. The size of the shadow of &amp;lt;math&amp;gt;\mathcal{R}(m,k)&amp;lt;/math&amp;gt; is closely related to the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:Suppose the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; is &lt;br /&gt;
::&amp;lt;math&amp;gt;m={m_k\choose k}+{m_{k-1}\choose k-1}+\cdots+{m_t\choose t}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;|\Delta\mathcal{R}(m,k)|={m_k\choose k-1}+{m_{k-1}\choose k-2}+\cdots+{m_t\choose t-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Given &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; and its &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-cascade representation &amp;lt;math&amp;gt;m={m_k\choose k}+{m_{k-1}\choose k-1}+\cdots+{m_t\choose t}&amp;lt;/math&amp;gt;, the &amp;lt;math&amp;gt;\mathcal{R}(m,k)&amp;lt;/math&amp;gt; is constructed as:&lt;br /&gt;
* all sets in &amp;lt;math&amp;gt;{[m_k]\choose k}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* all sets in &amp;lt;math&amp;gt;{[m_{k-1}]\choose k-1}&amp;lt;/math&amp;gt;, unioned with &amp;lt;math&amp;gt;\{1+m_k\}\,&amp;lt;/math&amp;gt;;&lt;br /&gt;
::&amp;lt;math&amp;gt;\vdots&amp;lt;/math&amp;gt;&lt;br /&gt;
* all sets in &amp;lt;math&amp;gt;{[m_{t}]\choose t}&amp;lt;/math&amp;gt;, unioned with &amp;lt;math&amp;gt;\{1+m_k,1+m_{k-1},\ldots,1+m_{t+1}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The shadow &amp;lt;math&amp;gt;\Delta\mathcal{R}(m,k)&amp;lt;/math&amp;gt; is the collection of all &amp;lt;math&amp;gt;(k-1)&amp;lt;/math&amp;gt;-sets contained by the above sets, which are&lt;br /&gt;
* all sets in &amp;lt;math&amp;gt;{[m_k]\choose k-1}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* all sets in &amp;lt;math&amp;gt;{[m_{k-1}]\choose k-2}&amp;lt;/math&amp;gt;, unioned with &amp;lt;math&amp;gt;\{1+m_k\}\,&amp;lt;/math&amp;gt;;&lt;br /&gt;
::&amp;lt;math&amp;gt;\vdots&amp;lt;/math&amp;gt;&lt;br /&gt;
* all sets in &amp;lt;math&amp;gt;{[m_{t}]\choose t-1}&amp;lt;/math&amp;gt;, unioned with &amp;lt;math&amp;gt;\{1+m_k,1+m_{k-1},\ldots,1+m_{t+1}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta\mathcal{R}(m,k)|={m_k\choose k-1}+{m_{k-1}\choose k-2}+\cdots+{m_t\choose t-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== The Kruskal–Katona theorem ===&lt;br /&gt;
The Kruskal–Katona theorem states that among all systems of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-sets, &amp;lt;math&amp;gt;\mathcal{R}(m,k)&amp;lt;/math&amp;gt;, i.e., the first &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-sets in the colex order, has the smallest shadow.&lt;br /&gt;
&lt;br /&gt;
The theorem is proved independently by Joseph Kruskal in 1963 and G.O.H. Katona in 1966, and is a fundamental result in finite set theory and combinatorial topology.&lt;br /&gt;
{{Theorem|Theorem (Kruskal 1963, Katona 1966)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|\mathcal{F}|=m&amp;lt;/math&amp;gt;, and suppose that&lt;br /&gt;
::&amp;lt;math&amp;gt;m={m_k\choose k}+{m_{k-1}\choose k-1}+\cdots+{m_t\choose t}&amp;lt;/math&amp;gt;&lt;br /&gt;
:where &amp;lt;math&amp;gt;m_k&amp;gt;m_{k-1}&amp;gt;\cdots&amp;gt;m_t\ge t\ge 1&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;|\Delta\mathcal{F}|\ge {m_k\choose k-1}+{m_{k-1}\choose k-2}+\cdots+{m_t\choose t-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;-vector of a set system &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^X&amp;lt;/math&amp;gt; is a vector &amp;lt;math&amp;gt;(|\mathcal{F}_0|,|\mathcal{F}_1|,\ldots,|\mathcal{F}_n|)&amp;lt;/math&amp;gt; where&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{F}_k=\{S\mid S\in\mathcal{F}, |S|=k\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
i.e., the &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;-vector gives the number of member sets of each size.&lt;br /&gt;
&lt;br /&gt;
In a hereditary family (also called an [http://en.wikipedia.org/wiki/Abstract_simplicial_complex abstract simplicial complex]), &amp;lt;math&amp;gt;\mathcal{F}_k&amp;lt;/math&amp;gt; is formed by the shadow of &amp;lt;math&amp;gt;\mathcal{F}_{k+1}&amp;lt;/math&amp;gt; as well as some additional &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-sets introduced in this level. The Kruskal-Katona theorem gives a lower bound on &amp;lt;math&amp;gt;|\mathcal{F}_i|&amp;lt;/math&amp;gt;, given the &amp;lt;math&amp;gt;|\mathcal{F}_{i-1}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The original proof of the theorem is rather complicated. In the following years, several different proofs were discovered. Here we present a proof dueto Frankl by the shifting technique.&lt;br /&gt;
&lt;br /&gt;
;Frankl&#039;s shifting proof of Kruskal-Katonal theorem (Frankl 1984)&lt;br /&gt;
&lt;br /&gt;
We take the classic &amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shift operator &amp;lt;math&amp;gt;S_{ij}&amp;lt;/math&amp;gt; defined in the original proof of the Erdős-Ko-Rado theorem.&lt;br /&gt;
{{Theorem|Definition (&amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shift)|&lt;br /&gt;
: Assume &amp;lt;math&amp;gt;\mathcal{F}\subseteq 2^{[n]}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;0\le i&amp;lt;j\le n-1&amp;lt;/math&amp;gt;. Define the &#039;&#039;&#039;&amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shift&#039;&#039;&#039; &amp;lt;math&amp;gt;S_{ij}&amp;lt;/math&amp;gt; as an operator on &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
:*for each &amp;lt;math&amp;gt;T\in\mathcal{F}&amp;lt;/math&amp;gt;, write &amp;lt;math&amp;gt;T_{ij}=(T\setminus\{j\})\cup\{i\} &amp;lt;/math&amp;gt;, and let&lt;br /&gt;
::&amp;lt;math&amp;gt;S_{ij}(T)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
T_{ij} &amp;amp; \mbox{if }j\in T, i\not\in T, \mbox{ and }T_{ij} \not\in\mathcal{F},\\&lt;br /&gt;
T &amp;amp; \mbox{otherwise;}&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
:* let &amp;lt;math&amp;gt;S_{ij}(\mathcal{F})=\{S_{ij}(T)\mid T\in \mathcal{F}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is immediate that shifting does not change the size of the set or the size of the system, i.e., &amp;lt;math&amp;gt;|S_{ij}(T)|=|T|\,&amp;lt;/math&amp;gt; and  &amp;lt;math&amp;gt;|S_{ij}(\mathcal{F})|=\mathcal{F}&amp;lt;/math&amp;gt;. And due to the finiteness, repeatedly applying &amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shifts for any &amp;lt;math&amp;gt;1\le i&amp;lt;j\le n&amp;lt;/math&amp;gt;, eventually &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; does not changing any more. We called such an &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; with  &amp;lt;math&amp;gt;\mathcal{F}=S_{ij}(\mathcal{F})&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;1\le i&amp;lt;j\le n&amp;lt;/math&amp;gt; &#039;&#039;&#039;shifted&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
In order to make the shifting technique work for shadows, we have to prove that shifting does not increase the size of the shadow.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:&amp;lt;math&amp;gt;\Delta S_{ij}(\mathcal{F})\subseteq S_{ij}(\Delta\mathcal{F})&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We abuse the notation &amp;lt;math&amp;gt;\Delta&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\Delta A=\Delta\{A\}&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is a set instead of a set system.&lt;br /&gt;
&lt;br /&gt;
It is sufficient to show that for any &amp;lt;math&amp;gt;A\in\mathcal{F}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\Delta S_{ij}(A)\subseteq S_{ij}(\Delta\mathcal{F})&amp;lt;/math&amp;gt;, which can be proved by case analysis. We omit the proof.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
An immediate corollary of the above proposition is that the &amp;lt;math&amp;gt;(i,j)&amp;lt;/math&amp;gt;-shifts &amp;lt;math&amp;gt;S_{ij}&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;1\le i&amp;lt;j\le n&amp;lt;/math&amp;gt; do not increase the size of the shadow.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta S_{ij}(\mathcal{F})|\le |\Delta\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
By the above proposition, &amp;lt;math&amp;gt;|\Delta S_{ij}(\mathcal{F})|\le|S_{ij}(\Delta\mathcal{F})|&amp;lt;/math&amp;gt;, and we know that &amp;lt;math&amp;gt;S_{ij}&amp;lt;/math&amp;gt; does not change the cardinality of a set family, that is, &amp;lt;math&amp;gt;|S_{ij}(\Delta\mathcal{F})|=|\Delta\mathcal{F}|&amp;lt;/math&amp;gt;, therefore &amp;lt;math&amp;gt;\Delta S_{ij}(\mathcal{F})|\le|\Delta\mathcal{F}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;Proof of Kruskal-Katona theorem&lt;br /&gt;
We know that shifts never enlarge the shadow, thus it is sufficient to prove the theorem for shifted &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;. We then assume &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is shifted.&lt;br /&gt;
&lt;br /&gt;
Apply induction on &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; and for given &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; on &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. The theorem holds trivially for the case that &amp;lt;math&amp;gt;k=1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; is arbitrary.&lt;br /&gt;
&lt;br /&gt;
Define &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{F}_0=\{A\in\mathcal{F}\mid 1\not\in A\}&amp;lt;/math&amp;gt;, &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{F}_1=\{A\in\mathcal{F}\mid 1\in A\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
And let &amp;lt;math&amp;gt;\mathcal{F}_1&#039;=\{A\setminus\{1\}\mid A\in\mathcal{F}_1\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Clearly &amp;lt;math&amp;gt;\mathcal{F}_0,\mathcal{F}_1\subseteq{X\choose k}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\mathcal{F}_1&#039;\subseteq{X\choose k-1}&amp;lt;/math&amp;gt;, and&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|=|\mathcal{F}_0|+|\mathcal{F}_1|=|\mathcal{F}_0|+|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Our induction is based on the following observation regarding the size of the shadow.&lt;br /&gt;
{{Theorem|Lemma 1|&lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta\mathcal{F}|\ge|\Delta\mathcal{F}_1&#039;|+|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Obviously &amp;lt;math&amp;gt;\mathcal{F}\supseteq\mathcal{F}_1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\Delta\mathcal{F}\supseteq\Delta\mathcal{F}_1&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Delta\mathcal{F}_1&lt;br /&gt;
&amp;amp;=\left\{A\in{X\choose k-1}\,\,\bigg|\,\, \exists B\in\mathcal{F}_1, A\subset B\right\}\\&lt;br /&gt;
&amp;amp;=\left\{A\in{X\choose k-1}\,\,\bigg|\,\, 1\in A, \exists B\in\mathcal{F}_1, A\subset B\right\}\\&lt;br /&gt;
&amp;amp;\quad\, \cup \left\{A\in{X\choose k-1}\,\,\bigg|\,\, 1\not\in A, \exists B\in\mathcal{F}_1, A\subset B\right\}\\&lt;br /&gt;
&amp;amp;=\{S\cup\{1\}\mid S\in\Delta\mathcal{F}_1&#039;\}\cup\mathcal{F}_1&#039;.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
The union is taken over two disjoint families. Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta\mathcal{F}|\ge|\Delta\mathcal{F}_1|=|\Delta\mathcal{F}_1&#039;|+|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The following property of shifted &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is essential for our proof.&lt;br /&gt;
{{Theorem|Lemma 2|&lt;br /&gt;
:For shifted &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;\Delta\mathcal{F}_0\subseteq \mathcal{F}_1&#039;&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
If &amp;lt;math&amp;gt;A\in\Delta\mathcal{F}_0&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;A\cup\{j\}\in\mathcal{F}_0\subseteq\mathcal{F}&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;j&amp;gt;1&amp;lt;/math&amp;gt; so that, since &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is shifted, applying the &amp;lt;math&amp;gt;(1,j)&amp;lt;/math&amp;gt;-shift &amp;lt;math&amp;gt;S_{1j}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;A\cup\{1\}\in\mathcal{F}&amp;lt;/math&amp;gt;, thus, &amp;lt;math&amp;gt;A\in\mathcal{F}_1&#039;&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We then bound the size of &amp;lt;math&amp;gt;\mathcal{F}_1&#039;&amp;lt;/math&amp;gt; as:&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma 3|&lt;br /&gt;
:If &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is shifted, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|\mathcal{F}_1&#039;|\ge{m_k-1\choose k-1}+{m_{k-1}-1\choose k-2}+\cdots+{m_t-1\choose t-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
By contradiction, assume that&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}_1&#039;|&amp;lt;{m_k-1\choose k-1}+{m_{k-1}-1\choose k-2}+\cdots+{m_t-1\choose t-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Then by &amp;lt;math&amp;gt;|\mathcal{F}|=|\mathcal{F}_0|+|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;, it holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
|\mathcal{F}_0| &amp;amp;=m-|\mathcal{F}_1&#039;|\\&lt;br /&gt;
&amp;amp;&amp;gt;\left\{{m_k\choose k}- {m_k-1\choose k-1}\right\}+\cdots+\left\{{m_t\choose t}- {m_t-1\choose t-1}\right\}\\&lt;br /&gt;
&amp;amp;={m_k-1\choose k}+\cdots+{m_t-1\choose t},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
so that, by the induction hypothesis, &lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta\mathcal{F}_0|\ge{m_k-1\choose k-1}+{m_{k-1}-1\choose k-1}+\cdots+{m_t-1\choose t-1}&amp;gt;|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;.&lt;br /&gt;
On the other hand, by Lemma 2, &amp;lt;math&amp;gt;|\mathcal{F}_1&#039;|\ge|\Delta\mathcal{F}_0|&amp;lt;/math&amp;gt;. Thus &amp;lt;math&amp;gt;|\mathcal{F}_1&#039;|\ge|\Delta\mathcal{F}_0|&amp;gt;|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;, a contradiction.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Now we officially apply the induction. By Lemma 1, &lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta\mathcal{F}|\ge|\Delta\mathcal{F}_1&#039;|+|\mathcal{F}_1&#039;|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;\mathcal{F}_1&#039;\subseteq{X\choose k-1}&amp;lt;/math&amp;gt;, and due to Lemma 3, &lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}_1&#039;|\ge{m_k-1\choose k-1}+{m_{k-1}-1\choose k-2}+\cdots+{m_t-1\choose t-1}&amp;lt;/math&amp;gt;, &lt;br /&gt;
thus by the induction hypothesis, &lt;br /&gt;
:&amp;lt;math&amp;gt;|\Delta\mathcal{F}_1&#039;|\ge{m_k-1\choose k-2}+{m_{k-1}-1\choose k-3}+\cdots+{m_t-1\choose t-2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
Combining them together, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
|\Delta\mathcal{F}|&lt;br /&gt;
&amp;amp;\ge |\Delta\mathcal{F}_1&#039;|+|\mathcal{F}_1&#039;|\\&lt;br /&gt;
&amp;amp;\ge {m_k-1\choose k-2}+\cdots+{m_t-1\choose t-2}+{m_k-1\choose k-1}+\cdots+{m_t-1\choose t-1}\\&lt;br /&gt;
&amp;amp;= {m_k\choose k-1}+{m_{k-1}\choose k-2}+\cdots+{m_t\choose t-1}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Shadows of specific sizes ===&lt;br /&gt;
The definition of shadow can be generalized to the subsets of any size.&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:The &#039;&#039;&#039;&amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-shadow&#039;&#039;&#039; of &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; is defined as&lt;br /&gt;
::&amp;lt;math&amp;gt;\Delta_r\mathcal{F}=\left\{S\mid |S|=r\text{ and }\exists T\in\mathcal{F}, S\subseteq T\right\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
And the general version of the Kruskal-Katona theorem can be deduced.&lt;br /&gt;
{{Theorem|Kruskal-Katona Theorem (general version)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|\mathcal{F}|=m&amp;lt;/math&amp;gt;, and suppose that&lt;br /&gt;
::&amp;lt;math&amp;gt;m={m_k\choose k}+{m_{k-1}\choose k-1}+\cdots+{m_t\choose t}&amp;lt;/math&amp;gt;&lt;br /&gt;
:where &amp;lt;math&amp;gt;m_k&amp;gt;m_{k-1}&amp;gt;\cdots&amp;gt;m_t\ge t\ge 1&amp;lt;/math&amp;gt;. Then for all &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le r\le k&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\left|\Delta_r\mathcal{F}\right|\ge {m_k\choose r}+{m_{k-1}\choose r-1}+\cdots+{m_t\choose t-k+r}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Note that for &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Delta_r\mathcal{F}=\underbrace{\Delta\cdots\Delta}_{k-r}\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The theorem follows by repeatedly applying the Kruskal-Katona theorem for &amp;lt;math&amp;gt;\Delta\mathcal{F}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== The Erdős–Ko–Rado theorem, revisited===&lt;br /&gt;
To demonstrate the power of the Krulskal-Katona theorem, we show that it actually includes the Erdős–Ko–Rado theorem as a special case. The following elegant proof of the Erdős–Ko–Rado theorem is due to Daykin and Clements independently.&lt;br /&gt;
{{Theorem|Erdős–Ko–Rado Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathcal{F}\subseteq {X\choose k}&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;|X|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;n\ge 2k&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;S\cap T\neq\emptyset&amp;lt;/math&amp;gt; for any &amp;lt;math&amp;gt;S,T\in\mathcal{F}&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|\mathcal{F}|\le{n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Proof by the Kruskal-Katona theorem|(Daykin 1974, Clements 1976)&lt;br /&gt;
By contradiction, suppose that &amp;lt;math&amp;gt;|\mathcal{F}|&amp;gt;{n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We define the dual system &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathcal{G}=\{\bar{S}\mid S\in\mathcal{F}\}&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;\bar{S}=X\setminus S&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;S,T\in\mathcal{F}&amp;lt;/math&amp;gt;, the condition &amp;lt;math&amp;gt;S\cap T\neq \emptyset&amp;lt;/math&amp;gt; is equivalent to &amp;lt;math&amp;gt;S\not\subseteq \bar{T}&amp;lt;/math&amp;gt;, so &amp;lt;math&amp;gt;\mathcal{F}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\Delta_k\mathcal{G}&amp;lt;/math&amp;gt; are disjoint, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|+|\Delta_k\mathcal{G}|\le{n\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Clearly, &amp;lt;math&amp;gt;\mathcal{G}\subseteq {X\choose n-k}&amp;lt;/math&amp;gt; and &lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{G}|=|\mathcal{F}|&amp;gt;{n-1\choose k-1}={n-1\choose n-k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By the Kruskal-Katona theorem, &amp;lt;math&amp;gt;|\Delta_k\mathcal{G}|\ge{n-1\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;|\mathcal{F}|+|\Delta_k\mathcal{G}|&amp;gt;{n-1\choose k-1}+{n-1\choose k}={n\choose k}&amp;lt;/math&amp;gt;,&lt;br /&gt;
a contradiction.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13735</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13735"/>
		<updated>2026-05-13T08:48:27Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/04/21)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第二次作业已发布&amp;lt;/font&amp;gt;，请在 2026/05/13 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A2.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 2|Problem Set 2]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal graph theory|Extremal graph theory | 极值图论]] ([http://tcs.nju.edu.cn/slides/comb2026/ExtremalGraphs.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal set theory|Extremal set theory | 极值集合论]]（[http://tcs.nju.edu.cn/slides/comb2026/ExtremalSets.pdf slides]）&lt;br /&gt;
#* [https://mathweb.ucsd.edu/~ronspubs/90_03_erdos_ko_rado.pdf Old and new proofs of the Erdős–Ko–Rado theorem] by Frankl and Graham&lt;br /&gt;
#* An [http://tcs.nju.edu.cn/slides/comb2026/sunflower-note.pdf LLM-generated lecture note] on Alweiss-Lovet-Wu-Zhang&#039;s improvement over the sunflower lemma, with simplified proofs by Rao-Tao&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13734</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13734"/>
		<updated>2026-05-13T08:46:57Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/04/21)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第二次作业已发布&amp;lt;/font&amp;gt;，请在 2026/05/13 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A2.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 2|Problem Set 2]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal graph theory|Extremal graph theory | 极值图论]] ([http://tcs.nju.edu.cn/slides/comb2026/ExtremalGraphs.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal set theory|Extremal set theory | 极值集合论]]（[http://tcs.nju.edu.cn/slides/comb2026/ExtremalSets.pdf slides]）&lt;br /&gt;
#* [https://mathweb.ucsd.edu/~ronspubs/90_03_erdos_ko_rado.pdf Old and new proofs of the Erdős–Ko–Rado theorem] by Frankl and Graham&lt;br /&gt;
#* A [http://tcs.nju.edu.cn/slides/comb2026/sunflower-note.pdf note] generated by ChatGPT on Alweiss-Lovet-Wu-Zhang&#039;s improved bound on the sunflower lemma, simplified by Rao-Tao&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13733</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13733"/>
		<updated>2026-05-13T08:42:49Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/04/21)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第二次作业已发布&amp;lt;/font&amp;gt;，请在 2026/05/13 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A2.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 2|Problem Set 2]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal graph theory|Extremal graph theory | 极值图论]] ([http://tcs.nju.edu.cn/slides/comb2026/ExtremalGraphs.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Extremal_graph_theory&amp;diff=13718</id>
		<title>组合数学 (Fall 2026)/Extremal graph theory</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Extremal_graph_theory&amp;diff=13718"/>
		<updated>2026-05-06T05:01:32Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== Forbidden Cliques == Extremal graph theory studies the problems like  &amp;quot;how many edges that a graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; can have, if &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has some property?&amp;quot; === Mantel&amp;#039;s theorem === We consider a typical extremal problem for graphs: the largest possible number of edges of &amp;#039;&amp;#039;&amp;#039;triangle-free&amp;#039;&amp;#039;&amp;#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.  {{Theorem|Theorem (Mantel 1907)| :Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;m...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Forbidden Cliques ==&lt;br /&gt;
Extremal graph theory studies the problems like  &amp;quot;how many edges that a graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; can have, if &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has some property?&amp;quot;&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We give three different proofs of the theorem. The first one uses induction and an argument based on pigeonhole principle. The second proof uses the famous Cauchy-Schwarz inequality in analysis. And the third proof uses another famous inequality: the inequality of the arithmetic and geometric mean.&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge \sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
The famous Turán&#039;s theorem generalizes the Mantel&#039;s theorem for triangles to cliques of any specific size. This theorem is one of the most important results in extremal combinatorics, which initiates the studies of extremal graph theory.&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We give an example of graphs with many edges which does not contain &amp;lt;math&amp;gt;K_r&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Partition &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;r-1&amp;lt;/math&amp;gt; disjoint classes &amp;lt;math&amp;gt;V=V_1\cup V_2\cup\cdots\cup V_{r-1}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;n_i=|V_i|&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;n_1+n_2+\cdots+n_{r-1}=n&amp;lt;/math&amp;gt;. For every two vertice &amp;lt;math&amp;gt;u,v&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;u\in V_i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v\in V_j&amp;lt;/math&amp;gt; for distinct &amp;lt;math&amp;gt;V_i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;V_j&amp;lt;/math&amp;gt;. The resulting graph is a &#039;&#039;&#039;complete &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;-partite graph&#039;&#039;&#039;, denoted &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt;. It is obvious that any &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;-partite graph contains no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique since only those vertices from different classes can be adjacent. &lt;br /&gt;
&lt;br /&gt;
A &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\sum_{i&amp;lt;j}n_i n_j\,&amp;lt;/math&amp;gt; edges, which is maximized when the numbers &amp;lt;math&amp;gt;n_i&amp;lt;/math&amp;gt; are divided as evenly as possible, that is, if &amp;lt;math&amp;gt;n_i\in\left\{\left\lfloor\frac{n}{r-1}\right\rfloor,\left\lceil\frac{n}{r-1}\right\rceil\right\}&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;1\le i\le r-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:We call a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_i\in\left\{\left\lfloor\frac{n}{r-1}\right\rfloor,\left\lceil\frac{n}{r-1}\right\rceil\right\}&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; a &#039;&#039;&#039; Turán graph&#039;&#039;&#039;, denoted &amp;lt;math&amp;gt;T(n,r-1)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
;Example:Turán graph &amp;lt;math&amp;gt;T(13,4)&amp;lt;/math&amp;gt;&lt;br /&gt;
[[File:Turan 13-4.svg|center|260px|Turán graph &amp;lt;math&amp;gt;T(13,4)&amp;lt;/math&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
Turán&#039;s theorem has been proved for many times by different mathematicians, with different tools. We show just a few.&lt;br /&gt;
&lt;br /&gt;
The first proof uses induction;  the second proof uses a technique called &amp;quot;weight shifting&amp;quot;; and the third proof uses the probabilistic method. All of them are very powerful and frequently used proof techniques.&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
&lt;br /&gt;
Induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. It is easy to verify that the theorem holds for &amp;lt;math&amp;gt;n&amp;lt;r&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; be a graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices without &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-cliques where &amp;lt;math&amp;gt;n\ge r&amp;lt;/math&amp;gt;. Suppose that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a maximum number of edges among such graphs. &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; certainly has &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;-cliques, since otherwise we could add edges to &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;-clique and let &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt;. Clearly &amp;lt;math&amp;gt;|A|=r-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|B|=n-r+1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By the  induction hypothesis, since &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-cliques, &amp;lt;math&amp;gt;|E(B)|\le\frac{r-2}{2(r-1)}(n-r+1)^2&amp;lt;/math&amp;gt;. And &amp;lt;math&amp;gt;E(A)={r-1\choose 2}&amp;lt;/math&amp;gt;. Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, every &amp;lt;math&amp;gt;v\in B&amp;lt;/math&amp;gt; is adjacent to at most &amp;lt;math&amp;gt;r-2&amp;lt;/math&amp;gt; vertices in &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, since otherwise &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; would form an &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique. We obtain that the number edges crossing between &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;|E(A,B)|\le (r-2)|B|=(r-2)(n-r+1)&amp;lt;/math&amp;gt;. Combining everything together,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|=|E(A)|+|E(B)|+|E(A,B)|\le {r-1\choose 2}+\frac{r-2}{2(r-1)}(n-r+1)^2+(r-2)(n-r+1)=\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
&lt;br /&gt;
Assign each vertex &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt; a nonnegative weight &amp;lt;math&amp;gt;w_v\ge 0&amp;lt;/math&amp;gt;, and assume that &amp;lt;math&amp;gt;\sum_{v\in V}w_v=1&amp;lt;/math&amp;gt;. We try to maximize the quantity&lt;br /&gt;
:&amp;lt;math&amp;gt;S=\sum_{uv\in E}w_uw_v&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let &amp;lt;math&amp;gt;W_u=\sum_{v:v\sim u}w_v\,&amp;lt;/math&amp;gt; be the sum of the weights of &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;&#039;s neighbors.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; can also be computed as &amp;lt;math&amp;gt;S=\frac{1}{2}\sum_{u\in V}w_uW_u&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any nonadjacent pair of vertices &amp;lt;math&amp;gt;u\not\sim v&amp;lt;/math&amp;gt;, supposed that &amp;lt;math&amp;gt;W_u\ge W_v&amp;lt;/math&amp;gt;, then for any &amp;lt;math&amp;gt;\epsilon\ge 0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;(w_u+\epsilon)W_u+(w_v-\epsilon)W_v\ge w_uW_u+w_vW_v&amp;lt;/math&amp;gt;.&lt;br /&gt;
This means that we do not decrease &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; by shifting all of the weight of the vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; to the vertex &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;. It follows that &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is maximized when all of the weight is concentrated on a complete subgraph, i.e., a clique.&lt;br /&gt;
&lt;br /&gt;
Now if &amp;lt;math&amp;gt;w_u&amp;gt;w_v&amp;gt;0&amp;lt;/math&amp;gt;, then choose &amp;lt;math&amp;gt;\epsilon&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;0&amp;lt;\epsilon&amp;lt;w_u-w_v&amp;lt;/math&amp;gt; and change &amp;lt;math&amp;gt;w_u&#039;=w_u-\epsilon&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;w_v&#039;=w_v+\epsilon&amp;lt;/math&amp;gt;. This changes &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;S&#039;=S+\epsilon(w_u-w_v)-\epsilon^2&amp;gt;S&amp;lt;/math&amp;gt;. Thus, the maximal value of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;  is attained when all nonzero weights are equal and concentrated on a clique.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has at most an &amp;lt;math&amp;gt;(r-1)&amp;lt;/math&amp;gt;-clique, thus &amp;lt;math&amp;gt;S\le{r-1\choose 2}\frac{1}{(r-1)^2}=\frac{r-2}{2(r-1)}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
As we argued above, this inequality hold for any nonnegative weight assignments with &amp;lt;math&amp;gt;\sum_{v\in V}w_v=1&amp;lt;/math&amp;gt;. In particular, for the case that all &amp;lt;math&amp;gt;w_v=\frac{1}{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;S=\sum_{uv\in E}w_uw_v=\frac{|E|}{n^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{|E|}{n^2}\le \frac{r-2}{2(r-1)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
which implies the theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
&lt;br /&gt;
Write &amp;lt;math&amp;gt;\omega(G)&amp;lt;/math&amp;gt; for the number of vertices in a largest clique, called the &#039;&#039;&#039;clique number&#039;&#039;&#039; of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. &lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;\omega(G)\ge\sum_{v\in V}\frac{1}{n-d_v}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We prove this by the probabilistic method. Fix a random ordering of vertices in &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;, say &amp;lt;math&amp;gt;v_1,v_2,\ldots,v_n&amp;lt;/math&amp;gt;. We construct a clique as follows:&lt;br /&gt;
*for &amp;lt;math&amp;gt;i=1,2,\ldots, n&amp;lt;/math&amp;gt;, add &amp;lt;math&amp;gt;v_i&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; iff all vertices in current &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; are adjacent to &amp;lt;math&amp;gt;v_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is obvious that an &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; constructed in this way is a clique. We now show that &amp;lt;math&amp;gt;\mathbf{E}[|S|]\ge\sum_{v\in V}\frac{1}{n-d_v}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X_v&amp;lt;/math&amp;gt; be the random variable that indicates whether &amp;lt;math&amp;gt;v\in S&amp;lt;/math&amp;gt;, i.e.,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X_v=\begin{cases}&lt;br /&gt;
1 &amp;amp; v\in S,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Note that a vertex &amp;lt;math&amp;gt;v\in S&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; is ranked before all its &amp;lt;math&amp;gt;n-d_v-1&amp;lt;/math&amp;gt; non-neighbors in the random ordering. The probability that this event occurs is &amp;lt;math&amp;gt;\frac{1}{n-d_v}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_v]=\Pr[v\in S]\ge\frac{1}{n-d_v}.&amp;lt;/math&amp;gt;&lt;br /&gt;
Observe that &amp;lt;math&amp;gt;|S|=\sum_{v\in V}X_v&amp;lt;/math&amp;gt;. Due to linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[|S|]=\sum_{v\in V}\mathbf{E}[X_v]\ge\sum_{v\in V}\frac{1}{n-d_v}&amp;lt;/math&amp;gt;.&lt;br /&gt;
There must exists a clique of at least such size, so that &amp;lt;math&amp;gt;\omega(G)\ge\sum_{v\in V}\frac{1}{n-d_v}&amp;lt;/math&amp;gt;. The claim is proved.&lt;br /&gt;
&lt;br /&gt;
Apply the Cauchy-Schwarz inequality&lt;br /&gt;
:&amp;lt;math&amp;gt;\left(\sum_{v\in V}a_vb_v\right)^2\le\left(\sum_{v\in V}^na_v^2\right)\left(\sum_{v\in V}^nb_v^2\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Set &amp;lt;math&amp;gt;a_v=\sqrt{n-d_v}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;b_v=\frac{1}{\sqrt{n-d_v}}&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;a_vb_v=1&amp;lt;/math&amp;gt; and so&lt;br /&gt;
:&amp;lt;math&amp;gt;n^2\le\sum_{v\in V}(n-d_v)\sum_{v\in V}\frac{1}{n-d_v}\le\omega(G)\sum_{v\in V}(n-d_v).&amp;lt;/math&amp;gt;&lt;br /&gt;
By the assumption of Turán&#039;s theorem, &amp;lt;math&amp;gt;\omega(G)\le r-1&amp;lt;/math&amp;gt;. Recall the handshaking lemma &amp;lt;math&amp;gt;2|E|=\sum_{v\in V}d_v&amp;lt;/math&amp;gt;. The above inequality gives us&lt;br /&gt;
:&amp;lt;math&amp;gt;n^2\le (r-1)(n^2-2|E|)&amp;lt;/math&amp;gt;,&lt;br /&gt;
which implies the theorem.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Our last proof uses the idea of vertex duplication. It does not only prove the edge bound of Turán&#039;s theorem, but also shows that Turán graphs are the &amp;lt;font color=red&amp;gt;only&amp;lt;/font&amp;gt; possible extremal graphs.&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Forbidden Cycles ==&lt;br /&gt;
Another direction to generalize Mantel&#039;s theorem other than Turán&#039;s theorem is to see a triangle as a 3-cycle rather than 3-clique. We then ask for the extremal bound for graphs without certain cycle structures.&lt;br /&gt;
Recall that the &#039;&#039;&#039;girth&#039;&#039;&#039; of a graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is the length of the shortest cycle in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. A graph is triangle-free if and only if its girth &amp;lt;math&amp;gt;g(G)\ge 4&amp;lt;/math&amp;gt;.&lt;br /&gt;
Matel&#039;s theorem can be seen as a bound on the edge number of graphs with girth &amp;lt;math&amp;gt;g(G)\ge 4&amp;lt;/math&amp;gt;. The next theorem extends this bound to the graphs with &amp;lt;math&amp;gt;g(G)\ge 5&amp;lt;/math&amp;gt;, i.e., graphs without triangles and quadrilaterals (&amp;quot;squares&amp;quot;).&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. If girth &amp;lt;math&amp;gt;g(G)\ge 5&amp;lt;/math&amp;gt; then &amp;lt;math&amp;gt;|E|\le\frac{1}{2}n\sqrt{n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Suppose &amp;lt;math&amp;gt;g(G)\ge 5&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;v_1,v_2,\ldots,v_d&amp;lt;/math&amp;gt; be the neighbors of a vertex &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;d=d(u)&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;S_i=\{v\in V\mid v\sim v_i\wedge v\neq u\}&amp;lt;/math&amp;gt; be the set of neighbors of &amp;lt;math&amp;gt;v_i&amp;lt;/math&amp;gt; other than &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
* For any &amp;lt;math&amp;gt;v_i,v_j&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;v_iv_j\not\in E&amp;lt;/math&amp;gt; since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no triangle. Thus, &amp;lt;math&amp;gt;S_i\cap\{u,v_1,v_2,\ldots,v_d\}=\emptyset&amp;lt;/math&amp;gt; for every &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;.&lt;br /&gt;
* No vertex other than &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; can be adjacent to more than one vertices in &amp;lt;math&amp;gt;v_1,v_2,\ldots,v_d&amp;lt;/math&amp;gt; since there is no &amp;lt;math&amp;gt;C_4&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;S_i\cap S_j=\emptyset&amp;lt;/math&amp;gt; for any distinct &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;\{u,v_1,v_2,\ldots,v_d\}\cup S_1\cup S_2\cup\cdots\cup S_d\subseteq V&amp;lt;/math&amp;gt; implies that&lt;br /&gt;
:&amp;lt;math&amp;gt;(d+1)+|S_1|+|S_2|+\cdots+|S_d|=(d+1)+(d(v_1)-1)+(d(v_2)-1)+\cdots+(d(v_d)-1)\le n&amp;lt;/math&amp;gt;,&lt;br /&gt;
so that &amp;lt;math&amp;gt;\sum_{v:v\sim u}d(v)\le n-1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By Cauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;n(n-1)\ge \sum_{u\in V}\sum_{v:v\sim u}d(v)=\sum_{v\in V}d(v)^2\ge\frac{\left(\sum_{v\in V}d(v)\right)}{n}=\frac{4|E|^2}{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
which implies that &amp;lt;math&amp;gt;|E|\le\frac{1}{2}n\sqrt{n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Erdős–Stone theorem ==&lt;br /&gt;
We introduce a notation for the number of edges in extremal graphs with a specific forbidden substructure.&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\mathrm{ex}(n,H)&amp;lt;/math&amp;gt; denote the largest number of edges that a graph &amp;lt;math&amp;gt;G\not\supseteq H&amp;lt;/math&amp;gt; on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices can have.&lt;br /&gt;
}}&lt;br /&gt;
With this notation, Turán&#039;s theorem can be restated as&lt;br /&gt;
{{Theorem|Turán&#039;s theorem (restated)|&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathrm{ex}(n,K_r)\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;K_s^r=K_{\underbrace{s,s,\cdots,s}_{r}}&amp;lt;/math&amp;gt; be the complete &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-partite graph with &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; vertices in each class, i.e., the Turán graph &amp;lt;math&amp;gt;T(rs,r)&amp;lt;/math&amp;gt;.&lt;br /&gt;
The Erdős–Stone theorem (also referred as the &#039;&#039;&#039;fundamental theorem of extremal graph theory&#039;&#039;&#039;) gives an asymptotic bound on &amp;lt;math&amp;gt;\mathrm{ex}(n,K_s^r)&amp;lt;/math&amp;gt;, i.e., the largest number of edges that an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-vertex graph can have to not contain &amp;lt;math&amp;gt;K_s^r&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any integers &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and any &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is sufficiently large then every graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices and with at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph, i.e.,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\mathrm{ex}(n,K_s^r)= \left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The theorem is called fundamental because of its single most important corollary: it relate the extremal bound for an arbitrary subgraph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; to a very natural parameter of &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;, its chromatic number.&lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;\chi(G)&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;chromatic number&#039;&#039;&#039; of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, the smallest number of colors that one can use to color the vertices so that no adjacent vertices have the same color.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:For every nonempty graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\lim_{n\rightarrow\infty}\frac{\mathrm{ex}(n,H)}{{n\choose 2}}=\frac{\chi(H)-2}{\chi(H)-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Prooftitle|Proof of corollary|&lt;br /&gt;
Let &amp;lt;math&amp;gt;r=\chi(H)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;T(n,r-1)&amp;lt;/math&amp;gt; can be colored with &amp;lt;math&amp;gt;r-1&amp;lt;/math&amp;gt; colors, one color for each part. Thus, &amp;lt;math&amp;gt;H\not\subseteq T(n,r-1)&amp;lt;/math&amp;gt;, since otherwise &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; can also be colored with &amp;lt;math&amp;gt;r-1&amp;lt;/math&amp;gt; colors, contradicting that &amp;lt;math&amp;gt;\chi(H)=1&amp;lt;/math&amp;gt;. By definition, &amp;lt;math&amp;gt;\mathrm{ex}(n,H)&amp;lt;/math&amp;gt; is the maximum number of edges that an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-vertex graph &amp;lt;math&amp;gt;G\not\supseteq H&amp;lt;/math&amp;gt; can have. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;|T(n,r-1)|\le\mathrm{ex}(n,H)&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is not hard to see that &lt;br /&gt;
:&amp;lt;math&amp;gt;|T(n,r-1)|\ge {r-1\choose 2}\left\lfloor\frac{n}{r-1}\right\rfloor^2\ge{r-1\choose 2}\left(\frac{n}{r-1}-1\right)^2=\left(\frac{r-2}{2(r-1)}-o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
On the other hand, any finite graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; with chromatic number &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; has that &amp;lt;math&amp;gt;H\subseteq K_s^r&amp;lt;/math&amp;gt; for all sufficiently large &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt;. We just connect all pairs of vertices from different color classes. Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathrm{ex}(n,H)\le\mathrm{ex}(n,K_s^r)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to Erdős–Stone theorem,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathrm{ex}(n,K_s^r)=\left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Altogether, we have &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\frac{r-2}{r-1}-o(1)\le\frac{|T(n,r-1)|}{{n\choose 2}}\le \frac{\mathrm{ex}(n,H)}{{n\choose 2}} \le \frac{\mathrm{ex}(n,K_s^r)}{{n\choose 2}}=\frac{r-2}{r-1}+o(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
* van Lin and Wilson. &#039;&#039;A course in combinatorics.&#039;&#039; Cambridge Press. Chapter 4.&lt;br /&gt;
* Aigner and Ziegler. &#039;&#039;Proofs from THE BOOK, 4th Edition.&#039;&#039; Springer-Verlag. [[media:PFTB_chap36.pdf| Chapter 36]]. &lt;br /&gt;
* Diestel. &#039;&#039;Graph Theory, 3rd Edition&#039;&#039;. Springer-Verlag 2000. [[media:Diestel2ed_chap7.pdf|Chapter 7]].&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13717</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13717"/>
		<updated>2026-05-06T05:00:57Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
* &#039;&#039;&#039;(2026/04/21)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第二次作业已发布&amp;lt;/font&amp;gt;，请在 2026/05/13 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A2.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 2|Problem Set 2]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Extremal graph theory|Extremal graph theory | 极值图论]] ([http://tcs.nju.edu.cn/slides/comb2025/ExtremalGraphs.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/The_probabilistic_method&amp;diff=13641</id>
		<title>组合数学 (Fall 2026)/The probabilistic method</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/The_probabilistic_method&amp;diff=13641"/>
		<updated>2026-04-17T13:18:44Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== The Probabilistic Method == The probabilistic method provides another way of proving the existence of objects: instead of explicitly constructing an object, we define a probability space of objects in which the probability is positive that a randomly selected object has the required property.  The basic principle of the probabilistic method is very simple, and can be stated in intuitive ways: *If an object chosen randomly from a universe satisfies a property with posi...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== The Probabilistic Method ==&lt;br /&gt;
The probabilistic method provides another way of proving the existence of objects: instead of explicitly constructing an object, we define a probability space of objects in which the probability is positive that a randomly selected object has the required property.&lt;br /&gt;
&lt;br /&gt;
The basic principle of the probabilistic method is very simple, and can be stated in intuitive ways:&lt;br /&gt;
*If an object chosen randomly from a universe satisfies a property with positive probability, then there must be an object in the universe that satisfies that property.&lt;br /&gt;
:For example, for a ball(the object) randomly chosen from a box(the universe) of balls, if the probability that the chosen ball is blue(the property) is &amp;gt;0, then there must be a blue ball in the box.&lt;br /&gt;
*Any random variable assumes at least one value that is no smaller than its expectation, and at least one value that is no greater than the expectation.&lt;br /&gt;
:For example, if we know the average height of the students in the class is &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;, then we know there is a students whose height is at least &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;, and there is a student whose height is at most &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Although the idea of  the probabilistic method is simple, it provides us a powerful tool for existential proof.&lt;br /&gt;
&lt;br /&gt;
===Ramsey number===&lt;br /&gt;
&lt;br /&gt;
Recall the Ramsey theorem which states that in a meeting of at least six people, there are either three people knowing each other or three people not knowing each other. In graph theoretical terms, this means that no matter how we color the edges of &amp;lt;math&amp;gt;K_6&amp;lt;/math&amp;gt; (the complete graph on six vertices), there must be a &#039;&#039;&#039;monochromatic&#039;&#039;&#039; &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt; (a triangle whose edges have the same color).&lt;br /&gt;
&lt;br /&gt;
Generally, the &#039;&#039;&#039;Ramsey number&#039;&#039;&#039; &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is the smallest integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; such that in any two-coloring of the edges of a complete graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; by red and blue, either there is a red &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; or there is a blue &amp;lt;math&amp;gt;K_\ell&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Ramsey showed in 1929 that &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is finite for any &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;. It is extremely hard to compute the exact value of &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt;. Here we give a lower bound of &amp;lt;math&amp;gt;R(k,k)&amp;lt;/math&amp;gt; by the probabilistic method.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem (Erdős 1947)|&lt;br /&gt;
:If &amp;lt;math&amp;gt;{n\choose k}\cdot 2^{1-{k\choose 2}}&amp;lt;1&amp;lt;/math&amp;gt; then it is possible to color the edges of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; with two colors so that there is no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; subgraph.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Consider a random two-coloring of edges of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; obtained as follows:&lt;br /&gt;
* For each edge of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt;, independently flip a fair coin to decide the color of the edge.&lt;br /&gt;
&lt;br /&gt;
For any fixed set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vertices, let &amp;lt;math&amp;gt;\mathcal{E}_S&amp;lt;/math&amp;gt; be the event that the &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; subgraph induced by &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is monochromatic. There are &amp;lt;math&amp;gt;{k\choose 2}&amp;lt;/math&amp;gt; many edges in &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;, therefore&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_S]=2\cdot 2^{-{k\choose 2}}=2^{1-{k\choose 2}}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Since there are &amp;lt;math&amp;gt;{n\choose k}&amp;lt;/math&amp;gt; possible choices of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, by the union bound&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\exists S, \mathcal{E}_S]\le {n\choose k}\cdot\Pr[\mathcal{E}_S]={n\choose k}\cdot 2^{1-{k\choose 2}}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Due to the assumption, &amp;lt;math&amp;gt;{n\choose k}\cdot 2^{1-{k\choose 2}}&amp;lt;1&amp;lt;/math&amp;gt;, thus there exists a two coloring that none of &amp;lt;math&amp;gt;\mathcal{E}_S&amp;lt;/math&amp;gt; occurs, which means  there is no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; subgraph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;k\ge 3&amp;lt;/math&amp;gt; and we take &amp;lt;math&amp;gt;n=\lfloor2^{k/2}\rfloor&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
{n\choose k}\cdot 2^{1-{k\choose 2}}&lt;br /&gt;
&amp;amp;&amp;lt;&lt;br /&gt;
\frac{n^k}{k!}\cdot\frac{2^{1+\frac{k}{2}}}{2^{k^2/2}}\\&lt;br /&gt;
&amp;amp;\le &lt;br /&gt;
\frac{2^{k^2/2}}{k!}\cdot\frac{2^{1+\frac{k}{2}}}{2^{k^2/2}}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{2^{1+\frac{k}{2}}}{k!}\\&lt;br /&gt;
&amp;amp;&amp;lt;1.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
By the above theorem, there exists a two-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; that there is no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;. Therefore, the Ramsey number &amp;lt;math&amp;gt;R(k,k)&amp;gt;\lfloor2^{k/2}\rfloor&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;k\ge 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
===Tournament===&lt;br /&gt;
A &#039;&#039;&#039;[http://en.wikipedia.org/wiki/Tournament_(graph_theory) tournament]&#039;&#039;&#039; (竞赛图) on a set &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; players is an &#039;&#039;&#039;orientation&#039;&#039;&#039; of the edges of the complete graph on the set of vertices &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;. Thus for every two distinct vertices &amp;lt;math&amp;gt;u,v&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt;, either &amp;lt;math&amp;gt;(u,v)\in E&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;(v,u)\in E&amp;lt;/math&amp;gt;, but not both.&lt;br /&gt;
&lt;br /&gt;
We can think of the set &amp;lt;math&amp;gt;V&amp;lt;/math&amp;gt; as a set of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; players in which each pair participates in a single match, where &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; is in the tournament iff player &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; beats player &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
:We say that a tournament has &#039;&#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-paradoxical&#039;&#039;&#039; if for every set of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; players there is a player who beats them all.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Is it true for every finite &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;, there is a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-paradoxical tournament (on more than &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vertices, of course)? This problem was first raised by Schütte, and as shown by Erdős, can be solved almost trivially by the probabilistic method.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Erdős 1963)|&lt;br /&gt;
:If &amp;lt;math&amp;gt;{n\choose k}\left(1-2^{-k}\right)^{n-k}&amp;lt;1&amp;lt;/math&amp;gt; then there is a tournament on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices that is &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-paradoxical.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Consider a uniformly random tournament &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; on the set &amp;lt;math&amp;gt;V=[n]&amp;lt;/math&amp;gt;. For every fixed subset &amp;lt;math&amp;gt;S\in{V\choose k}&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vertices, let &amp;lt;math&amp;gt;A_S&amp;lt;/math&amp;gt; be the event defined as follows&lt;br /&gt;
:&amp;lt;math&amp;gt;A_S:\,&amp;lt;/math&amp;gt; there is no vertex in &amp;lt;math&amp;gt;V\setminus S&amp;lt;/math&amp;gt; that beats all vertices in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In a uniform random tournament, the orientations of edges are independent. For any &amp;lt;math&amp;gt;u\in V\setminus S&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[u\mbox{ beats all }v\in S]=2^{-k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;\Pr[u\mbox{ does not beats all }v\in S]=1-2^{-k}&amp;lt;/math&amp;gt; and&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[A_S]=\prod_{u\in V\setminus S}\Pr[u\mbox{ does not beats all }v\in S]=(1-2^{-k})^{n-k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\bigvee_{S\in{V\choose k}}A_S\right]\le \sum_{S\in{V\choose k}}\Pr[A_S]={n\choose k}(1-2^{-k})^{n-k}&amp;lt;1.&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\,T\mbox{ is }k\mbox{-paradoxical }]=\Pr\left[\bigwedge_{S\in{V\choose k}}\overline{A_S}\right]=1-\Pr\left[\bigvee_{S\in{V\choose k}}A_S\right]&amp;gt;0.&amp;lt;/math&amp;gt; &lt;br /&gt;
There is a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-paradoxical tournament.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Linearity of expectation ===&lt;br /&gt;
Let &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; be a discrete &#039;&#039;&#039;random variable&#039;&#039;&#039;.  The expectation of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is defined as follows.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition (Expectation)|&lt;br /&gt;
:The &#039;&#039;&#039;expectation&#039;&#039;&#039; of a discrete random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, denoted by &amp;lt;math&amp;gt;\mathbf{E}[X]&amp;lt;/math&amp;gt;, is given by&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}[X] &amp;amp;= \sum_{x}x\Pr[X=x],&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
:where the summation is over all values &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; in the range of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
A fundamental fact regarding the expectation is its &#039;&#039;&#039;linearity&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem (Linearity of Expectations)|&lt;br /&gt;
:For any discrete random variables &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt;, and any real constants &amp;lt;math&amp;gt;a_1, a_2, \ldots, a_n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[\sum_{i=1}^n a_iX_i\right] &amp;amp;= \sum_{i=1}^n a_i\cdot\mathbf{E}[X_i].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;Hamiltonian paths&lt;br /&gt;
The following result of Szele in 1943 is often considered the first use of the probabilistic method.&lt;br /&gt;
{{Theorem|Theorem (Szele 1943)|&lt;br /&gt;
:There is a tournament on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; players with at least &amp;lt;math&amp;gt;n!2^{-(n-1)}&amp;lt;/math&amp;gt; Hamiltonian paths.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Consider the uniform random tournament &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; on &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;. For any permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_{\pi}&amp;lt;/math&amp;gt; be the indicator random variable defined as &lt;br /&gt;
:&amp;lt;math&amp;gt;X_{\pi}=\begin{cases}&lt;br /&gt;
1 &amp;amp; \forall i\in[n-1], (\pi_i,\pi_{i+1})\in T,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise}.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
In other words, &amp;lt;math&amp;gt;X_{\pi}&amp;lt;/math&amp;gt; indicates whether &amp;lt;math&amp;gt;\pi_0\rightarrow\pi_1\rightarrow\pi_2\rightarrow\cdots\rightarrow\pi_{n-1}&amp;lt;/math&amp;gt; gives a Hamiltonian path. &lt;br /&gt;
It holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathrm{E}[X_\pi]=1\cdot\Pr[X_\pi=1]+0\cdot\Pr[X_\pi=0]=\Pr[\forall i\in[n-1], (\pi_i,\pi_{i+1})\in T]=2^{-(n-1)}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X=\sum_{\pi:\text{permutation of }[n]}X_\pi\,&amp;lt;/math&amp;gt;. Clearly &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is the number of Hamiltonian paths in the tournament &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt;. &lt;br /&gt;
Due to the linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathrm{E}[X]=\mathrm{E}\left[\sum_{\pi:\text{permutation of }[n]}X_\pi\right]=\sum_{\pi:\text{permutation of }[n]}\mathrm{E}[X_\pi]=n!2^{-(n-1)}.&amp;lt;/math&amp;gt;&lt;br /&gt;
This is the average number of Hamiltonian paths in a tournament, where the average is taken over all tournaments.&lt;br /&gt;
Thus some tournament has at least &amp;lt;math&amp;gt;n!2^{-(n-1)}&amp;lt;/math&amp;gt; Hamiltonian paths.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
===Independent sets===&lt;br /&gt;
An independent set of a graph is a set of vertices with no edges between them. The following theorem gives a lower bound on the size of the largest independent set.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; edges. Then &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has an independent set with at least &amp;lt;math&amp;gt;\frac{n^2}{4m}&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| Let &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; be a set of vertices constructed as follows:&lt;br /&gt;
:For each vertex &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;:&lt;br /&gt;
:* &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; is included in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; independently with probability &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;,&lt;br /&gt;
&amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; to be determined.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X=|S|&amp;lt;/math&amp;gt;. It is obvious that &amp;lt;math&amp;gt;\mathbf{E}[X]=np&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For each edge &amp;lt;math&amp;gt;e\in E&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;Y_{e}&amp;lt;/math&amp;gt; be the random variable which indicates whether both endpoints of &amp;lt;math&amp;gt;e=uv&amp;lt;/math&amp;gt; are in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y_{uv}]=\Pr[u\in S\wedge v\in S]=p^2.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; be the number of edges in the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. It holds that &amp;lt;math&amp;gt;Y=\sum_{e\in E}Y_e&amp;lt;/math&amp;gt;. By linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[Y]=\sum_{e\in E}\mathbf{E}[Y_e]=mp^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Note that although &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is not necessary an independent set, it can be modified to one if for each edge &amp;lt;math&amp;gt;e&amp;lt;/math&amp;gt; of the induced subgraph &amp;lt;math&amp;gt;G(S)&amp;lt;/math&amp;gt;, we delete one of the endpoint of &amp;lt;math&amp;gt;e&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;S^*&amp;lt;/math&amp;gt; be the resulting set. It is obvious that &amp;lt;math&amp;gt;S^*&amp;lt;/math&amp;gt; is an independent set since there is no edge left in the induced subgraph &amp;lt;math&amp;gt;G(S^*)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Since there are &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; edges in &amp;lt;math&amp;gt;G(S)&amp;lt;/math&amp;gt;, there are at most &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; vertices in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; are deleted to make it become &amp;lt;math&amp;gt;S^*&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;|S^*|\ge X-Y&amp;lt;/math&amp;gt;. By linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[|S^*|]\ge\mathbf{E}[X-Y]=\mathbf{E}[X]-\mathbf{E}[Y]=np-mp^2.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The expectation is maximized when &amp;lt;math&amp;gt;p=\frac{n}{2m}&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[|S^*|]\ge n\cdot\frac{n}{2m}-m\left(\frac{n}{2m}\right)^2=\frac{n^2}{4m}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
There exists an independent set which contains at least &amp;lt;math&amp;gt;\frac{n^2}{4m}&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Coloring large-girth graphs ==&lt;br /&gt;
The girth of a graph is the length of the shortest cycle of the graph.&lt;br /&gt;
{{Theorem|Definition|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be an undirected graph.&lt;br /&gt;
* A &#039;&#039;&#039;cycle&#039;&#039;&#039; of length &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is a sequence of distinct vertices &amp;lt;math&amp;gt;v_1,v_2,\ldots,v_{k}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;v_iv_{i+1}\in E&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,2,\ldots,k-1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v_kv_1\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
* The &#039;&#039;&#039;girth&#039;&#039;&#039; of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, dented &amp;lt;math&amp;gt;g(G)&amp;lt;/math&amp;gt;, is the length of the shortest cycle in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The chromatic number of a graph is the minimum number of colors with which the graph can be &#039;&#039;properly&#039;&#039; colored.&lt;br /&gt;
{{Theorem|Definition (chromatic number)|&lt;br /&gt;
* The &#039;&#039;&#039;chromatic number&#039;&#039;&#039; of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, denoted &amp;lt;math&amp;gt;\chi(G)&amp;lt;/math&amp;gt;, is the minimal number of colors which we need to color the vertices of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; so that no two adjacent vertices have the same color. Formally,&lt;br /&gt;
::&amp;lt;math&amp;gt;\chi(G)=\min\{C\in\mathbb{N}\mid \exists f:V\rightarrow[C]\mbox{ such that }\forall uv\in E, f(u)\neq f(v)\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
In 1959, Erdős proved the following theorem: for any fixed &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;, there exists a finite graph with girth at least &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and chromatic number at least &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;. This is considered a striking example of the probabilistic method. The statement of the theorem itself calls for nothing about probability or randomness. And the result is highly contra-intuitive: if the girth is large there is no simple reason why the graph could not be colored with a few colors. We can always &amp;quot;locally&amp;quot; color a cycle with 2 or 3 colors. Erdős&#039; result shows that there are &amp;quot;global&amp;quot; restrictions for coloring, and although such configurations are very difficult to explicitly construct, with the probabilistic method, we know that they commonly exist.&lt;br /&gt;
&lt;br /&gt;
{{Theorem| Theorem (Erdős 1959)|&lt;br /&gt;
: For all &amp;lt;math&amp;gt;k,\ell&amp;lt;/math&amp;gt; there exists a graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;g(G)&amp;gt;\ell&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\chi(G)&amp;gt;k\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is very hard to directly analyze the chromatic number of a graph. We find that the chromatic number can be related to the size of the maximum independent set.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Definition (independence number)|&lt;br /&gt;
* The &#039;&#039;&#039;independence number&#039;&#039;&#039; of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, denoted &amp;lt;math&amp;gt;\alpha(G)&amp;lt;/math&amp;gt;, is the size of the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Formally,&lt;br /&gt;
::&amp;lt;math&amp;gt;\alpha(G)=\max\{|S|\mid S\subseteq V\mbox{ and }\forall u,v\in S, uv\not\in E\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We observe the following relationship between the chromatic number and the independence number.&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-vertex graph,&lt;br /&gt;
::&amp;lt;math&amp;gt;\chi(G)\ge\frac{n}{\alpha(G)}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
*In the optimal coloring, &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices are partitioned into &amp;lt;math&amp;gt;\chi(G)&amp;lt;/math&amp;gt; color classes according to the vertex color.&lt;br /&gt;
*Every color class is an independent set, or otherwise there exist two adjacent vertice with the same color.&lt;br /&gt;
*By the pigeonhole principle, there is a color class (hence an independent set) of size &amp;lt;math&amp;gt;\frac{n}{\chi(G)}&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;\alpha(G)\ge\frac{n}{\chi(G)}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The lemma follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Therefore, it is sufficient to prove that &amp;lt;math&amp;gt;\alpha(G)\le\frac{n}{k}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;g(G)&amp;gt;\ell&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Proof of Erdős theorem|&lt;br /&gt;
Fix &amp;lt;math&amp;gt;\theta&amp;lt;\frac{1}{\ell}&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; be &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;p=n^{\theta-1}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For any length-&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; simple cycle &amp;lt;math&amp;gt;\sigma&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_\sigma&amp;lt;/math&amp;gt; be the indicator random variable such that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
X_\sigma=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1 &amp;amp; \sigma\mbox{ is a cycle in }G,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise}.&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The number of cycles of length at most &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt; in graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;X=\sum_{i=3}^\ell\sum_{\sigma:i\text{-cycle}}X_\sigma&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For any particular length-&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; simple cycle &amp;lt;math&amp;gt;\sigma&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_\sigma]=\Pr[X_\sigma=1]=\Pr[\sigma\mbox{ is a cycle in }G]=p^i=n^{\theta i-i}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;3\le i\le n&amp;lt;/math&amp;gt;, the number of length-&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; simple cycle is &amp;lt;math&amp;gt;\frac{n(n-1)\cdots (n-i+1)}{2i}&amp;lt;/math&amp;gt;. By the linearity of expectation,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X]=\sum_{i=3}^\ell\sum_{\sigma:i\text{-cycle}}\mathbf{E}[X_\sigma]=\sum_{i=3}^\ell\frac{n(n-1)\cdots (n-i+1)}{2i}n^{\theta i-i}\le \sum_{i=3}^\ell\frac{n^{\theta i}}{2i}=o(n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Markov&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[X\ge \frac{n}{2}\right]\le\frac{\mathbf{E}[X]}{n/2}=o(1).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, with high probability the random graph has less than &amp;lt;math&amp;gt;n/2&amp;lt;/math&amp;gt; short cycles.&lt;br /&gt;
&lt;br /&gt;
Now we proceed to analyze the independence number. Let &amp;lt;math&amp;gt;m=\left\lceil\frac{3\ln n}{p}\right\rceil&amp;lt;/math&amp;gt;, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\alpha(G)\ge m]&lt;br /&gt;
&amp;amp;\le\Pr\left[\exists S\in{V\choose m}\forall \{u,v\}\in{S\choose 2}, uv\not\in G\right]\\&lt;br /&gt;
&amp;amp;\le{n\choose m}(1-p)^{m\choose 2}\\&lt;br /&gt;
&amp;amp;&amp;lt;n^m\mathrm{e}^{-p{m\choose 2}}\\&lt;br /&gt;
&amp;amp;=\left(n\mathrm{e}^{-p(m-1)/2}\right)^m=o(1)&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The probability that either of the above events occurs is &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[X\ge\frac{n}{2}\vee \alpha(G)\ge m\right]&lt;br /&gt;
\le \Pr\left[X\ge \frac{n}{2}\right]+\Pr\left[\alpha(G)\ge m\right]&lt;br /&gt;
=o(1).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Therefore, there exists a graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with less than &amp;lt;math&amp;gt;n/2&amp;lt;/math&amp;gt; &amp;quot;short&amp;quot; cycles, i.e., cycles of length at most &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt;, and with &amp;lt;math&amp;gt;\alpha(G)&amp;lt;m\le 3n^{1-\theta}\ln n&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Take each &amp;quot;short&amp;quot; cycle in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and remove a vertex from the cycle (and also remove all adjacent edges to the removed vertex). This gives a graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; which has no short cycles, hence the girth &amp;lt;math&amp;gt;g(G&#039;)\ge\ell&amp;lt;/math&amp;gt;. And &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has at least &amp;lt;math&amp;gt;n/2&amp;lt;/math&amp;gt; vertices, because at most &amp;lt;math&amp;gt;n/2&amp;lt;/math&amp;gt; vertices are removed.&lt;br /&gt;
&lt;br /&gt;
Notice that removing vertices cannot makes &amp;lt;math&amp;gt;\alpha(G)&amp;lt;/math&amp;gt; grow. It holds that &amp;lt;math&amp;gt;\alpha(G&#039;)\le\alpha(G)&amp;lt;/math&amp;gt;. Thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\chi(G&#039;)\ge\frac{n/2}{\alpha(G&#039;)}\ge\frac{n}{2m}\ge\frac{n^\theta}{6\ln n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The theorem is proved by taking &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; sufficiently large so that this value is greater than &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The proof contains a very simple procedure which for any &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt; &#039;&#039;generates&#039;&#039; such a graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;g(G)&amp;gt;\ell&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\chi(G)&amp;gt;k&amp;lt;/math&amp;gt;. The procedure is as such:&lt;br /&gt;
* Fix some &amp;lt;math&amp;gt;\theta&amp;lt;\frac{1}{\ell}&amp;lt;/math&amp;gt;. Choose sufficiently large &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;\frac{n^\theta}{6\ln n}&amp;gt;k&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;p=n^{\theta-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Generate a random graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt;.&lt;br /&gt;
* For each cycle of length at most &amp;lt;math&amp;gt;\ell&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;, remove a vertex from the cycle.&lt;br /&gt;
The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; satisfying that &amp;lt;math&amp;gt;g(G)&amp;gt;\ell&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\chi(G)&amp;gt;k&amp;lt;/math&amp;gt; with high probability.&lt;br /&gt;
&lt;br /&gt;
== Lovász Local Lemma==&lt;br /&gt;
Consider a set of &amp;quot;bad&amp;quot; events &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt;. Suppose that &amp;lt;math&amp;gt;\Pr[A_i]\le p&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;. We want to show that there is a situation that none of the bad events occurs. Due to the probabilistic method, we need to prove that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]&amp;gt;0.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
;Case 1&amp;lt;nowiki&amp;gt;: mutually independent events.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
If all the bad events &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; are mutually independent, then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]\ge(1-p)^n&amp;gt;0,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
for any &amp;lt;math&amp;gt;p&amp;lt;1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
;Case 2&amp;lt;nowiki&amp;gt;: arbitrarily dependent events.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
On the other hand, if we put no assumption on the dependencies between the events, then by the union bound (which holds unconditionally),&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]=1-\Pr\left[\bigvee_{i=1}^n A_i\right]\ge 1-np,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is not an interesting bound for &amp;lt;math&amp;gt;p\ge\frac{1}{n}&amp;lt;/math&amp;gt;. We cannot improve bound without further information regarding the dependencies between the events.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
We would like to know what is going on between the two extreme cases: mutually independent events, and arbitrarily dependent events. The Lovász local lemma provides such a tool.&lt;br /&gt;
&lt;br /&gt;
The local lemma is powerful tool for showing the possibility of rare event under &#039;&#039;limited dependencies&#039;&#039;. The structure of dependencies between a set of events is described by a &#039;&#039;&#039;dependency graph&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Definition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; be a set of events. A graph &amp;lt;math&amp;gt;D=(V,E)&amp;lt;/math&amp;gt; on the set of vertices &amp;lt;math&amp;gt;V=\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; is called a &#039;&#039;&#039;dependency graph&#039;&#039;&#039; for the events &amp;lt;math&amp;gt;A_1,\ldots,A_n&amp;lt;/math&amp;gt; if for each &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, the event &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; is mutually independent of all the events &amp;lt;math&amp;gt;\{A_j\mid (i,j)\not\in E\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
;Example&lt;br /&gt;
:Let &amp;lt;math&amp;gt;X_1,X_2,\ldots,X_m&amp;lt;/math&amp;gt; be a set of &#039;&#039;mutually independent&#039;&#039; random variables. Each event &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; is a predicate defined on a number of variables among &amp;lt;math&amp;gt;X_1,X_2,\ldots,X_m&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;v(A_i)&amp;lt;/math&amp;gt; be the unique smallest set of variables which determine &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt;. The dependency graph &amp;lt;math&amp;gt;D=(V,E)&amp;lt;/math&amp;gt; is defined by &lt;br /&gt;
:::&amp;lt;math&amp;gt;(i,j)\in E&amp;lt;/math&amp;gt; iff &amp;lt;math&amp;gt;v(A_i)\cap v(A_j)\neq \emptyset&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The following lemma, known as the Lovász local lemma, first proved by Erdős and Lovász in 1975, is an extremely powerful tool, as it supplies a way for dealing with rare events.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lovász Local Lemma (symmetric case)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; be a set of events, and assume that the following hold:&lt;br /&gt;
:#for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\Pr[A_i]\le p&amp;lt;/math&amp;gt;;&lt;br /&gt;
:#the maximum degree of the dependency graph for the events &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt;, and &lt;br /&gt;
:::&amp;lt;math&amp;gt;ep(d+1)\le 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We will prove a general version of the local lemma, where the events &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; are not symmetric. This generalization is due to Spencer.&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Lovász Local Lemma (general case)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;D=(V,E)&amp;lt;/math&amp;gt; be the dependency graph of events &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt;. Suppose there exist real numbers &amp;lt;math&amp;gt;x_1,x_2,\ldots, x_n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;0\le x_i&amp;lt;1&amp;lt;/math&amp;gt; and for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[A_i]\le x_i\prod_{(i,j)\in E}(1-x_j)&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Then &lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]\ge\prod_{i=1}^n(1-x_i)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We can use the following probability identity to compute the probability of the intersection of events:&lt;br /&gt;
{{Theorem|Chain rule|&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]=\prod_{i=1}^n\Pr\left[\overline{A_i}\mid \bigwedge_{j=1}^{i-1}\overline{A_{j}}\right]&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
By definition of conditional probability,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\overline{A_n}\mid\bigwedge_{i=1}^{n-1}\overline{A_{i}}\right]&lt;br /&gt;
=\frac{\Pr\left[\bigwedge_{i=1}^n\overline{A_{i}}\right]}&lt;br /&gt;
{\Pr\left[\bigwedge_{i=1}^{n-1}\overline{A_{i}}\right]}&amp;lt;/math&amp;gt;,&lt;br /&gt;
so we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{i=1}^n\overline{A_{i}}\right]=\Pr\left[\bigwedge_{i=1}^{n-1}\overline{A_{i}}\right]\Pr\left[\overline{A_n}\mid\bigwedge_{i=1}^{n-1}\overline{A_{i}}\right]&amp;lt;/math&amp;gt;.&lt;br /&gt;
The lemma is proved by recursively applying this equation.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Next we prove by induction on &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; that for any set of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; events &amp;lt;math&amp;gt;i_1,\ldots,i_m&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[A_{i_1}\mid \bigwedge_{j=2}^m\overline{A_{i_j}}\right]\le x_{i_1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The local lemma is a direct consequence of this by applying the chain rule.&lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;m=1&amp;lt;/math&amp;gt;, this is obvious. For general &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;i_2,\ldots,i_k&amp;lt;/math&amp;gt; be the set of vertices adjacent to  &amp;lt;math&amp;gt;i_1&amp;lt;/math&amp;gt; in the dependency graph. Clearly &amp;lt;math&amp;gt;k-1\le d&amp;lt;/math&amp;gt;. And it holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[A_{i_1}\mid \bigwedge_{j=2}^m\overline{A_{i_j}}\right]&lt;br /&gt;
=\frac{\Pr\left[ A_i\wedge \bigwedge_{j=2}^k\overline{A_{i_j}}\mid \bigwedge_{j=k+1}^m\overline{A_{i_j}}\right]}&lt;br /&gt;
{\Pr\left[\bigwedge_{j=2}^k\overline{A_{i_j}}\mid \bigwedge_{j=k+1}^m\overline{A_{i_j}}\right]}&lt;br /&gt;
&amp;lt;/math&amp;gt;,&lt;br /&gt;
which is due to the basic conditional probability identity&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[A\mid BC]=\frac{\Pr[AB\mid C]}{\Pr[B\mid C]}&amp;lt;/math&amp;gt;.&lt;br /&gt;
We bound the numerator&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[ A_{i_1}\wedge \bigwedge_{j=2}^k\overline{A_{i_j}}\mid \bigwedge_{j=k+1}^m\overline{A_{i_j}}\right]&lt;br /&gt;
&amp;amp;\le\Pr\left[ A_{i_1}\mid \bigwedge_{j=k+1}^m\overline{A_{i_j}}\right]\\&lt;br /&gt;
&amp;amp;=\Pr[A_{i_1}]\\&lt;br /&gt;
&amp;amp;\le x_{i_1}\prod_{(i_1,j)\in E}(1-x_j).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The equation is due to the independence between &amp;lt;math&amp;gt;A_{i_1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;A_{i_k+1},\ldots,A_{i_m}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The denominator can be expanded using the chain rule as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[\bigwedge_{j=2}^k\overline{A_{i_j}}\mid \bigwedge_{j=k+1}^m\overline{A_{i_j}}\right]&lt;br /&gt;
=\prod_{j=2}^k\Pr\left[\overline{A_{i_j}}\mid \bigwedge_{\ell=j+1}^m\overline{A_{i_\ell}}\right]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which by the induction hypothesis, is at least &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\prod_{j=2}^k(1-x_{i_j})=\prod_{\{i_1,i_j\}\in E}(1-x_j)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; is the edge set of the dependency graph.&lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left[A_{i_1}\mid \bigwedge_{j=2}^m\overline{A_{i_j}}\right]&lt;br /&gt;
\le\frac{x_{i_1}\prod_{(i_1,j)\in E}(1-x_j)}{\prod_{\{i_1,i_j\}\in E}(1-x_j)}\le x_{i_1}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying the chain rule, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]&lt;br /&gt;
&amp;amp;=\prod_{i=1}^n\Pr\left[\overline{A_i}\mid \bigwedge_{j=1}^{i-1}\overline{A_{j}}\right]\\&lt;br /&gt;
&amp;amp;=\prod_{i=1}^n\left(1-\Pr\left[A_i\mid \bigwedge_{j=1}^{i-1}\overline{A_{j}}\right]\right)\\&lt;br /&gt;
&amp;amp;\ge\prod_{i=1}^n\left(1-x_i\right).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
To prove the symmetric case. Let &amp;lt;math&amp;gt;x_i=\frac{1}{d+1}&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i=1,2,\ldots,n&amp;lt;/math&amp;gt;. Note that &amp;lt;math&amp;gt;\left(1-\frac{1}{d+1}\right)^d&amp;gt;\frac{1}{\mathrm{e}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If the following conditions are satisfied:&lt;br /&gt;
:#for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\Pr[A_i]\le p&amp;lt;/math&amp;gt;;&lt;br /&gt;
:#&amp;lt;math&amp;gt;ep(d+1)\le 1&amp;lt;/math&amp;gt;;&lt;br /&gt;
then for all &amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[A_i]\le p\le\frac{1}{e(d+1)}&amp;lt;\frac{1}{d+1}\left(1-\frac{1}{d+1}\right)^d\le x_i\prod_{(i,j)\in E}(1-x_j)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the local lemma for general cases, this implies that&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{i=1}^n\overline{A_i}\right]\ge\prod_{i=1}^n(1-x_i)=\left(1-\frac{1}{d+1}\right)^n&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
This gives the symmetric version of local lemma.&lt;br /&gt;
&lt;br /&gt;
=== Ramsey number, revisited ===&lt;br /&gt;
{{Theorem|Ramsey number|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;k,\ell&amp;lt;/math&amp;gt; be positive integers. The Ramsey number &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is defined as the smallest integer satisfying:&lt;br /&gt;
:If &amp;lt;math&amp;gt;n\ge R(k,\ell)&amp;lt;/math&amp;gt;, for any coloring of edges of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; with two colors red and blue, there exists a red &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; or a blue &amp;lt;math&amp;gt;K_\ell&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The Ramsey theorem says that for any &amp;lt;math&amp;gt;k,\ell&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is finite. The actual value of &amp;lt;math&amp;gt;R(k,\ell)&amp;lt;/math&amp;gt; is extremely difficult to compute.&lt;br /&gt;
We can use the local lemma to prove a lower bound for the diagonal Ramsey number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,k)\ge Ck2^{k/2}&amp;lt;/math&amp;gt; for some constant &amp;lt;math&amp;gt;C&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
To prove a lower bound &amp;lt;math&amp;gt;R(k,k)&amp;gt;n&amp;lt;/math&amp;gt;, it is sufficient to show that there exists a 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; without a monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;. We prove this by the probabilistic method.&lt;br /&gt;
&lt;br /&gt;
Pick a random 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; by coloring each edge uniformly and independently with one of the two colors. For any set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; vertices, let &amp;lt;math&amp;gt;A_S&amp;lt;/math&amp;gt; denote the event that &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; forms a monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;. It is easy to see that &amp;lt;math&amp;gt;\Pr[A_s]=2^{1-{k\choose 2}}=p&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subset &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; of vertices, &amp;lt;math&amp;gt;A_S&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;A_T&amp;lt;/math&amp;gt; are dependent if and only if &amp;lt;math&amp;gt;|S\cap T|\ge 2&amp;lt;/math&amp;gt;. For each &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, the number of &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;|S\cap T|\ge 2&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;{k\choose 2}{n\choose k-2}&amp;lt;/math&amp;gt;, so the max degree of the dependency graph is &amp;lt;math&amp;gt;d\le{k\choose 2}{n\choose k-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;n=Ck2^{k/2}&amp;lt;/math&amp;gt; for some appropriate constant &amp;lt;math&amp;gt;C&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\mathrm{e}p(d+1)&lt;br /&gt;
&amp;amp;\le \mathrm{e}2^{1-{k\choose 2}}\left({k\choose 2}{n\choose k-2}+1\right)\\&lt;br /&gt;
&amp;amp;\le 2^{3-{k\choose 2}}{k\choose 2}{n\choose k-2}\\&lt;br /&gt;
&amp;amp;\le 1&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying the local lemma, the probability that there is no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr\left[\bigwedge_{S\in{[n]\choose k}}\overline{A_S}\right]&amp;gt;0&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, there exists a 2-coloring of &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt; which has no monochromatic &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;, which means&lt;br /&gt;
:&amp;lt;math&amp;gt;R(k,k)&amp;gt;n=Ck2^{k/2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13640</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13640"/>
		<updated>2026-04-17T13:18:00Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2026/ProbMethod.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13639</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13639"/>
		<updated>2026-04-17T13:17:48Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
# [[组合数学 (Fall 2026)/The probabilistic method|The probabilistic method | 概率法]] ([http://tcs.nju.edu.cn/slides/comb2025/ProbMethod.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)/Weierstrass_Approximation_Theorem&amp;diff=13638</id>
		<title>概率论与数理统计 (Spring 2026)/Weierstrass Approximation Theorem</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)/Weierstrass_Approximation_Theorem&amp;diff=13638"/>
		<updated>2026-04-16T08:13:07Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;[https://en.wikipedia.org/wiki/Stone%E2%80%93Weierstrass_theorem &amp;#039;&amp;#039;&amp;#039;魏尔施特拉斯逼近定理&amp;#039;&amp;#039;&amp;#039;]（&amp;#039;&amp;#039;&amp;#039;Weierstrass approximation theorem&amp;#039;&amp;#039;&amp;#039;）陈述了这样一个事实：闭区间上的连续函数总可以用多项式一致逼近。 {{Theorem|魏尔施特拉斯逼近定理| :设 &amp;lt;math&amp;gt;f:[a,b]\to\mathbb{R}&amp;lt;/math&amp;gt; 为定义在实数区间 &amp;lt;math&amp;gt;[a,b]&amp;lt;/math&amp;gt; 上的连续实值函数。对每个 &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;，存在一个多项式 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 使得对...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[https://en.wikipedia.org/wiki/Stone%E2%80%93Weierstrass_theorem &#039;&#039;&#039;魏尔施特拉斯逼近定理&#039;&#039;&#039;]（&#039;&#039;&#039;Weierstrass approximation theorem&#039;&#039;&#039;）陈述了这样一个事实：闭区间上的连续函数总可以用多项式一致逼近。&lt;br /&gt;
{{Theorem|魏尔施特拉斯逼近定理|&lt;br /&gt;
:设 &amp;lt;math&amp;gt;f:[a,b]\to\mathbb{R}&amp;lt;/math&amp;gt; 为定义在实数区间 &amp;lt;math&amp;gt;[a,b]&amp;lt;/math&amp;gt; 上的连续实值函数。对每个 &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;，存在一个多项式 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 使得对于 &amp;lt;math&amp;gt;[a,b]&amp;lt;/math&amp;gt; 中所有 &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;，均有 &amp;lt;math&amp;gt;|p(x)-f(x)|\le \epsilon&amp;lt;/math&amp;gt;，即&lt;br /&gt;
:::&amp;lt;math&amp;gt;\sup_{x\in[a,b]}\|f(x)-p(x)\|\le \epsilon&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
不失一般性地，可以仅考虑区间&amp;lt;math&amp;gt;[a,b]=[0,1]&amp;lt;/math&amp;gt;。因为对于定义在一般的实数区间&amp;lt;math&amp;gt;[a,b]&amp;lt;/math&amp;gt;上的任意函数&amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt;，可通过变量变换&amp;lt;math&amp;gt;t\mapsto a+(b-a)t&amp;lt;/math&amp;gt;将其转化为定义在&amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt;上的新函数&amp;lt;math&amp;gt;g(t)=f(a+(b-a)t)&amp;lt;/math&amp;gt;，这并不会改变函数的连续性以及是否为多项式。&lt;br /&gt;
&lt;br /&gt;
因此，可假设连续函数&amp;lt;math&amp;gt;f:[0,1]\to\mathbb{R}&amp;lt;/math&amp;gt;。令&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;为足够大的正整数，取值待定。&lt;br /&gt;
&lt;br /&gt;
对于任意 &amp;lt;math&amp;gt;x\in [0,1]&amp;lt;/math&amp;gt;，令 &amp;lt;math&amp;gt;Y_x\sim\text{Bin}\left(n,x\right)&amp;lt;/math&amp;gt; 为以&amp;lt;math&amp;gt;n,x&amp;lt;/math&amp;gt;为参数的二项分布随机变量。&lt;br /&gt;
&lt;br /&gt;
将多项式 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 定义如下。对每个 &amp;lt;math&amp;gt;x\in[0,1]&amp;lt;/math&amp;gt;，令：&lt;br /&gt;
:&amp;lt;math&amp;gt;p(x)=\mathbb{E}\left[f\left(\frac{Y_x}{n}\right)\right]=\sum_{k=0}^nf\left(\frac{k}{n}\right)\Pr(Y_x=k)=\sum_{k=0}^nf\left(\frac{k}{n}\right){n\choose k}x^k(1-x)^{n-k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
容易看出，这是一个关于变量 &amp;lt;math&amp;gt;x\in[0,1]&amp;lt;/math&amp;gt; 的（&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;次）多项式。&lt;br /&gt;
&lt;br /&gt;
根据[https://en.wikipedia.org/wiki/Heine%E2%80%93Cantor_theorem &#039;&#039;&#039;一致连续性定理&#039;&#039;&#039;](&#039;&#039;&#039;海涅-康托尔定理&#039;&#039;&#039;)，紧空间 &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt; 上的连续函数 &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; 必然也是&#039;&#039;&#039;一致连续&#039;&#039;&#039;的。即，对任意 &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;，总存在 &amp;lt;math&amp;gt;\delta_\epsilon&amp;gt;0&amp;lt;/math&amp;gt;，使得对于 &amp;lt;math&amp;gt;[0,1]&amp;lt;/math&amp;gt; 中任意满足 &amp;lt;math&amp;gt;|x-y|\le\delta_\epsilon&amp;lt;/math&amp;gt; 的 &amp;lt;math&amp;gt;x,y&amp;lt;/math&amp;gt;，都有 &amp;lt;math&amp;gt;|f(x)-f(y)|\le\frac{\epsilon}{2}&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
不妨定下任意的&amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;，以及一致连续性定理因此保证的 &amp;lt;math&amp;gt;\delta_\epsilon&amp;gt;0&amp;lt;/math&amp;gt;。&lt;br /&gt;
并且定下任意的 &amp;lt;math&amp;gt;x\in[0,1]&amp;lt;/math&amp;gt;。我们希望验证 &amp;lt;math&amp;gt;|p(x)-f(x)|\le\epsilon&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
根据 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 的定义，有：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
|p(x)-f(x)|&lt;br /&gt;
=&amp;amp;&lt;br /&gt;
\left|\mathbb{E}\left[f\left(\frac{Y_x}{n}\right)\right]-f(x)\right|\\&lt;br /&gt;
=&amp;amp;&lt;br /&gt;
\left|\mathbb{E}\left[f\left(\frac{Y_x}{n}\right)-f(x)\right]\right| &amp;amp;&amp;amp; \text{(期望的线性)}\\&lt;br /&gt;
\le&amp;amp; &lt;br /&gt;
\mathbb{E}\left[\left|f\left(\frac{Y_x}{n}\right)-f(x)\right|\right] &amp;amp;&amp;amp; \text{(琴生不等式)}\\&lt;br /&gt;
=&amp;amp;&lt;br /&gt;
\mathbb{E}\left[\left|f\left(\frac{Y_x}{n}\right)-f(x)\right|\,\,\bigg{|}\,\, \left|\frac{Y_x}{n}-x\right|\le \delta_\epsilon\right]&lt;br /&gt;
\cdot&lt;br /&gt;
\Pr\left[\left|\frac{Y_x}{n}-x\right|\le \delta_\epsilon\right] \\&lt;br /&gt;
&amp;amp;+&lt;br /&gt;
\mathbb{E}\left[\left|f\left(\frac{Y_x}{n}\right)-f(x)\right|\,\,\bigg{|}\,\, \left|\frac{Y_x}{n}-x\right|&amp;gt; \delta_\epsilon\right]&lt;br /&gt;
\cdot&lt;br /&gt;
\Pr\left[\left|\frac{Y_x}{n}-x\right|&amp;gt; \delta_\epsilon\right] &amp;amp;&amp;amp; \text{(全期望法则)}\\&lt;br /&gt;
\le&amp;amp;&lt;br /&gt;
\frac{\epsilon}{2}&lt;br /&gt;
+2\|f\|_{\infty}\cdot \Pr\left[\left|\frac{Y_x}{n}-x\right|&amp;gt; \delta_\epsilon\right] &amp;amp;&amp;amp; (\star)&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
上述最后一个不等式 &amp;lt;math&amp;gt;(\star)&amp;lt;/math&amp;gt; 成立是因为根据一致连续性，条件 &amp;lt;math&amp;gt;\left|\frac{Y_x}{n}-x\right|\le \delta_\epsilon&amp;lt;/math&amp;gt; 保证了 &amp;lt;math&amp;gt;\left|f\left(\frac{Y_x}{n}\right)-f(x)\right|\le\frac{\epsilon}{2}&amp;lt;/math&amp;gt;，因此&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}\left[\left|f\left(\frac{Y_x}{n}\right)-f(x)\right|\,\,\bigg{|}\,\, \left|\frac{Y_x}{n}-x\right|\le \delta_\epsilon\right]\le\frac{\epsilon}{2}&amp;lt;/math&amp;gt;&lt;br /&gt;
另一方面，有 &amp;lt;math&amp;gt;\left|f\left({Y_x}/{n}\right)-f(x)\right|\le \sup_{y\in[0,1]}|f(y)-f(x)|\le 2\sup_{y\in[0,1]}|f(y)| = 2\|f\|_{\infty}&amp;lt;/math&amp;gt; 无条件成立，因此&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}\left[\left|f\left(\frac{Y_x}{n}\right)-f(x)\right|\,\,\bigg{|}\,\, \left|\frac{Y_x}{n}-x\right|&amp;gt; \delta_\epsilon\right]\le 2\|f\|_{\infty}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
公式 &amp;lt;math&amp;gt;(\star)&amp;lt;/math&amp;gt; 中的概率可由切比雪夫不等式得出如下上界：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left[\left|\frac{Y_x}{n}-x\right|&amp;gt; \delta_\epsilon\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[\left|{Y_x}-\mathbb{E}[Y_x]\right|&amp;gt; n\delta_\epsilon\right]  &amp;amp;&amp;amp; (\mathbb{E}[Y_x]=nx)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[Y_x]}{n^2\delta_\epsilon^2} &amp;amp;&amp;amp; \text{(切比雪夫不等式)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{nx(1-x)}{n^2\delta_\epsilon^2} &amp;amp;&amp;amp; (\mathbf{Var}[Y_x]=nx(1-x))\\&lt;br /&gt;
&amp;amp;\le &lt;br /&gt;
\frac{1}{4n\delta_\epsilon^2}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
因此可以选择 &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; 为任意满足 &amp;lt;math&amp;gt;n\ge\frac{\|f\|_{\infty}}{\epsilon\delta_\epsilon^2}&amp;lt;/math&amp;gt; 的正整数（注意到这一 &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; 的选取是与  &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; 无关的，因此可对所有  &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; 选取一致的  &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;）。如此可保证以上的概率上界为 &amp;lt;math&amp;gt;\frac{1}{4n\delta_\epsilon^2}\le\frac{\epsilon}{4\|f\|_{\infty}}&amp;lt;/math&amp;gt;。&lt;br /&gt;
将其代回至 &amp;lt;math&amp;gt;(\star)&amp;lt;/math&amp;gt;式，得到如下结论：&lt;br /&gt;
:&amp;lt;math&amp;gt;|p(x)-f(x)|\le \frac{\epsilon}{2}+2\|f\|_{\infty}\cdot \frac{\epsilon}{4\|f\|_{\infty}}\le\epsilon&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)/Threshold_of_k-clique_in_random_graph&amp;diff=13637</id>
		<title>概率论与数理统计 (Spring 2026)/Threshold of k-clique in random graph</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)/Threshold_of_k-clique_in_random_graph&amp;diff=13637"/>
		<updated>2026-04-16T08:12:47Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;在 Erdős-Rényi 随机图模型 &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt; 中，一个随机无向图 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 以如下的方式生成：图 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 包含 &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; 个顶点，每一对顶点之间都独立同地以概率 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 连一条无向边。如此生成的随机图记为 &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt;。  固定整数 &amp;lt;math&amp;gt;k\ge 3&amp;lt;/math&amp;gt;，考虑随机图 &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt; 包含 &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;（&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团，&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-clique）子图...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;在 Erdős-Rényi 随机图模型 &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt; 中，一个随机无向图 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 以如下的方式生成：图 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 包含 &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; 个顶点，每一对顶点之间都独立同地以概率 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 连一条无向边。如此生成的随机图记为 &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
固定整数 &amp;lt;math&amp;gt;k\ge 3&amp;lt;/math&amp;gt;，考虑随机图 &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt; 包含 &amp;lt;math&amp;gt;K_k&amp;lt;/math&amp;gt;（&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团，&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-clique）子图概率。&lt;br /&gt;
&lt;br /&gt;
特别地，对于任意固定（即，常数）的整数 &amp;lt;math&amp;gt;k\ge 3&amp;lt;/math&amp;gt;，我们将证明存在函数 &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt; 使得对所有足够大的 &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; 和随机图 &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt; 有&lt;br /&gt;
:&amp;lt;math&amp;gt;({\color{red}\star})\qquad&amp;lt;/math&amp;gt; &amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr\left(G\text{ 包含子图 }K_k\right)&lt;br /&gt;
=\begin{cases}&lt;br /&gt;
o(1) &amp;amp; \text{如果} p=o(p_k(n))\\&lt;br /&gt;
1-o(1) &amp;amp; \text{如果} p=\omega(p_k(n))&lt;br /&gt;
\end{cases}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
这描述了“包含 &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团“这一性质，在随机图 &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt; 上，展现出的一种所谓的&#039;&#039;&#039;阈值现象&#039;&#039;&#039;：随着参数 &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt; 从远小于 &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt; 增长到远大于 &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;，“包含 &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团“这一事件在随机图 &amp;lt;math&amp;gt;G(n,p)&amp;lt;/math&amp;gt; 上从渐进几乎从不（asymptotically almost never）发生变为渐进几乎一定（&#039;&#039;&#039;a.a.s.&#039;&#039;&#039;）发生。&lt;br /&gt;
&lt;br /&gt;
固定整数 &amp;lt;math&amp;gt;k\ge 3&amp;lt;/math&amp;gt;，假设&amp;lt;math&amp;gt;k=O(1)&amp;lt;/math&amp;gt;为常数。抽取随机图 &amp;lt;math&amp;gt;G\sim G(n,p)&amp;lt;/math&amp;gt;。令随机变量 &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; 表示 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 中包含的 &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团的个数，因此有：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left(G\text{ 包含子图 }K_k\right)&lt;br /&gt;
=&lt;br /&gt;
\Pr(X\ge 1)&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=一阶矩方法=&lt;br /&gt;
根据马尔可夫不等式：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr(X\ge 1)\le\mathbb{E}[X]&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
因此，通过计算 &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; 的期望（一阶矩），可得到概率 &amp;lt;math&amp;gt;\Pr\left(G\text{ 包含子图 }K_k\right)&amp;lt;/math&amp;gt; 的上界。&lt;br /&gt;
&lt;br /&gt;
对于顶点集合 &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt; 的任意 &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-子集 &amp;lt;math&amp;gt;S\in{[n]\choose k}&amp;lt;/math&amp;gt;，定义指示随机变量 &amp;lt;math&amp;gt;I_S=I(K_S\subseteq G)&amp;lt;/math&amp;gt;，即 &amp;lt;math&amp;gt;I_S=1&amp;lt;/math&amp;gt; 当且仅当点集 &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; 在随机图 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 中构成了一个 &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团。根据这个定义，有：&lt;br /&gt;
*&amp;lt;math&amp;gt;X=\sum_{S\in{[n]\choose k}}I_S&amp;lt;/math&amp;gt;&lt;br /&gt;
*&amp;lt;math&amp;gt;\mathbb{E}[I_S]=\Pr(K_S\subseteq G)=p^{{k\choose 2}}&amp;lt;/math&amp;gt;&lt;br /&gt;
于是根据期望的线性：&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}[X]=\sum_{S\in{[n]\choose k}}\mathbb{E}[I_S]={n\choose k}p^{{k\choose 2}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
对于常数 &amp;lt;math&amp;gt;k=O(1)&amp;lt;/math&amp;gt;，当 &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; 足够大时，&amp;lt;math&amp;gt;{n\choose k}p^{{k\choose 2}}=\Theta\left(n^kp^{k(k-1)/2}\right)&amp;lt;/math&amp;gt;。这提示我们将函数 &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt; 做如下定义：&lt;br /&gt;
:&amp;lt;math&amp;gt;p_k(n)=n^{-2/(k-1)}&amp;lt;/math&amp;gt;.&lt;br /&gt;
于是，容易验证有：&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}[X]=\Theta\left(n^kp^{k(k-1)/2}\right)=\begin{cases}&lt;br /&gt;
o(1) &amp;amp; \text{如果 } p=o(p_k(n))=o\left(n^{-2/(k-1)}\right)\\&lt;br /&gt;
\omega(1) &amp;amp; \text{如果 } p=\omega(p_k(n))=\omega\left(n^{-2/(k-1)}\right)&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
特别地，当 &amp;lt;math&amp;gt;p=o(p_k(n))=o\left(n^{-2/(k-1)}\right)&amp;lt;/math&amp;gt; 时，&amp;lt;math&amp;gt;\mathbb{E}[X]=o(1)&amp;lt;/math&amp;gt;。根据马尔可夫不等式：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left(G\text{ 包含子图 }K_k\right)=\Pr(X\ge 1)\le \mathbb{E}[X]=o(1).&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
因此，&amp;lt;math&amp;gt;({\color{red}\star})&amp;lt;/math&amp;gt; 中的概率上界已得到证明。&lt;br /&gt;
&lt;br /&gt;
事实上，这一结果可以更加直接地使用 union bound 证明：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr\left(G\text{ 包含子图 }K_k\right)=\mbox{$\Pr\left(\bigcup_{S\in{[n]\choose k}}(K_S\subseteq G)\right)$}\le {n\choose k}p^{{k\choose 2}}=o(1)\qquad\left(\text{如果 $p=o(p_k(n))$}\right)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
尽管如此，对于一阶矩的分析，有助于我们猜测出正确的阈值 &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;，而且对于更高阶矩的估计也是必要的。&lt;br /&gt;
&lt;br /&gt;
=二阶矩方法=&lt;br /&gt;
为了完全证明 &amp;lt;math&amp;gt;({\color{red}\star})&amp;lt;/math&amp;gt;，我们需要补全其中的概率下界，即在 &amp;lt;math&amp;gt;\mathbb{E}[X]=\omega(1)&amp;lt;/math&amp;gt; 的情况下，证明&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr(X\ge 1)\ge 1-o(1)&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
很遗憾，为实现这一目的，仅知道 &amp;lt;math&amp;gt;\mathbb{E}[X]=\omega(1)&amp;lt;/math&amp;gt; 是不够的。我们还需要知道关于随机变量 &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; 更多的信息，比方说它的方差（二阶中心矩）。&lt;br /&gt;
&lt;br /&gt;
根据切比雪夫不等式：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr(X=0)\le \Pr(|X-\mathbb{E}[X]|\ge \mathbb{E}[X])\le \frac{\mathbf{Var}[X]}{\mathbb{E}[X]^2}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
于是，上述目标归结为证明方差和期望之间有如下渐进关系： &amp;lt;math&amp;gt;\mathbf{Var}[X]=o\left(\mathbb{E}[X]^2\right)&amp;lt;/math&amp;gt;。实际上我们证明了更强的性质&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}[X^2]=o\left(\mathbb{E}[X]^2\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
这足以推出 &amp;lt;math&amp;gt;\mathbf{Var}[X]=\mathbb{E}[X^2]-\mathbb{E}[X]^2\le \mathbb{E}[X^2]=o\left(\mathbb{E}[X]^2\right)&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
还记得 &amp;lt;math&amp;gt;X=\sum_{S\in{[n]\choose k}}I_S&amp;lt;/math&amp;gt;，于是有：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbb{E}\left[X^2\right]&lt;br /&gt;
=&lt;br /&gt;
\mathbb{E}\left[\left(\sum_{S\in{[n]\choose k}}I_S\right)^2\right]&lt;br /&gt;
=&lt;br /&gt;
\sum_{S\in{[n]\choose k}}\mathbb{E}[I_S^2] + \sum_{\substack{S,T\in{[n]\choose k}\\S\neq T}}\mathbb{E}[I_SI_T]&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
我们将分别计算这两项。首先，由于随机变量 &amp;lt;math&amp;gt;I_S&amp;lt;/math&amp;gt; 取值为0或1，因此 &amp;lt;math&amp;gt;I_S^2=I_S&amp;lt;/math&amp;gt;，于是有：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\sum_{S\in{[n]\choose k}}\mathbb{E}[I_S^2]&lt;br /&gt;
=&lt;br /&gt;
\sum_{S\in{[n]\choose k}}\mathbb{E}[I_S]&lt;br /&gt;
=&lt;br /&gt;
\mathbb{E}[X]&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
另一方面，&amp;lt;math&amp;gt;I_SI_T=1&amp;lt;/math&amp;gt; 当且仅当 &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; 和 &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; 在随机图 &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; 中都是 &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-团，于是：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\sum_{\substack{S,T\in{[n]\choose k}\\S\neq T}}\mathbb{E}[I_SI_T]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\substack{S,T\in{[n]\choose k}\\S\neq T}}\Pr((K_S\cup K_T)\subseteq G)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\ell=0}^{k-1}\sum_{|S\cap T|=\ell}p^{2{k\choose 2}-{\ell\choose 2}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\sum_{\ell=0}^{k-1}{n\choose 2k-\ell} {2k-\ell\choose k}{k\choose \ell} p^{2{k\choose 2}-{\ell\choose 2}} &amp;amp;&amp;amp; \left(\text{因为 }{2k-\ell\choose k}{k\choose \ell}\le 2^{2k}=O(1)\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
O\left(n^{2k}p^{2{k\choose 2}}\cdot\sum_{\ell=0}^{k-1} n^{-\ell}p^{-{\ell\choose 2}}\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbb{E}[X]^2\cdot O\left(\sum_{\ell=0}^{k-1} n^{-\ell}p^{-{\ell\choose 2}}\right). &amp;amp;&amp;amp; \left(\text{因为 }\mathbb{E}[X]=\Theta\left(n^kp^{{k\choose 2}}\right)\right)&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
综上，我们有 &amp;lt;math&amp;gt;\mathbf{Var}[X]\le \mathbb{E}[X^2]\le \mathbb{E}[X]+\mathbb{E}[X]^2\cdot O\left(\sum_{\ell=0}^{k-1} n^{-\ell}p^{-{\ell\choose 2}}\right)&amp;lt;/math&amp;gt;，于是：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{\mathbf{Var}[X]}{\mathbb{E}[X]^2}&lt;br /&gt;
&amp;amp;\le &lt;br /&gt;
\frac{\mathbb{E}[X^2]}{\mathbb{E}[X]^2}\\&lt;br /&gt;
&amp;amp;\le &lt;br /&gt;
\frac{1}{\mathbb{E}[X]} + O\left(\sum_{\ell=0}^{k-1} n^{-\ell}p^{-{\ell\choose 2}}\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
O\left(\sum_{\ell=2}^{k}n^{-\ell}p^{-{\ell\choose 2}} \right)  &amp;amp;&amp;amp; \left(\text{因为 }\mathbb{E}[X]=\Theta\left(n^kp^{{k\choose 2}}\right)\right)\\&lt;br /&gt;
&amp;amp;=o(1). &amp;amp;&amp;amp; (\text{当 }p=\omega(p_k(n)))&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
正如我们之前分析的，根据切比雪夫不等式，这证明了当 &amp;lt;math&amp;gt;p=\omega(p_k(n))=\omega\left(n^{-2/(k-1)}\right)&amp;lt;/math&amp;gt; 时，有 &amp;lt;math&amp;gt;\Pr(X=0)=o(1)&amp;lt;/math&amp;gt;。和一阶矩方法的部分合起来，这证明了 &amp;lt;math&amp;gt;({\color{red}\star})&amp;lt;/math&amp;gt;。&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)&amp;diff=13636</id>
		<title>概率论与数理统计 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)&amp;diff=13636"/>
		<updated>2026-04-16T08:12:04Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lectures */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;&#039;&#039;&#039;概率论与数理统计&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;Probability Theory&#039;&#039;&#039; &amp;lt;br&amp;gt; &amp;amp; &#039;&#039;&#039;Mathematical Statistics&#039;&#039;&#039;&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = &#039;&#039;&#039;尹一通&#039;&#039;&#039;&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4  = office&lt;br /&gt;
|data4   = 计算机系 804&lt;br /&gt;
|header5 = &lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &#039;&#039;&#039;刘景铖&#039;&#039;&#039;&lt;br /&gt;
|header6 = &lt;br /&gt;
|label6  = Email&lt;br /&gt;
|data6   = liu@nju.edu.cn  &lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = office&lt;br /&gt;
|data7   = 计算机系 516&lt;br /&gt;
|header8 = Class&lt;br /&gt;
|label8  = &lt;br /&gt;
|data8   = &lt;br /&gt;
|header9 =&lt;br /&gt;
|label9  = Class meeting&lt;br /&gt;
|data9   = Wednesday, 9am-12am&amp;lt;br&amp;gt;&lt;br /&gt;
仙Ⅱ-212&lt;br /&gt;
|header10=&lt;br /&gt;
|label10 = Office hour&lt;br /&gt;
|data10  = TBA &amp;lt;br&amp;gt;计算机系 804（尹一通）&amp;lt;br&amp;gt;计算机系 516（刘景铖）&lt;br /&gt;
|header11= Textbook&lt;br /&gt;
|label11 = &lt;br /&gt;
|data11  = &lt;br /&gt;
|header12=&lt;br /&gt;
|label12 = &lt;br /&gt;
|data12  = [[File:概率导论.jpeg|border|100px]]&lt;br /&gt;
|header13=&lt;br /&gt;
|label13 = &lt;br /&gt;
|data13  = &#039;&#039;&#039;概率导论&#039;&#039;&#039;（第2版·修订版）&amp;lt;br&amp;gt; Dimitri P. Bertsekas and John N. Tsitsiklis&amp;lt;br&amp;gt; 郑忠国 童行伟 译；人民邮电出版社 (2022)&lt;br /&gt;
|header14=&lt;br /&gt;
|label14 = &lt;br /&gt;
|data14  = [[File:Grimmett_probability.jpg|border|100px]]&lt;br /&gt;
|header15=&lt;br /&gt;
|label15 = &lt;br /&gt;
|data15  = &#039;&#039;&#039;Probability and Random Processes&#039;&#039;&#039; (4E) &amp;lt;br&amp;gt; Geoffrey Grimmett and David Stirzaker &amp;lt;br&amp;gt;  Oxford University Press (2020)&lt;br /&gt;
|header16=&lt;br /&gt;
|label16 = &lt;br /&gt;
|data16  = [[File:Probability_and_Computing_2ed.jpg|border|100px]]&lt;br /&gt;
|header17=&lt;br /&gt;
|label17 = &lt;br /&gt;
|data17  = &#039;&#039;&#039;Probability and Computing&#039;&#039;&#039; (2E) &amp;lt;br&amp;gt; Michael Mitzenmacher and Eli Upfal &amp;lt;br&amp;gt;   Cambridge University Press (2017)&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Probability Theory and Mathematical Statistics&#039;&#039; (概率论与数理统计) class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* TBA&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: &lt;br /&gt;
:* [http://tcs.nju.edu.cn/yinyt/ 尹一通]：[mailto:yinyt@nju.edu.cn &amp;lt;yinyt@nju.edu.cn&amp;gt;]，计算机系 804 &lt;br /&gt;
:* [https://liuexp.github.io 刘景铖]：[mailto:liu@nju.edu.cn &amp;lt;liu@nju.edu.cn&amp;gt;]，计算机系 516 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 鞠哲：[mailto:juzhe@smail.nju.edu.cn &amp;lt;juzhe@smail.nju.edu.cn&amp;gt;]，计算机系 426&lt;br /&gt;
** 祝永祺：[mailto:652025330045@smail.nju.edu.cn &amp;lt;652025330045@smail.nju.edu.cn&amp;gt;]，计算机系 426&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;:&lt;br /&gt;
** 周三：9am-12am，仙Ⅱ-212&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: &lt;br /&gt;
:* TBA, 计算机系 804（尹一通）&lt;br /&gt;
:* TBA, 计算机系 516（刘景铖）&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090092561（申请加入需提供姓名、院系、学号）&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
课程内容分为三大部分：&lt;br /&gt;
* &#039;&#039;&#039;经典概率论&#039;&#039;&#039;：概率空间、随机变量及其数字特征、多维与连续随机变量、极限定理等内容&lt;br /&gt;
* &#039;&#039;&#039;概率与计算&#039;&#039;&#039;：测度集中现象 (concentration of measure)、概率法 (the probabilistic method)、离散随机过程的相关专题&lt;br /&gt;
* &#039;&#039;&#039;数理统计&#039;&#039;&#039;：参数估计、假设检验、贝叶斯估计、线性回归等统计推断等概念&lt;br /&gt;
&lt;br /&gt;
对于第一和第二部分，要求清楚掌握基本概念，深刻理解关键的现象与规律以及背后的原理，并可以灵活运用所学方法求解相关问题。对于第三部分，要求熟悉数理统计的若干基本概念，以及典型的统计模型和统计推断问题。&lt;br /&gt;
&lt;br /&gt;
经过本课程的训练，力求使学生能够熟悉掌握概率的语言，并会利用概率思维来理解客观世界并对其建模，以及驾驭概率的数学工具来分析和求解专业问题。&lt;br /&gt;
&lt;br /&gt;
=== 教材与参考书 Course Materials ===&lt;br /&gt;
* &#039;&#039;&#039;[BT]&#039;&#039;&#039; 概率导论（第2版·修订版），[美]伯特瑟卡斯（Dimitri P.Bertsekas）[美]齐齐克利斯（John N.Tsitsiklis）著，郑忠国 童行伟 译，人民邮电出版社（2022）。&lt;br /&gt;
* &#039;&#039;&#039;[GS]&#039;&#039;&#039; &#039;&#039;Probability and Random Processes&#039;&#039;, by Geoffrey Grimmett and David Stirzaker; Oxford University Press; 4th edition (2020).&lt;br /&gt;
* &#039;&#039;&#039;[MU]&#039;&#039;&#039; &#039;&#039;Probability and Computing: Randomization and Probabilistic Techniques in Algorithms and Data Analysis&#039;&#039;, by Michael Mitzenmacher, Eli Upfal; Cambridge University Press; 2nd edition (2017).&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grading Policy ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩和期末考试成绩综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。符合规则的讨论与致谢将不会影响得分。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
*[[概率论与数理统计 (Spring 2026)/Problem Set 1|Problem Set 1]]  请在 2026/4/1 上课之前(9am UTC+8)提交到 [mailto:pr2026_nju@163.com pr2026_nju@163.com] (文件名为&#039;&amp;lt;font color=red &amp;gt;学号_姓名_A1.pdf&amp;lt;/font&amp;gt;&#039;).&lt;br /&gt;
** [[概率论与数理统计 (Spring 2026)/第一次作业提交名单|第一次作业提交名单]]&lt;br /&gt;
&lt;br /&gt;
*[[概率论与数理统计 (Spring 2026)/Problem Set 2|Problem Set 2]]  请在 2026/4/22 上课之前(9am UTC+8)提交到 [mailto:pr2026_nju@163.com pr2026_nju@163.com] (文件名为&#039;&amp;lt;font color=red &amp;gt;学号_姓名_A2.pdf&amp;lt;/font&amp;gt;&#039;).&lt;br /&gt;
&lt;br /&gt;
= Lectures =&lt;br /&gt;
# [http://tcs.nju.edu.cn/slides/prob2026/Intro.pdf 课程简介]&lt;br /&gt;
# [http://tcs.nju.edu.cn/slides/prob2026/ProbSpace.pdf 概率空间]&lt;br /&gt;
#* 阅读：&#039;&#039;&#039;[BT] 第1章&#039;&#039;&#039; 或 &#039;&#039;&#039;[GS] Chapter 1&#039;&#039;&#039;&lt;br /&gt;
#* [[概率论与数理统计 (Spring 2026)/Entropy and volume of Hamming balls|Entropy and volume of Hamming balls]]&lt;br /&gt;
#* [[概率论与数理统计 (Spring 2026)/Karger&#039;s min-cut algorithm| Karger&#039;s min-cut algorithm]]&lt;br /&gt;
# [http://tcs.nju.edu.cn/slides/prob2026/RandVar.pdf 随机变量]&lt;br /&gt;
#* 阅读：&#039;&#039;&#039;[BT] 第2章&#039;&#039;&#039; 或 &#039;&#039;&#039;[GS] Chapter 2, Sections 3.1~3.5, 3.7&#039;&#039;&#039;&lt;br /&gt;
#* 阅读：&#039;&#039;&#039;[MU] Chapter 2&#039;&#039;&#039;&lt;br /&gt;
#* [[概率论与数理统计 (Spring 2026)/Average-case analysis of QuickSort|Average-case analysis of &#039;&#039;&#039;&#039;&#039;QuickSort&#039;&#039;&#039;&#039;&#039;]]&lt;br /&gt;
# [http://tcs.nju.edu.cn/slides/prob2026/Deviation.pdf 矩与偏差]&lt;br /&gt;
#* 阅读：&#039;&#039;&#039;[MU] Chapter 3&#039;&#039;&#039;&lt;br /&gt;
#* 阅读：&#039;&#039;&#039;[BT] 章节 2.4, 4.2, 4.3, 5.1&#039;&#039;&#039; 或 &#039;&#039;&#039;[GS] Sections 3.3, 3.6, 7.3&#039;&#039;&#039;&lt;br /&gt;
#* [[概率论与数理统计 (Spring 2026)/Threshold of k-clique in random graph|Threshold of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-clique in random graph]]&lt;br /&gt;
#* [[概率论与数理统计 (Spring 2026)/Weierstrass Approximation Theorem|Weierstrass approximation]]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [https://plato.stanford.edu/entries/probability-interpret/ Interpretations of probability]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/History_of_probability History of probability]&lt;br /&gt;
* Example problems:&lt;br /&gt;
** [https://dornsifecms.usc.edu/assets/sites/520/docs/VonNeumann-ams12p36-38.pdf von Neumann&#039;s Bernoulli factory] and other [https://peteroupc.github.io/bernoulli.html Bernoulli factory algorithms]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Boy_or_Girl_paradox Boy or Girl paradox]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Monty_Hall_problem Monty Hall problem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Bertrand_paradox_(probability) Bertrand paradox]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Hard_spheres Hard spheres model] and [https://en.wikipedia.org/wiki/Ising_model Ising model]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/PageRank &#039;&#039;PageRank&#039;&#039;] and stationary [https://en.wikipedia.org/wiki/Random_walk random walk]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Diffusion_process Diffusion process] and [https://en.wikipedia.org/wiki/Diffusion_model diffusion model]&lt;br /&gt;
*[https://en.wikipedia.org/wiki/Probability_space Probability space]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sample_space Sample space]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Event_(probability_theory) Event] and [https://en.wikipedia.org/wiki/Σ-algebra &amp;lt;math&amp;gt;\sigma&amp;lt;/math&amp;gt;-algebra]&lt;br /&gt;
** Kolmogorov&#039;s [https://en.wikipedia.org/wiki/Probability_axioms axioms of probability]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Discrete_uniform_distribution Classical] and [https://en.wikipedia.org/wiki/Geometric_probability goemetric probability]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Boole%27s_inequality Union bound]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Inclusion%E2%80%93exclusion_principle Inclusion-Exclusion principle]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Boole%27s_inequality#Bonferroni_inequalities Bonferroni inequalities]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Conditional_probability Conditional probability]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Chain_rule_(probability) Chain rule]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Law_of_total_probability Law of total probability]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Bayes%27_theorem Bayes&#039; law]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Independence_(probability_theory) Independence] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Pairwise_independence Pairwise independence]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Random_variable Random variable]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cumulative_distribution_function Cumulative distribution function]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Probability_mass_function Probability mass function]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Probability_density_function Probability density function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Multivariate_random_variable Random vector]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Joint_probability_distribution Joint probability distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Conditional_probability_distribution Conditional probability distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Marginal_distribution Marginal distribution]&lt;br /&gt;
* Some &#039;&#039;&#039;discrete&#039;&#039;&#039; probability distributions&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Bernoulli_trial Bernoulli trial] and [https://en.wikipedia.org/wiki/Bernoulli_distribution Bernoulli distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Discrete_uniform_distribution Discrete uniform distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Binomial_distribution Binomial distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Geometric_distribution Geometric distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Negative_binomial_distribution Negative binomial distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Hypergeometric_distribution Hypergeometric distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Poisson_distribution Poisson distribution]&lt;br /&gt;
** and [https://en.wikipedia.org/wiki/List_of_probability_distributions#Discrete_distributions others]&lt;br /&gt;
* Balls into bins model&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Multinomial_distribution Multinomial distribution]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Birthday_problem Birthday problem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Coupon_collector%27s_problem Coupon collector]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Balls_into_bins_problem Occupancy problem]&lt;br /&gt;
* Random graphs&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi random graph model]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Galton%E2%80%93Watson_process Galton–Watson branching process]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Expected_value Expectation]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Law_of_the_unconscious_statistician Law of the unconscious statistician, &#039;&#039;LOTUS&#039;&#039;]&lt;br /&gt;
** [https://dlsun.github.io/probability/linearity.html Linearity of expectation]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Conditional_expectation Conditional expectation]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Law_of_total_expectation Law of total expectation]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Existence_problems&amp;diff=13623</id>
		<title>组合数学 (Fall 2026)/Existence problems</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Existence_problems&amp;diff=13623"/>
		<updated>2026-04-08T08:59:32Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== Existence by Counting == === Shannon&amp;#039;s circuit lower bound=== This is a fundamental problem in in Computer Science.  A &amp;#039;&amp;#039;&amp;#039;boolean function&amp;#039;&amp;#039;&amp;#039; is a function in the form &amp;lt;math&amp;gt;f:\{0,1\}^n\rightarrow \{0,1\}&amp;lt;/math&amp;gt;.  [http://en.wikipedia.org/wiki/Boolean_circuit Boolean circuit] is a mathematical model of computation. Formally, a boolean circuit is a directed acyclic graph. Nodes with indegree zero are input nodes, labeled &amp;lt;math&amp;gt;x_1, x_2, \ldots , x_n&amp;lt;/math&amp;gt;. A circuit h...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Existence by Counting ==&lt;br /&gt;
=== Shannon&#039;s circuit lower bound===&lt;br /&gt;
This is a fundamental problem in in Computer Science.&lt;br /&gt;
&lt;br /&gt;
A &#039;&#039;&#039;boolean function&#039;&#039;&#039; is a function in the form &amp;lt;math&amp;gt;f:\{0,1\}^n\rightarrow \{0,1\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
[http://en.wikipedia.org/wiki/Boolean_circuit Boolean circuit] is a mathematical model of computation.&lt;br /&gt;
Formally, a boolean circuit is a directed acyclic graph. Nodes with indegree zero are input nodes, labeled &amp;lt;math&amp;gt;x_1, x_2, \ldots , x_n&amp;lt;/math&amp;gt;. A circuit has a unique node with outdegree zero, called the output node. Every other node is a gate. There are three types of gates: AND, OR (both with indegree two), and NOT (with indegree one).&lt;br /&gt;
&lt;br /&gt;
Computations in Turing machines can be simulated by circuits, and any boolean function in &#039;&#039;&#039;P&#039;&#039;&#039; can be computed by a circuit with polynomially many gates. Thus, if we can find a function in &#039;&#039;&#039;NP&#039;&#039;&#039; that cannot be computed by any circuit with polynomially many gates, then &#039;&#039;&#039;NP&#039;&#039;&#039;&amp;lt;math&amp;gt;\neq&amp;lt;/math&amp;gt;&#039;&#039;&#039;P&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
The following theorem due to Shannon says that functions with exponentially large circuit complexity do exist.&lt;br /&gt;
&lt;br /&gt;
{{Theorem&lt;br /&gt;
|Theorem (Shannon 1949)|&lt;br /&gt;
:There is a boolean function &amp;lt;math&amp;gt;f:\{0,1\}^n\rightarrow \{0,1\}&amp;lt;/math&amp;gt; with circuit complexity greater than &amp;lt;math&amp;gt;\frac{2^n}{3n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof| &lt;br /&gt;
We first count the number of boolean functions &amp;lt;math&amp;gt;f:\{0,1\}^n\rightarrow \{0,1\}&amp;lt;/math&amp;gt;. There are &amp;lt;math&amp;gt;2^{2^n}&amp;lt;/math&amp;gt; boolean functions &amp;lt;math&amp;gt;f:\{0,1\}^n\rightarrow \{0,1\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Then we count the number of boolean circuit with fixed number of gates.&lt;br /&gt;
Fix an integer &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt;, we count the number of circuits with &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; gates. By the [http://en.wikipedia.org/wiki/De_Morgan&#039;s_laws De Morgan&#039;s laws], we can assume that all NOTs are pushed back to the inputs. Each gate has one of the two types (AND or OR), and has two inputs. Each of the inputs to a gate is either a constant 0 or 1, an input variable &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt;, an inverted input variable &amp;lt;math&amp;gt;\neg x_i&amp;lt;/math&amp;gt;, or the output of another gate; thus, there are at most &amp;lt;math&amp;gt;2+2n+t-1&amp;lt;/math&amp;gt; possible gate inputs. It follows that the number of circuits with &amp;lt;math&amp;gt;t&amp;lt;/math&amp;gt; gates is at most &amp;lt;math&amp;gt;2^t(t+2n+1)^{2t}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;t=2^n/3n&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\frac{2^t(t+2n+1)^{2t}}{2^{2^n}}=o(1)&amp;lt;1,&amp;lt;/math&amp;gt;      thus, &amp;lt;math&amp;gt;2^t(t+2n+1)^{2t} &amp;lt; 2^{2^n}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Each boolean circuit computes one boolean function. Therefore, there must exist a boolean function &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; which cannot be computed by any circuits with &amp;lt;math&amp;gt;2^n/3n&amp;lt;/math&amp;gt; gates.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Note that by Shannon&#039;s theorem, not only there exists a boolean function with exponentially large circuit complexity, but &#039;&#039;almost all&#039;&#039; boolean functions have exponentially large circuit complexity.&lt;br /&gt;
&lt;br /&gt;
=== Double counting ===&lt;br /&gt;
The double counting principle states the following obvious fact: if the elements of a set are counted in two different ways, the answers are the same.&lt;br /&gt;
==== Handshaking lemma ====&lt;br /&gt;
The following lemma is a standard demonstration of double counting.&lt;br /&gt;
{{Theorem|Handshaking Lemma|&lt;br /&gt;
:At a party, the number of guests who shake hands an odd number of times is even.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
We model this scenario as an undirected graph &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; standing for the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; guests. There is an edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; shake hands. Let &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; be the degree of vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, which represents the number of times that &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; shakes hand. The handshaking lemma states that in any undirected graph, the number of vertices whose degrees are odd is even. It is sufficient to show that the sum of odd degrees is even.&lt;br /&gt;
&lt;br /&gt;
The handshaking lemma is a direct consequence of the following lemma, which is proved by Euler in his 1736 paper on [http://en.wikipedia.org/wiki/Seven_Bridges_of_K%C3%B6nigsberg Seven Bridges of Königsberg] that began the study of graph theory.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma (Euler 1736)|&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{v\in V}d(v)=2|E|&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We count the number of &#039;&#039;&#039;directed&#039;&#039;&#039; edges. A directed edge is an ordered pair &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\{u,v\}\in E&amp;lt;/math&amp;gt;. There are two ways to count the directed edges.&lt;br /&gt;
&lt;br /&gt;
First, we can enumerate by edges. Pick every edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; and apply two directions &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;(v,u)&amp;lt;/math&amp;gt; to the edge. This gives us &amp;lt;math&amp;gt;2|E|&amp;lt;/math&amp;gt; directed edges.&lt;br /&gt;
&lt;br /&gt;
On the other hand, we can enumerate by vertices. Pick every vertex &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt; and for each of its &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; neighbors, say &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;, generate a directed edge &amp;lt;math&amp;gt;(v,u)&amp;lt;/math&amp;gt;. This gives us &amp;lt;math&amp;gt;\sum_{v\in V}d(v)&amp;lt;/math&amp;gt; directed edges.&lt;br /&gt;
&lt;br /&gt;
It is obvious that the two terms are equal, since we just count the same thing twice with different methods. The lemma follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The handshaking lemma is implied directly by the above lemma, since the sum of even degrees is even.&lt;br /&gt;
==== Sperner&#039;s lemma ====&lt;br /&gt;
A &#039;&#039;&#039;triangulation&#039;&#039;&#039; of a triangle &amp;lt;math&amp;gt;abc&amp;lt;/math&amp;gt; is a decomposition of &amp;lt;math&amp;gt;abc&amp;lt;/math&amp;gt; to small triangles (called &#039;&#039;cells&#039;&#039;), such that any two different cells are either disjoint, or share an edge, or a vertex.&lt;br /&gt;
&lt;br /&gt;
A &#039;&#039;&#039;proper coloring&#039;&#039;&#039; of a triangulation of triangle &amp;lt;math&amp;gt;abc&amp;lt;/math&amp;gt; is a coloring of all vertices in the triangulation with three colors: &amp;lt;font color=red&amp;gt;red&amp;lt;/font&amp;gt;, &amp;lt;font color=blue&amp;gt;blue&amp;lt;/font&amp;gt;, and &amp;lt;font color=green&amp;gt;green&amp;lt;/font&amp;gt;, such that the following constraints are satisfied:&lt;br /&gt;
* The three vertices &amp;lt;math&amp;gt;a,b&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;c&amp;lt;/math&amp;gt; of the big triangle receive all three colors.&lt;br /&gt;
* The vertices in each of the three lines &amp;lt;math&amp;gt;ab&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;bc&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;ac&amp;lt;/math&amp;gt; receive two colors. &lt;br /&gt;
&lt;br /&gt;
The following figure is an example of a properly colored triangulation.&lt;br /&gt;
[[Image:sperner-triangle.png|260px|center]]&lt;br /&gt;
&lt;br /&gt;
In 1928 young Emanuel Sperner gave a combinatorial proof of the famous Brouwer&#039;s fixed point theorem by proving the following lemma (now called Sperner&#039;s lemma), with an extremely elegant proof.&lt;br /&gt;
{{Theorem|Sperner&#039;s Lemma (1928)|&lt;br /&gt;
:For any properly colored triangulation, there exists a cell receiving all three colors.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
The proof is done by appropriately constructing a dual graph of the triangulation. &lt;br /&gt;
&lt;br /&gt;
The dual graph is defined as follows:&lt;br /&gt;
* Each cell in the triangulation corresponds to a distinct vertex in the dual graph.&lt;br /&gt;
* The outer space corresponds to a distinct vertex in the dual graph.&lt;br /&gt;
* An edge is added between two vertices in the dual graph if the corresponding cells share a &amp;lt;math&amp;gt;{\color{Red}\mbox{red}}\mbox{--}{\color{Blue}\mbox{blue}}&amp;lt;/math&amp;gt; edge.&lt;br /&gt;
&lt;br /&gt;
The following is an example of the dual graph of a properly colored triangulation:&lt;br /&gt;
[[Image:sperner-dual.png|260px|center]]&lt;br /&gt;
&lt;br /&gt;
For vertices in the dual graph:&lt;br /&gt;
* If a cell receives all three colors, the corresponding vertex in the dual graph has degree 1;&lt;br /&gt;
* if a cell receives only &amp;lt;math&amp;gt;{\color{Red}\mbox{red}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;{\color{Blue}\mbox{blue}}&amp;lt;/math&amp;gt;, the corresponding vertex has degree 2;&lt;br /&gt;
* for all other cases (the cell is monochromatic, or does not have blue or red), the corresponding vertex has degree 0.&lt;br /&gt;
&lt;br /&gt;
Besides, the unique vertex corresponding to the outer space must have odd degree, since the number of &amp;lt;math&amp;gt;{\color{Red}\mbox{red}}\mbox{--}{\color{Blue}\mbox{blue}}&amp;lt;/math&amp;gt; transitions between a &amp;lt;math&amp;gt;{\color{Red}\mbox{red}}&amp;lt;/math&amp;gt; endpoint and a &amp;lt;math&amp;gt;{\color{Blue}\mbox{blue}}&amp;lt;/math&amp;gt; endpoint must be odd.&lt;br /&gt;
&lt;br /&gt;
By handshaking lemma, the number of odd-degree vertices in the dual graph is even, thus the number of cells receiving all three colors must be odd, which cannot be zero.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== The Pigeonhole Principle ==&lt;br /&gt;
The &#039;&#039;&#039;pigeonhole principle&#039;&#039;&#039; states the following &amp;quot;obvious&amp;quot; fact:&lt;br /&gt;
:&#039;&#039;&amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt; pigeons cannot sit in &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; holes so that every pigeon is alone in its hole.&#039;&#039;&lt;br /&gt;
This is one of the oldest &#039;&#039;&#039;non-constructive&#039;&#039;&#039; principles: it states only the &#039;&#039;existence&#039;&#039; of a pigeonhole with more than one pigeons and says nothing about how to &#039;&#039;find&#039;&#039; such a pigeonhole.&lt;br /&gt;
&lt;br /&gt;
The general form of pigeonhole principle, also known as the &#039;&#039;&#039;averaging principle&#039;&#039;&#039;, is stated as follows.&lt;br /&gt;
{{Theorem|Generalized pigeonhole principle|&lt;br /&gt;
:If a set consisting of more than &amp;lt;math&amp;gt;mn&amp;lt;/math&amp;gt; objects is partitioned into &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; classes, then some class receives more than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; objects.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Inevitable divisors ===&lt;br /&gt;
The following is one of Erdős&#039; favorite initiation questions to mathematics. The proof uses the Pigeonhole Principle.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:For any subset &amp;lt;math&amp;gt;S\subseteq\{1,2,\ldots,2n\}&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;|S|&amp;gt;n\,&amp;lt;/math&amp;gt;, there are two numbers &amp;lt;math&amp;gt;a,b\in S&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;a|b\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
For every odd number &amp;lt;math&amp;gt;m\in\{1,2,\ldots,2n\}&amp;lt;/math&amp;gt;, let &lt;br /&gt;
:&amp;lt;math&amp;gt;C_m=\{2^km\mid k\ge 0, 2^km\le 2n\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is easy to see that for any &amp;lt;math&amp;gt;b&amp;lt;a&amp;lt;/math&amp;gt; from the same &amp;lt;math&amp;gt;C_m&amp;lt;/math&amp;gt;, it holds that &amp;lt;math&amp;gt;a|b&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Every number &amp;lt;math&amp;gt;a\in S&amp;lt;/math&amp;gt; can be uniquely represented as &amp;lt;math&amp;gt;a=2^km&amp;lt;/math&amp;gt; for some odd number &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;, thus belongs to exactly one of &amp;lt;math&amp;gt;C_m&amp;lt;/math&amp;gt;, for odd &amp;lt;math&amp;gt;m\in\{1,2,\ldots, 2n\}&amp;lt;/math&amp;gt;.  There are &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; odd numbers in &amp;lt;math&amp;gt;\{1,2,\ldots,2n\}&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; different &amp;lt;math&amp;gt;C_m&amp;lt;/math&amp;gt;, but &amp;lt;math&amp;gt;|S|&amp;gt;n&amp;lt;/math&amp;gt;, thus there must exist distinct &amp;lt;math&amp;gt;a,b\in S&amp;lt;/math&amp;gt;, supposed that &amp;lt;math&amp;gt;b&amp;lt;a&amp;lt;/math&amp;gt;, belonging to the same &amp;lt;math&amp;gt;C_m&amp;lt;/math&amp;gt;, which implies that &amp;lt;math&amp;gt;a|b&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Monotonic subsequences ===&lt;br /&gt;
Let &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; be a sequence of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct real numbers. A &#039;&#039;&#039;subsequence&#039;&#039;&#039; is a sequence of distinct terms of &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; appearing in the same order in which they appear in &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt;. Formally, a subsequence of &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;(a_{i_1},a_{i_2},\ldots,a_{i_k})&amp;lt;/math&amp;gt;, with &amp;lt;math&amp;gt;i_1&amp;lt;i_2&amp;lt;\cdots&amp;lt;i_k&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A sequence &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_n)&amp;lt;/math&amp;gt; is &#039;&#039;&#039;increasing&#039;&#039;&#039; if &amp;lt;math&amp;gt;a_1&amp;lt;a_2&amp;lt;\cdots&amp;lt;a_n&amp;lt;/math&amp;gt;, and &#039;&#039;&#039;decreasing&#039;&#039;&#039; if &amp;lt;math&amp;gt;a_1&amp;gt;a_2&amp;gt;\cdots&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We are interested in the &#039;&#039;longest&#039;&#039; increasing and decreasing subsequences of an &amp;lt;math&amp;gt;a_1&amp;lt;a_2&amp;lt;\cdots&amp;lt;a_n&amp;lt;/math&amp;gt;. It is intuitive that the length of both the longest increasing subsequence and the longest decreasing subsequence cannot be small simultaneously. A famous result of Erdős and Szekeres formally justifies this intuition. This is one of the first results in extremal combinatorics, published in the influential 1935 paper of Erdős and Szekeres.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Erdős-Szekeres 1935)|&lt;br /&gt;
:A sequence of more than &amp;lt;math&amp;gt;mn&amp;lt;/math&amp;gt; different real numbers must contain either an increasing subsequence of length &amp;lt;math&amp;gt;m+1&amp;lt;/math&amp;gt;, or a decreasing subsequence of length &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|(due to Seidenberg 1959)&lt;br /&gt;
Let &amp;lt;math&amp;gt;(a_1,a_2,\ldots,a_{N})&amp;lt;/math&amp;gt; be the original sequence of &amp;lt;math&amp;gt;N&amp;gt;mn&amp;lt;/math&amp;gt; distinct real numbers. Associate each &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; a pair &amp;lt;math&amp;gt;(x_i,y_i)&amp;lt;/math&amp;gt;, defined as:&lt;br /&gt;
*&amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt;: the length of the longest &#039;&#039;increasing&#039;&#039; subsequence &#039;&#039;ending&#039;&#039; at &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;;&lt;br /&gt;
*&amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt;: the length of the longest &#039;&#039;decreasing&#039;&#039; subsequence &#039;&#039;starting&#039;&#039; at &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
A key observation is that &amp;lt;math&amp;gt;(x_i,y_i)\neq (x_j,y_j)&amp;lt;/math&amp;gt; whenever &amp;lt;math&amp;gt;i\neq j&amp;lt;/math&amp;gt;. This is proved as follows:&lt;br /&gt;
: &#039;&#039;&#039;Case 1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;a_i&amp;lt;a_j&amp;lt;/math&amp;gt;, then the longest increasing subsequence ending at &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; can be extended by adding on &amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;, so &amp;lt;math&amp;gt;x_i&amp;lt;x_j&amp;lt;/math&amp;gt;.&lt;br /&gt;
: &#039;&#039;&#039;Case 2:&#039;&#039;&#039;  If &amp;lt;math&amp;gt;a_i&amp;gt;a_j&amp;lt;/math&amp;gt;, then the longest decreasing subsequence starting at &amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt; can be preceded by &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;, so &amp;lt;math&amp;gt;y_i&amp;gt;y_j&amp;lt;/math&amp;gt;.&lt;br /&gt;
Now we put &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; &amp;quot;pigeons&amp;quot; &amp;lt;math&amp;gt;a_1,a_2,\ldots,a_N&amp;lt;/math&amp;gt; into &amp;quot;pigeonholes&amp;quot; &amp;lt;math&amp;gt;\{1,2,\ldots,N\}\times\{1,2,\ldots,N\}&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; is put into hole &amp;lt;math&amp;gt;(x_i,y_i)&amp;lt;/math&amp;gt;, with at most one pigeon per each hole (since different &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; has different &amp;lt;math&amp;gt;(x_i,y_i)&amp;lt;/math&amp;gt;). &lt;br /&gt;
&lt;br /&gt;
The number of pigeons is &amp;lt;math&amp;gt;N&amp;gt;mn&amp;lt;/math&amp;gt;. Due to pigeonhole principle, there must be a pigeon which is outside the region &amp;lt;math&amp;gt;\{1,2,\ldots,m\}\times\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, which implies that there exists an &amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt; with either &amp;lt;math&amp;gt;x_i&amp;gt;m&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;y_i&amp;gt;n&amp;lt;/math&amp;gt;. Due to our definition of &amp;lt;math&amp;gt;(x_i,y_i)&amp;lt;/math&amp;gt;, there must be either an increasing subsequence of length &amp;lt;math&amp;gt;m+1&amp;lt;/math&amp;gt;, or a decreasing subsequence of length &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Dirichlet&#039;s approximation ===&lt;br /&gt;
Let &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; be an irrational number. We now want to approximate &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; be a rational number (a fraction).&lt;br /&gt;
&lt;br /&gt;
Since every real interval &amp;lt;math&amp;gt;[a,b]&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;a&amp;lt;b&amp;lt;/math&amp;gt; contains infinitely many rational numbers, there must exist rational numbers arbitrarily close to &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;. The trick is to let the denominator of the fraction sufficiently large.&lt;br /&gt;
&lt;br /&gt;
Suppose however we restrict the rationals we may select to have denominators bounded by &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. How closely we can approximate &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; now?&lt;br /&gt;
&lt;br /&gt;
The following important theorem is due to Dirichlet and his &#039;&#039;Schubfachprinzip&#039;&#039; (&amp;quot;drawer principle&amp;quot;). The theorem is fundamental in numer theory and real analysis, but the proof is combinatorial.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Dirichlet 1879)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; be an irrational number. For any natural number &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, there is a rational number &amp;lt;math&amp;gt;\frac{p}{q}&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;1\le q\le n&amp;lt;/math&amp;gt; and &lt;br /&gt;
::&amp;lt;math&amp;gt;\left|x-\frac{p}{q}\right|&amp;lt;\frac{1}{nq}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;\{x\}=x-\lfloor x\rfloor&amp;lt;/math&amp;gt; denote the &#039;&#039;&#039;fractional part&#039;&#039;&#039; of the real number &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;. It is obvious that &amp;lt;math&amp;gt;\{x\}\in[0,1)&amp;lt;/math&amp;gt; for any real number &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Consider the &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt; numbers &amp;lt;math&amp;gt;\{kx\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;k=1,2,\ldots,n+1&amp;lt;/math&amp;gt;. These &amp;lt;math&amp;gt;n+1&amp;lt;/math&amp;gt; numbers (pigeons) belong to the following &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; intervals (pigeonholes):&lt;br /&gt;
:&amp;lt;math&amp;gt;\left(0,\frac{1}{n}\right),\left(\frac{1}{n},\frac{2}{n}\right),\ldots,\left(\frac{n-1}{n},1\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; is irrational, &amp;lt;math&amp;gt;\{kx\}&amp;lt;/math&amp;gt; cannot coincide with any endpoint of the above intervals.&lt;br /&gt;
&lt;br /&gt;
By the pigeonhole principle, there exist &amp;lt;math&amp;gt;1\le a&amp;lt;b\le n+1&amp;lt;/math&amp;gt;, such that &amp;lt;math&amp;gt;\{ax\},\{bx\}&amp;lt;/math&amp;gt; are in the same interval, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;|\{bx\}-\{ax\}|&amp;lt;\frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;|(b-a)x-\left(\lfloor bx\rfloor-\lfloor ax\rfloor\right)|&amp;lt;\frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let &amp;lt;math&amp;gt;q=b-a&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;p=\lfloor bx\rfloor-\lfloor ax\rfloor&amp;lt;/math&amp;gt;. We have &amp;lt;math&amp;gt;|qx-p|&amp;lt;\frac{1}{n}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;1\le q\le n&amp;lt;/math&amp;gt;. Dividing both sides by &amp;lt;math&amp;gt;q&amp;lt;/math&amp;gt;, the theorem is proved.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13622</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13622"/>
		<updated>2026-04-08T08:59:09Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Existence problems|Existence problems | 存在性问题]]&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13621</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13621"/>
		<updated>2026-04-08T08:54:40Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on entropy and counting ([http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf notes]) &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13620</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13620"/>
		<updated>2026-04-08T08:52:23Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# Guest lecture by Prof. Penghui Yao on [http://tcs.nju.edu.cn/slides/comb2026/entropy.pdf entropy and counting] &lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Cayley%27s_formula&amp;diff=13619</id>
		<title>组合数学 (Fall 2026)/Cayley&#039;s formula</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Fall_2026)/Cayley%27s_formula&amp;diff=13619"/>
		<updated>2026-04-08T08:45:55Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;== Cayley&amp;#039;s Formula == We now present a theorem of the number of labeled trees on a fixed number of vertices. It is due to [http://en.wikipedia.org/wiki/Arthur_Cayley Cayley] in 1889. The theorem is often referred by the name [http://en.wikipedia.org/wiki/Cayley&amp;#039;s_formula Cayley&amp;#039;s formula].  {{Theorem|Cayley&amp;#039;s formula for trees| : There are &amp;lt;math&amp;gt;n^{n-2}&amp;lt;/math&amp;gt; different trees on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct vertices. }}  The theorem has several proofs, including the bijectio...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Cayley&#039;s Formula ==&lt;br /&gt;
We now present a theorem of the number of labeled trees on a fixed number of vertices. It is due to [http://en.wikipedia.org/wiki/Arthur_Cayley Cayley] in 1889. The theorem is often referred by the name [http://en.wikipedia.org/wiki/Cayley&#039;s_formula Cayley&#039;s formula].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Cayley&#039;s formula for trees|&lt;br /&gt;
: There are &amp;lt;math&amp;gt;n^{n-2}&amp;lt;/math&amp;gt; different trees on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct vertices.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The theorem has several proofs, including the bijection which encodes a tree by a [http://en.wikipedia.org/wiki/Pr%C3%BCfer_sequence Prüfer code], through the [http://en.wikipedia.org/wiki/Kirchhoff&#039;s_matrix_tree_theorem Kirchhoff&#039;s matrix tree theorem], and by double counting.&lt;br /&gt;
&lt;br /&gt;
=== Proof of Cayley&#039;s formula by double counting ===&lt;br /&gt;
We now present a double counting proof, which is considered by the [http://en.wikipedia.org/wiki/Proofs_from_THE_BOOK Proofs from THE BOOK] &amp;quot;the most beautiful of them all&amp;quot;.&lt;br /&gt;
{{Prooftitle|Proof of Cayley&#039;s formula by double counting|&lt;br /&gt;
(Due to Pitman 1999)&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;T_n&amp;lt;/math&amp;gt; be the number of different trees defined on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct vertices.&lt;br /&gt;
&lt;br /&gt;
A &#039;&#039;&#039;rooted tree&#039;&#039;&#039; is a tree with a special vertex. That is, one of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices is marked as the &amp;quot;root&amp;quot; of the tree. A rooted tree defines a natural direction of all edges, such that an edge &amp;lt;math&amp;gt;uv&amp;lt;/math&amp;gt; of the tree is directed from &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; if &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; is before &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; along the unique path from the root.&lt;br /&gt;
&lt;br /&gt;
We count the number of different &#039;&#039;sequences&#039;&#039; of directed edges that can be added to an empty graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices to form from it a &#039;&#039;rooted&#039;&#039; tree. We note that such a sequence can be formed in two ways:&lt;br /&gt;
# Starting with an unrooted tree, choose one of its vertices as root, and fix an total order of edges to specify the order in which the edges are added.&lt;br /&gt;
# Starting from an empty graph, add the edges one by one in steps.&lt;br /&gt;
&lt;br /&gt;
In the first method, we pick one of the &amp;lt;math&amp;gt;T_n&amp;lt;/math&amp;gt; unrooted trees, choose one of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices as the root, and pick one of the &amp;lt;math&amp;gt;(n-1)!&amp;lt;/math&amp;gt; total orders of the &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; edges. This gives us &amp;lt;math&amp;gt;T_nn(n-1)!=T_nn!&amp;lt;/math&amp;gt; ways.&lt;br /&gt;
&lt;br /&gt;
In the second method, we consider the number of choices in one step, and multiply the numbers of choices in all steps. This is done as follows.&lt;br /&gt;
&lt;br /&gt;
Given a sequence of &#039;&#039;adding&#039;&#039; &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; edges to an empty graph to form a rooted tree, we reverse this sequence and get a sequence of &#039;&#039;removing&#039;&#039; edges one by one from the final rooted tree until no edge left. We observe that:&lt;br /&gt;
* At first, we remove an edge from the rooted tree. Suppose that the root of the tree is &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;, and the removed directed edge is &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt;.  After removing &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt;, the original rooted tree is disconnected into two rooted trees, one rooted at &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; and the other rooted at &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;.&lt;br /&gt;
* After removing &amp;lt;math&amp;gt;k-1&amp;lt;/math&amp;gt; edges, there are &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; rooted trees. In the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;th step, a directed edge &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; in the current forest is removed and the tree containing &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt; is disconnected into two trees, one rooted at the old root of that tree, and the other rooted at &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We now again reverse the above procedure, and consider the sequence of adding directed edges to an empty graph to form a rooted tree.&lt;br /&gt;
* At first, we have &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; rooted trees, each of 0 edge (&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; isolated vertices).&lt;br /&gt;
* After adding &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; edges, there are &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; rooted trees. Denoting the directed edge added next as &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt;. As observed above, &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; can be any one of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices; but &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; must be the root of one of the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; trees, except the tree which contains &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;. There are &amp;lt;math&amp;gt;n(k-1)&amp;lt;/math&amp;gt; choices of such &amp;lt;math&amp;gt;(u,v)&amp;lt;/math&amp;gt;.&lt;br /&gt;
Multiplying the numbers of choices in all steps, the number of sequences of adding directed edges to an empty graph to form a rooted tree is given by&lt;br /&gt;
:&amp;lt;math&amp;gt;\prod_{k=2}^nn(k-1)=n^{n-2}n!&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By the principle of double counting, counting the same thing by different methods yield the same result.&lt;br /&gt;
:&amp;lt;math&amp;gt;T_nn!=n^{n-2}n!&amp;lt;/math&amp;gt;,&lt;br /&gt;
which gives that &amp;lt;math&amp;gt;T_n=n^{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
==  Prüfer code ==&lt;br /&gt;
The Prüfer code encodes a labeled tree to a sequence of labels. This gives a bijections between trees and tuples.&lt;br /&gt;
&lt;br /&gt;
=== Encoding  ===&lt;br /&gt;
In a tree, the vertices of degree 1 are called leaves. It is easy to see that:&lt;br /&gt;
* each tree has at least two leaves; and&lt;br /&gt;
* after removing a leaf (along with the edge adjacent to it) from a tree, the resulting graph is still a tree. &lt;br /&gt;
&lt;br /&gt;
The following algorithm transforms a tree &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices &amp;lt;math&amp;gt;1,2,\ldots,n&amp;lt;/math&amp;gt;, to a tuple &amp;lt;math&amp;gt;(v_1,v_2,\ldots,v_{n-2})\in\{1,2,\ldots,n\}^{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Theorem| Prüfer code (encoder)|&lt;br /&gt;
:&#039;&#039;&#039;Input&#039;&#039;&#039;: A tree &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct vertices, labeled by &amp;lt;math&amp;gt;1,2,\ldots,n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&lt;br /&gt;
:let &amp;lt;math&amp;gt;T_1=T&amp;lt;/math&amp;gt;;&lt;br /&gt;
:for &amp;lt;math&amp;gt;i=1&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, do&lt;br /&gt;
::let &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt; be the leaf in &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; with the smallest label, and &amp;lt;math&amp;gt;v_i&amp;lt;/math&amp;gt; be its neighbor;&lt;br /&gt;
::let &amp;lt;math&amp;gt;T_{i+1}&amp;lt;/math&amp;gt; be the new tree obtained from deleting the leaf &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt;;&lt;br /&gt;
:end&lt;br /&gt;
:return &amp;lt;math&amp;gt;(v_1,v_2,\ldots,v_{n-2})&amp;lt;/math&amp;gt;;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Decoding ===&lt;br /&gt;
It is trivial to observe the following lemma:&lt;br /&gt;
{{Theorem|Lemma 1|&lt;br /&gt;
:For each &amp;lt;math&amp;gt;1\le i\le n-1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; is a tree of &amp;lt;math&amp;gt;n-i+1&amp;lt;/math&amp;gt; vertices. In particular, the vertices of &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; are  &amp;lt;math&amp;gt;u_i,u_{i+1},\ldots,u_{n-1},v_{n-1}&amp;lt;/math&amp;gt;, and the edges of &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; are precisely &amp;lt;math&amp;gt;\{u_j,v_j\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;i\le j\le n-1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
And there is a reason that we do not need to store &amp;lt;math&amp;gt;v_{n-1}&amp;lt;/math&amp;gt; in the Prüfer code.&lt;br /&gt;
{{Theorem|Lemma 2|&lt;br /&gt;
:It always holds that &amp;lt;math&amp;gt;v_{n-1}=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Every tree (of at least two vertices) has at least two leaves. The &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n-1&amp;lt;/math&amp;gt;, are the leaf of the smallest label in &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt;, which can never be &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is never deleted.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Lemma 1 and 2 together imply that given a Prüfer code &amp;lt;math&amp;gt;(v_1,v_2,\ldots,v_{n-2})&amp;lt;/math&amp;gt;, the only remaining task to reconstruct the tree &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; is to figure out those &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n-1&amp;lt;/math&amp;gt;. The following lemma state how to obtain &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n-1&amp;lt;/math&amp;gt;, from a Prüfer code &amp;lt;math&amp;gt;(v_1,v_2,\ldots,v_{n-2})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma 3|&lt;br /&gt;
:For &amp;lt;math&amp;gt;i=1,2,\ldots,n-1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt; is the smallest element of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; not in &amp;lt;math&amp;gt;\{u_1,\ldots,u_{i-1}\}\cup\{v_i,\ldots,v_{n-1}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Note that &amp;lt;math&amp;gt;u_1,u_2,\ldots,u_{n-1},v_{n-1}&amp;lt;/math&amp;gt; is a sequence of distinct vertices, because &amp;lt;math&amp;gt;u_1,u_2,\ldots,u_{n-1}&amp;lt;/math&amp;gt; are deleted one by one from the tree, and &amp;lt;math&amp;gt;v_{n-1}=n&amp;lt;/math&amp;gt; is never deleted. Thus, each vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; appears among &amp;lt;math&amp;gt;u_1,u_2,\ldots,u_{n-1},v_{n-1}&amp;lt;/math&amp;gt; exactly once. And each vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; appears for &amp;lt;math&amp;gt;deg(v)&amp;lt;/math&amp;gt; times among the edges &amp;lt;math&amp;gt;\{u_i,v_i\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;1\le i\le n-1&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;deg(v)&amp;lt;/math&amp;gt; denotes the degree of vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; in the original tree &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt;. Therefore, each vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; appears among  &amp;lt;math&amp;gt;v_1,v_2,\ldots,v_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;deg(v)-1&amp;lt;/math&amp;gt; times.&lt;br /&gt;
&lt;br /&gt;
Similarly, each vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; appears among &amp;lt;math&amp;gt;v_i,v_{i+1},\ldots,v_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;deg_i(v)-1&amp;lt;/math&amp;gt; times, where &amp;lt;math&amp;gt;deg_i(v)&amp;lt;/math&amp;gt; is the degree of vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt; in the tree &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt;. In particular, the leaves of &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; are not among &amp;lt;math&amp;gt;\{v_i,v_{i+1},\ldots,v_{n-2}\}&amp;lt;/math&amp;gt;. Recall that the vertices of &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; are &amp;lt;math&amp;gt;u_i,u_{i+1},\ldots,u_{n-1},v_{n-1}&amp;lt;/math&amp;gt;. Then the leaves of &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; are the elements of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; not in &amp;lt;math&amp;gt;\{u_1,\ldots,u_{i-1}\}\cup\{v_i,\ldots,v_{n-1}\}&amp;lt;/math&amp;gt;. By definition of Prüfer code, &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt; is the leaf in &amp;lt;math&amp;gt;T_i&amp;lt;/math&amp;gt; of smallest label, hence the smallest element of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; not in &amp;lt;math&amp;gt;\{u_1,\ldots,u_{i-1}\}\cup\{v_i,\ldots,v_{n-1}\}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Applying Lemma 3, we have the following decoder for the Prüfer code:&lt;br /&gt;
{{Theorem| Prüfer code (decoder)|&lt;br /&gt;
:&#039;&#039;&#039;Input&#039;&#039;&#039;: A tuple &amp;lt;math&amp;gt;(v_1,v_2,\ldots,v_{n-2})\in\{1,2,\ldots,n\}^{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:&lt;br /&gt;
:let &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt; be empty graph, and &amp;lt;math&amp;gt;v_{n-1}=n&amp;lt;/math&amp;gt;;&lt;br /&gt;
:for &amp;lt;math&amp;gt;i=1&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, do&lt;br /&gt;
::let &amp;lt;math&amp;gt;u_i&amp;lt;/math&amp;gt; be the smallest label not in &amp;lt;math&amp;gt;\{u_1,\ldots,u_{i-1}\}\cup\{v_i,\ldots,v_{n-1}\}&amp;lt;/math&amp;gt;;&lt;br /&gt;
::add an edge &amp;lt;math&amp;gt;\{u_i,v_i\}&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt;;&lt;br /&gt;
:end&lt;br /&gt;
:return &amp;lt;math&amp;gt;T&amp;lt;/math&amp;gt;; &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
In other words, the encoding of trees to tuples by the Prüfer code is reversible, thus the mapping is injective (1-1). To see it is also surjective, we need to show that for every possible &amp;lt;math&amp;gt;(v_1,v_2,\ldots,v_{n-2})\in\{1,2,\ldots,n\}^{n-2}&amp;lt;/math&amp;gt;, the above decoder recovers a tree from it. &lt;br /&gt;
&lt;br /&gt;
It is easy to see that the decoder always returns a graph of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; edges on the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The only thing remaining to verify is that the returned graph has no cycle in it, which can be easily proved by a timeline argument (left as an exercise).&lt;br /&gt;
&lt;br /&gt;
=== Bijection proof of Cayley&#039;s formula ===&lt;br /&gt;
Therefore, the Prüfer code establishes a bijection between the set of trees on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct vertices and the tuples from &amp;lt;math&amp;gt;\{1,2,\ldots,n\}^{n-2}&amp;lt;/math&amp;gt;. This proves Cayley&#039;s formula.&lt;br /&gt;
&lt;br /&gt;
== Kirchhoff&#039;s Matrix-Tree Theorem ==&lt;br /&gt;
Given an undirected graph &amp;lt;math&amp;gt;G([n],E)&amp;lt;/math&amp;gt;, the &#039;&#039;&#039;adjacency matrix&#039;&#039;&#039; &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; of graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;n\times n&amp;lt;/math&amp;gt; matrix such that&lt;br /&gt;
:&amp;lt;math&amp;gt;A(i,j)=\begin{cases}&lt;br /&gt;
1 &amp;amp; \{i,j\}\in E,\\&lt;br /&gt;
0 &amp;amp; \{i,j\}\not\in E.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Graph Laplacian===&lt;br /&gt;
Let &amp;lt;math&amp;gt;D&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;n\times n&amp;lt;/math&amp;gt; diagonal matrix such that&lt;br /&gt;
:&amp;lt;math&amp;gt;D(i,j)=\begin{cases}&lt;br /&gt;
\text{deg}(i) &amp;amp; i=j,\\&lt;br /&gt;
0 &amp;amp; i\neq j,&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\text{deg}(i)&amp;lt;/math&amp;gt; denotes the degree of vertex &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;Laplacian matrix&#039;&#039;&#039; &amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt; of graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is defined as &amp;lt;math&amp;gt;L=D-A&amp;lt;/math&amp;gt;, that is,&lt;br /&gt;
:&amp;lt;math&amp;gt;L(i,j)=\begin{cases}&lt;br /&gt;
\text{deg}(i) &amp;amp; i=j,\\&lt;br /&gt;
-1 &amp;amp; i\neq j\text{ and } \{i,j\}\in E,\\&lt;br /&gt;
0 &amp;amp; \text{otherwise}.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Suppose &amp;lt;math&amp;gt;G([n],E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; edges. The &#039;&#039;&#039;incidence matrix&#039;&#039;&#039; &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; of graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix such that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\forall i\in[n], \forall e\in E,\quad B(i,e)=\begin{cases}&lt;br /&gt;
1 &amp;amp; e=\{i,j\}\text{ and } i&amp;lt;j,\\&lt;br /&gt;
-1 &amp;amp; e=\{i,j\}\text{ and } i&amp;gt;j,\\&lt;br /&gt;
0 &amp;amp; \text{otherwise}.&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The following proposition is easy to verify.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:&amp;lt;math&amp;gt;L=BB^T&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
For any &amp;lt;math&amp;gt;i,j\in[n]&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;(BB^T)(i,j)=\sum_{e\in E}B(i,e)B^T(e,j)=\sum_{e\in E}B(i,e)B(j,e)&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is easy to verify that &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\sum_{e\in E}B(i,e)B(j,e)=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
\text{deg}(i) &amp;amp; i=j,\\&lt;br /&gt;
-1 &amp;amp; i\neq j\text{ and } \{i,j\}\in E,\\&lt;br /&gt;
0 &amp;amp; \text{otherwise},&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is equal to the definition of &amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== The Matrix-tree Theorem ===&lt;br /&gt;
The matrix-tree theorem of Kirchhoff states a striking fact: The number of spanning trees in any connected graph can be computed as the determinant of some appropriate graph matrix. &lt;br /&gt;
{{Theorem|Kirchhoff&#039;s Matrix-Tree Theorem|&lt;br /&gt;
:For any connected graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, the number of spanning trees in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;\det(L_{i,i})\,&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;L_{i,i}\,&amp;lt;/math&amp;gt; is the &amp;lt;math&amp;gt;(n-1)\times(n-1)&amp;lt;/math&amp;gt; matrix resulting from the Laplacian matrix &amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt; of graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; by deleting the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th row and the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th column.&lt;br /&gt;
}}&lt;br /&gt;
The determinant can be computed as fast as matrix multiplication, thus is quite efficient, especially when compared to our task: counting the number of subgraphs satisfying certain nontrivial global constraint (e.g. spanning tree). Such efficient algorithm is rarely seen for counting problems, which are usually #P-hard to compute (e.g. number of matchings in a graph).&lt;br /&gt;
&lt;br /&gt;
The key to prove the matrix-tree theorem is the Cauchy-Binet theorem in linear algebra, whose proof is beyond the scope of this class.&lt;br /&gt;
{{Theorem|Cauchy-Binet Theorem|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; be, respectively, &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;m\times n&amp;lt;/math&amp;gt; matrix. For any &amp;lt;math&amp;gt;S\subseteq [m]&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;A_{[n], S}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B_{S,[n]}&amp;lt;/math&amp;gt; denote, respectively, the &amp;lt;math&amp;gt;n\times n&amp;lt;/math&amp;gt; submatrices of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, consisting of the columns of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, or the rows of &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, indexed by elements of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
::&amp;lt;math&amp;gt;\det(AB)=\sum_{S\in{[m]\choose n}}\det(A_{[n],S})\det(B_{S,[n]}).&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; be the incidence matrix of graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Fix any vertex &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;(n-1)\times m&amp;lt;/math&amp;gt; matrix resulting from &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; by deleting the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th row. &lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;L=BB^T&amp;lt;/math&amp;gt;. It is easy to verify that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_{i,i}=CC^T.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Due to the Cauchy-Binet Theorem, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\det(L_{i,i})=\det(CC^T)&lt;br /&gt;
&amp;amp;=\sum_{S\in{[m]\choose n-1}}\det(C_{[n-1],S})\det(C^T_{S,[n-1]})\\&lt;br /&gt;
&amp;amp;=\sum_{S\in{[m]\choose n-1}}\det(C_{[n-1],S})^2.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The next lemma gives a key observation to prove the matrix-tree theorem.&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;S\subseteq [m]&amp;lt;/math&amp;gt; of size &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;, the value of &amp;lt;math&amp;gt;\det(C_{[n-1],S})&amp;lt;/math&amp;gt; is either 1, or -1, or 0. Moreover, &amp;lt;math&amp;gt;\det(C_{[n-1],S})=\pm1&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; indicates a spanning tree in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We first show &amp;lt;math&amp;gt;\det(C_{[n-1],S})\in\{0,1,-1\}&amp;lt;/math&amp;gt; by induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. &lt;br /&gt;
Note that &amp;lt;math&amp;gt;C_{[n-1],S}&amp;lt;/math&amp;gt; is an &amp;lt;math&amp;gt;(n-1)\times (n-1)&amp;lt;/math&amp;gt; matrix such that each column contains at most one 1 and at most one -1, and all other entries are 0. For such matrix, when &amp;lt;math&amp;gt;n-1=1&amp;lt;/math&amp;gt;, the induction hypothesis &amp;lt;math&amp;gt;\det(C_{[n-1],S})\in\{0,1,-1\}&amp;lt;/math&amp;gt; is trivially true. And for general &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, if every column has a 1 and a -1, then the sum of all rows is the zero vector, so the matrix is singular.  Otherwise, expand the determinant by a column with one nonzero entry to find it is &amp;lt;math&amp;gt;\pm1&amp;lt;/math&amp;gt; times the determinant of a smaller matrix of the same property, which by induction hypothesis has value 0 or &amp;lt;math&amp;gt;\pm1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We then show that &amp;lt;math&amp;gt;\det(C_{[n-1],S})&amp;lt;/math&amp;gt; is nonzero if and only if &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is a spanning tree. &lt;br /&gt;
&lt;br /&gt;
Suppose the &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; edges corresponding to &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is not a spanning tree of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; must have more than one components, and there must be a component &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; not containing vertex &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. The rows of &amp;lt;math&amp;gt;C_{[n-1],S}&amp;lt;/math&amp;gt; corresponding to &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; add to 0, thus these rows are linearly dependent, and hence &amp;lt;math&amp;gt;\det(C_{[n-1],S})=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Suppose the &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; edges corresponding to &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is a spanning tree of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Then there is a vertex &amp;lt;math&amp;gt;j_1\neq i&amp;lt;/math&amp;gt; of degree 1 in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;; let &amp;lt;math&amp;gt;e_1&amp;lt;/math&amp;gt; be the edge in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; incident to &amp;lt;math&amp;gt;j_1&amp;lt;/math&amp;gt;. Deleting vertex &amp;lt;math&amp;gt;j_1&amp;lt;/math&amp;gt; and edge &amp;lt;math&amp;gt;e_1&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, we obtain a tree of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; edges. Again there is a vertex &amp;lt;math&amp;gt;j_2\neq i&amp;lt;/math&amp;gt; of degree 1 with incident edge &amp;lt;math&amp;gt;e_2&amp;lt;/math&amp;gt;. Continue this we enumerate all vertices except &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;j_1,j_2,\ldots,j_{n-1}&amp;lt;/math&amp;gt; and all edges in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;e_1,e_2,\ldots,e_{n-1}&amp;lt;/math&amp;gt;. Now permute the rows and columns of &amp;lt;math&amp;gt;C_{[n-1],S}&amp;lt;/math&amp;gt; such that the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-th row corresponds to vertex &amp;lt;math&amp;gt;j_k&amp;lt;/math&amp;gt; and the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-th column corresponds to &amp;lt;math&amp;gt;e_k&amp;lt;/math&amp;gt;. By our construction the permuted &amp;lt;math&amp;gt;C_{[n-1],S}&amp;lt;/math&amp;gt; is lower triangle with &amp;lt;math&amp;gt;\pm1&amp;lt;/math&amp;gt; diagonal entries, since when each &amp;lt;math&amp;gt;j_k&amp;lt;/math&amp;gt; is removed it is of degree 1 in the remaining tree and is incident to edge &amp;lt;math&amp;gt;e_k&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;\det(C_{[n-1],S})=\pm1&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The matrix-tree theorem follows as consequence. Recall we show that&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\det(L_{i,i})&lt;br /&gt;
&amp;amp;=\sum_{S\in{[m]\choose n-1}}\det(C_{[n-1],S})^2.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
The sum enumerates over all subgraphs &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; edges, and by the above lemma &amp;lt;math&amp;gt;\det(C_{[n-1],S})^2=1&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is a spanning tree in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;. Therefore, &amp;lt;math&amp;gt;\det(L_{i,i})&amp;lt;/math&amp;gt; gives the number of spanning trees in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Cayley&#039;s formula by the matrix-tree theorem ===&lt;br /&gt;
The number of trees of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; distinct vertices equals the number of spanning trees in the complete graph &amp;lt;math&amp;gt;K_n&amp;lt;/math&amp;gt;. For &amp;lt;math&amp;gt;G=K_n&amp;lt;/math&amp;gt;, for any &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt; the &amp;lt;math&amp;gt;(n-1)\times (n-1)&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;L_{ii}&amp;lt;/math&amp;gt; is given by&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
L_{i,i}&lt;br /&gt;
=&lt;br /&gt;
\begin{bmatrix}&lt;br /&gt;
n-1 &amp;amp; -1 &amp;amp; \cdots &amp;amp; -1\\&lt;br /&gt;
-1 &amp;amp; n-1 &amp;amp; \cdots &amp;amp; -1\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; -1\\&lt;br /&gt;
-1 &amp;amp; -1 &amp;amp; \cdots &amp;amp; n-1&lt;br /&gt;
\end{bmatrix},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
whose determinant is &amp;lt;math&amp;gt;\det(L_{i,i})=n^{n-2}&amp;lt;/math&amp;gt;.&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13618</id>
		<title>组合数学 (Spring 2026)</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E7%BB%84%E5%90%88%E6%95%B0%E5%AD%A6_(Spring_2026)&amp;diff=13618"/>
		<updated>2026-04-08T08:40:11Z</updated>

		<summary type="html">&lt;p&gt;Etone: /* Lecture Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{Infobox&lt;br /&gt;
|name         = Infobox&lt;br /&gt;
|bodystyle    = &lt;br /&gt;
|title        = &amp;lt;font size=3&amp;gt;组合数学  &amp;lt;br&amp;gt;&lt;br /&gt;
Combinatorics&amp;lt;/font&amp;gt;&lt;br /&gt;
|titlestyle   = &lt;br /&gt;
&lt;br /&gt;
|image        = &lt;br /&gt;
|imagestyle   = &lt;br /&gt;
|caption      = &lt;br /&gt;
|captionstyle = &lt;br /&gt;
|headerstyle  = background:#ccf;&lt;br /&gt;
|labelstyle   = background:#ddf;&lt;br /&gt;
|datastyle    = &lt;br /&gt;
&lt;br /&gt;
|header1 =Instructor&lt;br /&gt;
|label1  = &lt;br /&gt;
|data1   = &lt;br /&gt;
|header2 = &lt;br /&gt;
|label2  = &lt;br /&gt;
|data2   = 尹一通&lt;br /&gt;
|header3 = &lt;br /&gt;
|label3  = Email&lt;br /&gt;
|data3   = yinyt@nju.edu.cn  &lt;br /&gt;
|header4 =&lt;br /&gt;
|label4= office&lt;br /&gt;
|data4= 计算机系 804&lt;br /&gt;
|header5 = Class&lt;br /&gt;
|label5  = &lt;br /&gt;
|data5   = &lt;br /&gt;
|header6 =&lt;br /&gt;
|label6  = Class meetings&lt;br /&gt;
|data6   = Wednesday, 2pm-4pm &amp;lt;br&amp;gt; 逸B-313&lt;br /&gt;
|header7 =&lt;br /&gt;
|label7  = Place&lt;br /&gt;
|data7   = &lt;br /&gt;
|header8 =&lt;br /&gt;
|label8  = Office hours&lt;br /&gt;
|data8   = Tuesday, 2-3pm &amp;lt;br&amp;gt;计算机系 804&lt;br /&gt;
|header9 = Textbook&lt;br /&gt;
|label9  = &lt;br /&gt;
|data9   = &lt;br /&gt;
|header10 =&lt;br /&gt;
|label10  = &lt;br /&gt;
|data10   = [[File:LW-combinatorics.jpeg|border|100px]]&lt;br /&gt;
|header11 =&lt;br /&gt;
|label11  = &lt;br /&gt;
|data11   = van Lint and Wilson. &amp;lt;br&amp;gt; &#039;&#039;A course in Combinatorics, 2nd ed.&#039;&#039;, &amp;lt;br&amp;gt; Cambridge Univ Press, 2001.&lt;br /&gt;
|header12 =&lt;br /&gt;
|label12  = &lt;br /&gt;
|data12   = [[File:Jukna_book.jpg|border|100px]]&lt;br /&gt;
|header13 =&lt;br /&gt;
|label13  = &lt;br /&gt;
|data13   = Jukna. &#039;&#039;Extremal Combinatorics: &amp;lt;br&amp;gt; With Applications in Computer Science,&amp;lt;br&amp;gt;2nd ed.&#039;&#039;, Springer, 2011.&lt;br /&gt;
|belowstyle = background:#ddf;&lt;br /&gt;
|below = &lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
This is the webpage for the &#039;&#039;Combinatorics&#039;&#039; class of Spring 2026. Students who take this class should check this page periodically for content updates and new announcements. &lt;br /&gt;
&lt;br /&gt;
= Announcement =&lt;br /&gt;
* &#039;&#039;&#039;(2026/03/25)&#039;&#039;&#039;&amp;lt;font color=red size=4&amp;gt; 第一次作业已发布&amp;lt;/font&amp;gt;，请在 2026/04/08 上课之前提交到 [mailto:njucomb26@163.com njucomb26@163.com] (文件名为&#039;学号_姓名_A1.pdf&#039;)&lt;br /&gt;
&lt;br /&gt;
= Course info =&lt;br /&gt;
* &#039;&#039;&#039;Instructor &#039;&#039;&#039;: 尹一通 ([http://tcs.nju.edu.cn/yinyt/ homepage])&lt;br /&gt;
:*&#039;&#039;&#039;email&#039;&#039;&#039;: yinyt@nju.edu.cn&lt;br /&gt;
:*&#039;&#039;&#039;office&#039;&#039;&#039;: 计算机系 804 &lt;br /&gt;
* &#039;&#039;&#039;Teaching assistant&#039;&#039;&#039;:&lt;br /&gt;
** 丁天行([mailto:652024330006@smail.nju.edu.cn 652024330006@smail.nju.edu.cn])&lt;br /&gt;
** 周灿&lt;br /&gt;
** 方子伊&lt;br /&gt;
* &#039;&#039;&#039;Class meeting&#039;&#039;&#039;: Wednesday, 2pm-4pm, 逸A-313.&lt;br /&gt;
* &#039;&#039;&#039;Office hour&#039;&#039;&#039;: TBA&lt;br /&gt;
:* &#039;&#039;&#039;QQ群&#039;&#039;&#039;: 1090691552 (加入时需报姓名、专业、学号)&lt;br /&gt;
&lt;br /&gt;
= Syllabus =&lt;br /&gt;
&lt;br /&gt;
=== 先修课程 Prerequisites ===&lt;br /&gt;
* 离散数学（Discrete Mathematics）&lt;br /&gt;
* 线性代数（Linear Algebra）&lt;br /&gt;
* 概率论（Probability Theory）&lt;br /&gt;
&lt;br /&gt;
=== Course materials ===&lt;br /&gt;
* [[组合数学 (Spring 2025)/Course materials|&amp;lt;font size=3&amp;gt;教材和参考书清单&amp;lt;/font&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
=== 成绩 Grades ===&lt;br /&gt;
* 课程成绩：本课程将会有若干次作业和一次期末考试。最终成绩将由平时作业成绩 (≥ 60%) 和期末考试成绩 (≤ 40%) 综合得出。&lt;br /&gt;
* 迟交：如果有特殊的理由，无法按时完成作业，请提前联系授课老师，给出正当理由。否则迟交的作业将不被接受。&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;font color=red&amp;gt; 学术诚信 Academic Integrity &amp;lt;/font&amp;gt;===&lt;br /&gt;
学术诚信是所有从事学术活动的学生和学者最基本的职业道德底线，本课程将不遗余力的维护学术诚信规范，违反这一底线的行为将不会被容忍。&lt;br /&gt;
&lt;br /&gt;
作业完成的原则：署你名字的工作必须是你个人的贡献。在完成作业的过程中，允许讨论，前提是讨论的所有参与者均处于同等完成度。但关键想法的执行、以及作业文本的写作必须独立完成，并在作业中致谢（acknowledge）所有参与讨论的人。不允许其他任何形式的合作——尤其是与已经完成作业的同学“讨论”。&lt;br /&gt;
&lt;br /&gt;
本课程将对剽窃行为采取零容忍的态度。在完成作业过程中，对他人工作（出版物、互联网资料、其他人的作业等）直接的文本抄袭和对关键思想、关键元素的抄袭，按照 [http://www.acm.org/publications/policies/plagiarism_policy ACM Policy on Plagiarism]的解释，都将视为剽窃。剽窃者成绩将被取消。如果发现互相抄袭行为，&amp;lt;font color=red&amp;gt; 抄袭和被抄袭双方的成绩都将被取消&amp;lt;/font&amp;gt;。因此请主动防止自己的作业被他人抄袭。&lt;br /&gt;
&lt;br /&gt;
学术诚信影响学生个人的品行，也关乎整个教育系统的正常运转。为了一点分数而做出学术不端的行为，不仅使自己沦为一个欺骗者，也使他人的诚实努力失去意义。让我们一起努力维护一个诚信的环境。&lt;br /&gt;
&lt;br /&gt;
= Assignments =&lt;br /&gt;
* [[组合数学 (Spring 2026)/Problem Set 1|Problem Set 1]]&lt;br /&gt;
&lt;br /&gt;
= Lecture Notes =&lt;br /&gt;
# [[组合数学 (Spring 2026)/Basic enumeration|Basic enumeration | 基本计数]] ([http://tcs.nju.edu.cn/slides/comb2026/BasicEnumeration.pdf slides])&lt;br /&gt;
# [[组合数学 (Spring 2026)/Generating functions|Generating functions | 生成函数]] ([http://tcs.nju.edu.cn/slides/comb2026/GeneratingFunction.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Sieve methods|Sieve methods | 筛法]] ([http://tcs.nju.edu.cn/slides/comb2026/PIE.pdf slides])&lt;br /&gt;
# [[组合数学 (Fall 2026)/Cayley&#039;s formula|Cayley&#039;s formula | Cayley公式]]  ([http://tcs.nju.edu.cn/slides/comb2026/Cayley.pdf slides])&lt;br /&gt;
&lt;br /&gt;
= Resources =&lt;br /&gt;
* [http://math.mit.edu/~fox/MAT307.html Combinatorics course] by Jacob Fox&lt;br /&gt;
* [https://yufeizhao.com/pm/ Probabilistic Methods in Combinatorics] and [https://yufeizhao.com/gtacbook/ Graph Theory and Additive Combinatorics] by Yufei Zhao&lt;br /&gt;
* [https://www.math.uvic.ca/~noelj/combinatoricsLectures.html Combinatorics Lecture Videos online]&lt;br /&gt;
* [https://www.math.ucla.edu/~pak/lectures/Math-Videos/comb-videos.htm Collection of Combinatorics Videos]&lt;br /&gt;
&lt;br /&gt;
= Concepts =&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_coefficient Binomial coefficient]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Twelvefold_way The twelvefold way]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Composition_(number_theory) Composition of a number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multiset#Formal_definition Multiset]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Combination#Number_of_combinations_with_repetition Combinations with repetition], [http://en.wikipedia.org/wiki/Multiset#Counting_multisets &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on a set]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Multinomial_theorem#Multinomial_coefficients Multinomial coefficients]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Stirling_numbers_of_the_second_kind Stirling number of the second kind]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Partition_(number_theory) Partition of a number]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Young_tableau Young tableau]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Catalan_number Catalan number]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Generating_function Generating function] and [http://en.wikipedia.org/wiki/Formal_power_series formal power series]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Binomial_series Newton&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Inclusion-exclusion_principle The principle of inclusion-exclusion] (and more generally the [http://en.wikipedia.org/wiki/Sieve_theory sieve method])&lt;br /&gt;
* [http://en.wikipedia.org/wiki/M%C3%B6bius_inversion_formula Möbius inversion formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Derangement Derangement], and [http://en.wikipedia.org/wiki/M%C3%A9nage_problem Problème des ménages]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ryser%27s_formula#Ryser_formula Ryser&#039;s formula]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Euler_totient Euler totient function]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Burnside%27s_lemma Burnside&#039;s lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action Group action]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Group_action#Orbits_and_stabilizers Orbits]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/P%C3%B3lya_enumeration_theorem Pólya enumeration theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Permutation_group Permutation group]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Cycle_index Cycle index]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Cayley_formula Cayley&#039;s formula]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Prüfer_sequence Prüfer code for trees]&lt;br /&gt;
** [http://en.wikipedia.org/wiki/Kirchhoff%27s_matrix_tree_theorem Kirchhoff&#039;s matrix-tree theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Double_counting_(proof_technique) Double counting] and the [http://en.wikipedia.org/wiki/Handshaking_lemma handshaking lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Sperner&#039;s_lemma Sperner&#039;s lemma] and [http://en.wikipedia.org/wiki/Brouwer_fixed_point_theorem Brouwer fixed point theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Pigeonhole_principle Pigeonhole principle]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Dirichlet&#039;s_approximation_theorem Dirichlet&#039;s approximation theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Probabilistic_method The Probabilistic Method]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Lov%C3%A1sz_local_lemma Lovász local lemma]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93R%C3%A9nyi_model Erdős–Rényi model for random graphs]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Extremal_graph_theory Extremal graph theory]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Turan_theorem Turán&#039;s theorem], [http://en.wikipedia.org/wiki/Tur%C3%A1n_graph Turán graph]&lt;br /&gt;
* Two analytic inequalities: &lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Cauchy%E2%80%93Schwarz_inequality Cauchy–Schwarz inequality]&lt;br /&gt;
:* the [http://en.wikipedia.org/wiki/Inequality_of_arithmetic_and_geometric_means inequality of arithmetic and geometric means]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Stone_theorem Erdős–Stone theorem] (fundamental theorem of extremal graph theory)&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sunflower_(mathematics) Sunflower lemma and conjecture]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Ko%E2%80%93Rado_theorem Erdős–Ko–Rado theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sperner%27s_theorem Sperner&#039;s theorem]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Sperner_family Sperner system] or &#039;&#039;&#039;antichain&#039;&#039;&#039;&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Sauer%E2%80%93Shelah_lemma Sauer–Shelah lemma]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension Vapnik–Chervonenkis dimension]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Kruskal%E2%80%93Katona_theorem Kruskal–Katona theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Ramsey_theory Ramsey theory]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Ramsey&#039;s_theorem Ramsey&#039;s theorem]&lt;br /&gt;
:*[http://en.wikipedia.org/wiki/Happy_Ending_problem Happy Ending problem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Van_der_Waerden%27s_theorem Van der Waerden&#039;s theorem]&lt;br /&gt;
:*[https://en.wikipedia.org/wiki/Hales%E2%80%93Jewett_theorem Hales–Jewett theorem]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Hall%27s_marriage_theorem Hall&#039;s theorem ] (the marriage theorem)&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Doubly_stochastic_matrix Birkhoff–Von Neumann theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/K%C3%B6nig&#039;s_theorem_(graph_theory) König-Egerváry theorem]&lt;br /&gt;
* [http://en.wikipedia.org/wiki/Dilworth&#039;s_theorem Dilworth&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Erd%C5%91s%E2%80%93Szekeres_theorem Erdős–Szekeres theorem]&lt;br /&gt;
* The  [http://en.wikipedia.org/wiki/Max-flow_min-cut_theorem Max-Flow Min-Cut Theorem]&lt;br /&gt;
:* [https://en.wikipedia.org/wiki/Menger%27s_theorem Menger&#039;s theorem]&lt;br /&gt;
:* [http://en.wikipedia.org/wiki/Maximum_flow_problem Maximum flow]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Linear_programming Linear programming]&lt;br /&gt;
** [https://en.wikipedia.org/wiki/Dual_linear_program Duality] &lt;br /&gt;
** [https://en.wikipedia.org/wiki/Unimodular_matrix Unimodularity]&lt;br /&gt;
* [https://en.wikipedia.org/wiki/Matroid Matroid]&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)/Average-case_analysis_of_QuickSort&amp;diff=13617</id>
		<title>概率论与数理统计 (Spring 2026)/Average-case analysis of QuickSort</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=%E6%A6%82%E7%8E%87%E8%AE%BA%E4%B8%8E%E6%95%B0%E7%90%86%E7%BB%9F%E8%AE%A1_(Spring_2026)/Average-case_analysis_of_QuickSort&amp;diff=13617"/>
		<updated>2026-04-08T08:38:47Z</updated>

		<summary type="html">&lt;p&gt;Etone: Created page with &amp;quot;[http://en.wikipedia.org/wiki/Quicksort &amp;#039;&amp;#039;&amp;#039;快速排序&amp;#039;&amp;#039;&amp;#039;（&amp;#039;&amp;#039;&amp;#039;Quicksort&amp;#039;&amp;#039;&amp;#039;）]是由Tony Hoare发现的排序算法。该算法的伪代码描述如下（为方便起见，假设数组元素互不相同——更一般情况的分析易推广得到）：   &amp;#039;&amp;#039;&amp;#039;&amp;#039;&amp;#039;QSort&amp;#039;&amp;#039;&amp;#039;&amp;#039;&amp;#039;(A): 输入A[1...n]是存有n个不同数字的数组   if n&amp;gt;1 then        &amp;#039;&amp;#039;&amp;#039;pivot&amp;#039;&amp;#039;&amp;#039; = A[1];        将A中&amp;lt;pivot的元素存于数组L，将A中&amp;gt;pivot的元素存于数组R; \\保持内部元素之...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[http://en.wikipedia.org/wiki/Quicksort &#039;&#039;&#039;快速排序&#039;&#039;&#039;（&#039;&#039;&#039;Quicksort&#039;&#039;&#039;）]是由Tony Hoare发现的排序算法。该算法的伪代码描述如下（为方便起见，假设数组元素互不相同——更一般情况的分析易推广得到）：&lt;br /&gt;
  &#039;&#039;&#039;&#039;&#039;QSort&#039;&#039;&#039;&#039;&#039;(A): 输入A[1...n]是存有n个不同数字的数组&lt;br /&gt;
  if n&amp;gt;1 then&lt;br /&gt;
       &#039;&#039;&#039;pivot&#039;&#039;&#039; = A[1];&lt;br /&gt;
       将A中&amp;lt;pivot的元素存于数组L，将A中&amp;gt;pivot的元素存于数组R; \\保持内部元素之间相对顺序&lt;br /&gt;
       递归调用&#039;&#039;&#039;&#039;&#039;QSort&#039;&#039;&#039;&#039;&#039;(L)和&#039;&#039;&#039;&#039;&#039;QSort&#039;&#039;&#039;&#039;&#039;(R);&lt;br /&gt;
&lt;br /&gt;
该伪代码描述省略了对数组的存取归并等具体实现的交代。可以不妨认为算法中的&amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;和&amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;其实就分别是原数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;的前半部分&amp;lt;math&amp;gt;A[1...i-1]&amp;lt;/math&amp;gt;和后半部分&amp;lt;math&amp;gt;A[i...n]&amp;lt;/math&amp;gt;，其中&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;表示&#039;&#039;&#039;pivot&#039;&#039;&#039;（“&#039;&#039;&#039;基准&#039;&#039;&#039;”、“&#039;&#039;&#039;轴点&#039;&#039;&#039;”、“&#039;&#039;&#039;分水岭&#039;&#039;&#039;”元素）在数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中的序数，即&#039;&#039;&#039;pivot&#039;&#039;&#039;是数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中第&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;小的数。&lt;br /&gt;
&lt;br /&gt;
作为基于比较的（comparison-based）排序算法，我们将算法进行的元素间的&#039;&#039;&#039;比较次数&#039;&#039;&#039;作为算法的复杂性度量。在&#039;&#039;&#039;最坏情况&#039;&#039;&#039;（&#039;&#039;&#039;worst-case&#039;&#039;&#039;）输入下，我们描述的“快速排序”算法所使用的比较次数可以达到&amp;lt;math&amp;gt;\Theta(n^2)&amp;lt;/math&amp;gt;。而且有趣的是，这一&amp;lt;math&amp;gt;\Omega(n^2)&amp;lt;/math&amp;gt;的复杂度下界是在输入数组是一个已排好序的数组时达到的：此时&#039;&#039;&#039;pivot&#039;&#039;&#039;会将数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;划分为大小极不平衡的子数组&amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;和&amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;，递归树也因此极不平衡，而算法使用的比较次数可因此达到&amp;lt;math&amp;gt;(n-1)+(n-2)+(n-3)+\cdots+1=\Omega(n^2)&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
现在我们来分析这一算法的&#039;&#039;&#039;平均情况&#039;&#039;&#039;（&#039;&#039;&#039;average-case&#039;&#039;&#039;）复杂度。&lt;br /&gt;
令输入数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;为&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;个元素的均匀分布的随机排列。如下性质不难递归验证：&lt;br /&gt;
{{Theorem|性质一（输入分布的递归不变性）|&lt;br /&gt;
* 作为基于比较的排序算法，快速排序算法的行为仅与输入数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中元素的相对顺序有关，而与&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中元素的具体数值无关。&lt;br /&gt;
* 进一步地，在算法的每次递归调用中，其输入数组中元素的相对顺序，符合该数量元素的所有排列上的均匀分布。&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
= 快速排序算法的平均复杂度分析 I（基于全期望法则）=&lt;br /&gt;
令随机变量&amp;lt;math&amp;gt;X_n&amp;lt;/math&amp;gt;表示在随机输入&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;下，快速排序算法使用的总比较次数。同时令&amp;lt;math&amp;gt;t(n)=\mathbb{E}[X_n]&amp;lt;/math&amp;gt;表示其期望。正如刚刚解释的，&amp;lt;math&amp;gt;t(n)&amp;lt;/math&amp;gt;是良定义的，因为期望值&amp;lt;math&amp;gt;\mathbb{E}[X_n]&amp;lt;/math&amp;gt;仅与&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;有关。&lt;br /&gt;
&lt;br /&gt;
对&amp;lt;math&amp;gt;1\le i\le n&amp;lt;/math&amp;gt;，定义&amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt;为如下事件：&lt;br /&gt;
:*&#039;&#039;&#039;pivot&#039;&#039;&#039;在数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中是第&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;小的元素。&lt;br /&gt;
则&amp;lt;math&amp;gt;B_1,B_2,\ldots,B_n&amp;lt;/math&amp;gt;构成了对所有情况（样本空间）的一个划分。且每个事件&amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt;发生的概率有：&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr(B_i)=\frac{(n-1)!}{n!}=\frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
这是因为事件&amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt;等价于一个均匀分布的随机排列&amp;lt;math&amp;gt;\pi:[n]\xrightarrow[\text{onto}]{\text{1-1}}[n]&amp;lt;/math&amp;gt;有&amp;lt;math&amp;gt;\pi(1)=i&amp;lt;/math&amp;gt;，而这件事的概率为&amp;lt;math&amp;gt;\frac{1}{n}&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
因此，根据&#039;&#039;&#039;全期望法则&#039;&#039;&#039;（&#039;&#039;&#039;law of total expectation&#039;&#039;&#039;），有：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
t(n)&lt;br /&gt;
=&lt;br /&gt;
\mathbb{E}[X_n]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{i=1}^n\mathbb{E}[X_n\mid B_i]\Pr(B_i)&lt;br /&gt;
=&lt;br /&gt;
\frac{1}{n}\sum_{i=1}^n\mathbb{E}[X_n\mid B_i]&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
而当事件&amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt;发生时，即当&#039;&#039;&#039;pivot&#039;&#039;&#039;恰为数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中第&amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;小的元素，&lt;br /&gt;
算法会将&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中比&#039;&#039;&#039;pivot&#039;&#039;&#039;小的&amp;lt;math&amp;gt;(i-1)&amp;lt;/math&amp;gt;个元素放入子数组&amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;、将比&#039;&#039;&#039;pivot&#039;&#039;&#039;大的&amp;lt;math&amp;gt;(n-i)&amp;lt;/math&amp;gt;个元素放入子数组&amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;，这总共会花费&amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;次比较。&lt;br /&gt;
而且，根据&#039;&#039;&#039;性质一&#039;&#039;&#039;中所述的输入分布的递归不变性，递归调用&#039;&#039;&#039;&#039;&#039;QSort&#039;&#039;&#039;&#039;&#039;&amp;lt;math&amp;gt;(L)&amp;lt;/math&amp;gt;和&#039;&#039;&#039;&#039;&#039;QSort&#039;&#039;&#039;&#039;&#039;&amp;lt;math&amp;gt;(R)&amp;lt;/math&amp;gt;各自所使用的总比较次数分别为&amp;lt;math&amp;gt;X_{i-1}&amp;lt;/math&amp;gt;和&amp;lt;math&amp;gt;X_{n-i}&amp;lt;/math&amp;gt;。&lt;br /&gt;
综上所述，有如下恒等关系：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbb{E}[X_n\mid B_i]&lt;br /&gt;
=&lt;br /&gt;
\mathbb{E}[n-1+X_{i-1}+X_{n-i}]&lt;br /&gt;
&amp;lt;/math&amp;gt;.&lt;br /&gt;
因此，上述全期望因此可被计算如下：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
t(n)&lt;br /&gt;
=&lt;br /&gt;
\mathbb{E}[X_n]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1}{n}\sum_{i=1}^n\mathbb{E}[X_n\mid B_i]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1}{n}\sum_{i=1}^n\mathbb{E}[n-1+X_{i-1}+X_{n-i}]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
n-1+\frac{2}{n}\sum_{i=0}^{n-1}\mathbb{E}[X_{i}] &amp;amp;\text{(根据期望的线性)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
n-1+\frac{2}{n}\sum_{i=0}^{n-1}t(i).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
我们因此建立了平均复杂度&amp;lt;math&amp;gt;t(n)=\mathbb{E}[X_n]&amp;lt;/math&amp;gt;的如下递归式：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
t(n)&lt;br /&gt;
&amp;amp;= &lt;br /&gt;
n-1+\frac{2}{n}\sum_{i=0}^{n-1}t(i)&lt;br /&gt;
&amp;amp;&amp;amp; &lt;br /&gt;
\text{if }n&amp;gt;1;\\&lt;br /&gt;
t(n)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
0&lt;br /&gt;
&amp;amp;&amp;amp;&lt;br /&gt;
\text{if }0\le n\le 1.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
我们有若干种方法从这一递归式获得关于&amp;lt;math&amp;gt;t(n)=\mathbb{E}[X_n]&amp;lt;/math&amp;gt;的真相：&lt;br /&gt;
* 一种是使用&#039;&#039;&#039;生成函数&#039;&#039;&#039;等工具来求解这样的递归方程，例如使用[[组合数学_(Fall_2025)/Generating_functions#Analysis_of_Quicksort|&#039;&#039;&#039;本学期组合数学课上的分析&#039;&#039;&#039;]]，最终可精确解出&amp;lt;math&amp;gt;t(n)=2(n+1)H(n)-4n&amp;lt;/math&amp;gt;，此处&amp;lt;math&amp;gt;H(n)=\sum_{k=1}^{n}\frac{1}{k}= \ln n+O(1)&amp;lt;/math&amp;gt;为第&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;个[https://en.wikipedia.org/wiki/Harmonic_number 调和数]；&lt;br /&gt;
* 另一种是采用&#039;&#039;&#039;数学归纳法&#039;&#039;&#039;来对我们猜想的关于&amp;lt;math&amp;gt;t(n)&amp;lt;/math&amp;gt;的界进行验证，例如从&amp;lt;math&amp;gt;t(n)\le c n \ln n&amp;lt;/math&amp;gt;的归纳假设出发，尝试使&amp;lt;math&amp;gt;c&amp;lt;/math&amp;gt;尽量小。令该归纳假设对于所有&amp;lt;math&amp;gt;0\le i&amp;lt;n&amp;lt;/math&amp;gt;都成立，则对于&amp;lt;math&amp;gt;t(n)&amp;lt;/math&amp;gt;，根据递归式因此有：&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
t(n)&lt;br /&gt;
&amp;amp;= &lt;br /&gt;
n-1+\frac{2}{n}\sum_{i=0}^{n-1}t(i)\\&lt;br /&gt;
&amp;amp;\le &lt;br /&gt;
n-1+\frac{2c}{n}\sum_{i=1}^{n-1}i \ln i &amp;amp;&amp;amp;\text{(归纳假设)}\\&lt;br /&gt;
&amp;amp;\le &lt;br /&gt;
n-1+\frac{2c}{n}\int_{1}^nx \ln x\,\mathrm{d}x &amp;amp;&amp;amp;\text{(求和的积分上界)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
n-1+\frac{c}{2n}\left(2n^2\ln n-n^2+1\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
cn\ln n-\left(\frac{c}{2}-1\right)n-1+\frac{c}{2n}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
设&amp;lt;math&amp;gt;c=2&amp;lt;/math&amp;gt;时，有&amp;lt;math&amp;gt;t(n)\le cn\ln n-\left(\frac{c}{2}-1\right)n-1+\frac{c}{2n}\le 2 n \ln n&amp;lt;/math&amp;gt;对于所有&amp;lt;math&amp;gt;n\ge 1&amp;lt;/math&amp;gt;总是成立。&lt;br /&gt;
&lt;br /&gt;
= 快速排序算法的平均复杂度分析 II（基于期望的线性）=&lt;br /&gt;
令随机变量&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;表示在随机输入&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;下，快速排序算法使用的总比较次数。我们希望分析得到期望&amp;lt;math&amp;gt;\mathbb{E}[X]&amp;lt;/math&amp;gt;的上界。&lt;br /&gt;
&lt;br /&gt;
假设输入数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;中的元素，在从小到大排序之后为&amp;lt;math&amp;gt;a_1&amp;lt;a_2&amp;lt;\cdots a_n&amp;lt;/math&amp;gt;。&lt;br /&gt;
对&amp;lt;math&amp;gt;1\le i&amp;lt;j\le n&amp;lt;/math&amp;gt;，令布尔值随机变量&amp;lt;math&amp;gt;I_{ij}=I(A_{ij})\in\{0,1\}&amp;lt;/math&amp;gt;指示如下事件的发生：&lt;br /&gt;
* &amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;：元素&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;和&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;在算法运行过程中被比较了。&lt;br /&gt;
&lt;br /&gt;
容易验证：在算法运行过程中，任何元素对&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;和&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;之间至多只会进行一次比较。&lt;br /&gt;
因此总比较次数&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;可被计算如下：&lt;br /&gt;
:&amp;lt;math&amp;gt;X=\sum_{i&amp;lt;j}I_{ij}&amp;lt;/math&amp;gt;.&lt;br /&gt;
根据&#039;&#039;&#039;期望的线性&#039;&#039;&#039;（&#039;&#039;&#039;linearity of expectation&#039;&#039;&#039;）：&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbb{E}[X]=\sum_{i&amp;lt;j}\mathbb{E}[I_{ij}]=\sum_{i&amp;lt;j}\Pr(A_{ij})&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
因此接下来，仅需要针对每一对具体的&amp;lt;math&amp;gt;1\le i&amp;lt;j\le n&amp;lt;/math&amp;gt;，计算概率：&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr(A_{ij})=\Pr(a_i\text{和}a_j\text{在算法中进行了比较})&amp;lt;/math&amp;gt;.&lt;br /&gt;
而事件&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;发生，当且仅当：在&#039;&#039;&#039;某次&#039;&#039;&#039;递归调用中，&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;同属于当前输入数组，且&#039;&#039;&#039;pivot&#039;&#039;&#039;恰好选中&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;或&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;二者之一。&lt;br /&gt;
注意到：假如&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;同属于当前输入数组，则当前数组必然也同时包含它们之间的所有元素&amp;lt;math&amp;gt;\{a_k\mid i&amp;lt;k&amp;lt;j\}&amp;lt;/math&amp;gt;；&lt;br /&gt;
同时，事件&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;的发生与否，仅当&#039;&#039;&#039;首次&#039;&#039;&#039;在某递归调用中&#039;&#039;&#039;pivot&#039;&#039;&#039;选中了&amp;lt;math&amp;gt;\{a_k\mid i\le k\le j\}&amp;lt;/math&amp;gt;中的元素，才会确定。&lt;br /&gt;
综上，于是有：&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr(A_{ij})=\Pr(\mathsf{pivot}\in\{a_i,a_j\}\mid\mathsf{pivot}\in\{a_i,a_{i+1},\ldots,a_{j}\})=\frac{2}{j-i+1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|如果你认为该证明不够详细。这里再提供一个对于上述事实的详细证明。&lt;br /&gt;
&lt;br /&gt;
定下任意一对具体的&amp;lt;math&amp;gt;1\le i&amp;lt;j\le n&amp;lt;/math&amp;gt;。&lt;br /&gt;
在算法运行之初（顶层递归调用），元素&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;显然有&amp;lt;math&amp;gt;a_i&amp;lt;a_j&amp;lt;/math&amp;gt;，且同属于当前递归调用的输入数组&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;。&lt;br /&gt;
&lt;br /&gt;
事件&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;的发生与否，由如下的递归随机过程确定：&lt;br /&gt;
# 在当前输入元素中按均匀分布随机选择&#039;&#039;&#039;pivot&#039;&#039;&#039;元素。如果刚好选中&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;或&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;二者之一，那么&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;之间必被比较（因为&#039;&#039;&#039;pivot&#039;&#039;&#039;会与当前数组内的所有元素进行比较）。此情况下，事件&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;确定发生。&lt;br /&gt;
# 如果&amp;lt;math&amp;gt;a_i&amp;lt;&amp;lt;/math&amp;gt;&#039;&#039;&#039;pivot&#039;&#039;&#039;&amp;lt;math&amp;gt;&amp;lt;a_j&amp;lt;/math&amp;gt;，即&#039;&#039;&#039;pivot&#039;&#039;&#039;取值在&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;之间，此时&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;会被大小居于其中间的&#039;&#039;&#039;pivot&#039;&#039;&#039;元素分别分流到子数组&amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;中，且在之后的递归调用中永不再见，因此不会被比较。此情况下，事件&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;确定不发生。&lt;br /&gt;
# 如果&#039;&#039;&#039;pivot&#039;&#039;&#039;选中了比&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;更小或者比&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;更大的元素，在此情况下&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;，连同它们之间的所有元素&amp;lt;math&amp;gt;\{a_k\mid i&amp;lt;k&amp;lt;j\}&amp;lt;/math&amp;gt;，都会被分到同一个子数组&amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;（假如&#039;&#039;&#039;pivot&#039;&#039;&#039;&amp;lt;math&amp;gt;&amp;gt;a_j&amp;lt;/math&amp;gt;）或者&amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;（假如&#039;&#039;&#039;pivot&#039;&#039;&#039;&amp;lt;math&amp;gt;&amp;lt;a_i&amp;lt;/math&amp;gt;）中。而此时&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;的发生与否尚未能在这一层递归调用中确定，则需进入下一层递归调用，将当前输入数组改为&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;所同属的那个子数组（即&amp;lt;math&amp;gt;L&amp;lt;/math&amp;gt;如果&#039;&#039;&#039;pivot&#039;&#039;&#039;&amp;lt;math&amp;gt;&amp;gt;a_j&amp;lt;/math&amp;gt;，&amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;如果&#039;&#039;&#039;pivot&#039;&#039;&#039;&amp;lt;math&amp;gt;&amp;lt;a_i&amp;lt;/math&amp;gt;），然后回到（1）重复。&lt;br /&gt;
不难看出该递归过程跟踪模拟了事件&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;的发生与否，在快速排序算法的递归调用过程中，是如何被确定下来的。&lt;br /&gt;
&lt;br /&gt;
因此，可以得出：&amp;lt;math&amp;gt;A_{ij}&amp;lt;/math&amp;gt;发生，当且仅当在上述情况（1）或（2）之一发生的前提下有（1）发生。&lt;br /&gt;
&lt;br /&gt;
除此之外，如下两个观察通过递归验证易得：&lt;br /&gt;
* 当&amp;lt;math&amp;gt;a_i&amp;lt;/math&amp;gt;与&amp;lt;math&amp;gt;a_j&amp;lt;/math&amp;gt;同属于当前递归调用的输入数组时，它们之间的所有元素&amp;lt;math&amp;gt;\{a_k\mid i&amp;lt;k&amp;lt;j\}&amp;lt;/math&amp;gt;也属于该数组；&lt;br /&gt;
* 在任何递归调用中，&#039;&#039;&#039;pivot&#039;&#039;&#039;始终在当前输入数组中均匀分布——这也可由&#039;&#039;&#039;性质一&#039;&#039;&#039;所述的输入分布的递归不变性得到。&lt;br /&gt;
综上可得：&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr(A_{ij})=\Pr(\mathsf{pivot}\in\{a_i,a_j\}\mid\mathsf{pivot}\in\{a_i,a_{i+1},\ldots,a_{j}\})=\frac{2}{j-i+1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
算法总比较次数的期望&amp;lt;math&amp;gt;\mathbb{E}[X]&amp;lt;/math&amp;gt;可被计算如下：&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbb{E}\left[X\right] &lt;br /&gt;
&amp;amp;= &lt;br /&gt;
\sum_{i&amp;lt;j}\Pr(A_{ij})\\&lt;br /&gt;
&amp;amp;=\sum_{i=1}^{n-1}\sum_{j=i+1}^n\frac{2}{j-i+1}\\&lt;br /&gt;
&amp;amp;= \sum_{i=1}^{n-1}\sum_{k=2}^{n-i+1}\frac{2}{k} &amp;amp; &amp;amp; (\text{令 }k=j-i+1)\\&lt;br /&gt;
&amp;amp;\le \sum_{i=1}^n\sum_{k=1}^{n}\frac{2}{k}\\&lt;br /&gt;
&amp;amp;= 2n\sum_{k=1}^{n}\frac{1}{k}\\&lt;br /&gt;
&amp;amp;= 2n H(n).&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
此处&amp;lt;math&amp;gt;H(n)=\sum_{k=1}^{n}\frac{1}{k}= \ln n+O(1)&amp;lt;/math&amp;gt;为第&amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;个[http://en.wikipedia.org/wiki/Harmonic_number 调和数]。&lt;br /&gt;
&lt;br /&gt;
如果更加仔细一些，我们甚至可以计算出总比较次数期望&amp;lt;math&amp;gt;\mathbb{E}[X]&amp;lt;/math&amp;gt;的精确值：&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbb{E}\left[X\right] &lt;br /&gt;
&amp;amp;= &lt;br /&gt;
\sum_{i=1}^{n-1}\sum_{k=2}^{n-i+1}\frac{2}{k}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{j=2}^n\sum_{k=2}^{j}\frac{2}{k} &amp;amp;&amp;amp; \text{(令$j=n-i+1$)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k=2}^{n}\frac{2(n-k+1)}{k} &amp;amp;&amp;amp; \text{(交换求和顺序)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k=2}^{n}\left(\frac{2(n+1)}{k}-2\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
2(n+1)\sum_{k=2}^{n}\frac{1}{k}-2(n-1)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
2(n+1)\sum_{k=1}^{n}\frac{1}{k}-4n\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
2(n+1)H(n)-4n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Etone</name></author>
	</entry>
</feed>