<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://tcs.nju.edu.cn/wiki/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=172.21.1.0%2F24</id>
	<title>TCS Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://tcs.nju.edu.cn/wiki/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=172.21.1.0%2F24"/>
	<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Special:Contributions/172.21.1.0/24"/>
	<updated>2026-09-15T15:56:06Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.46.0</generator>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3518</id>
		<title>Combinatorics (Fall 2010)/Extremal graphs</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3518"/>
		<updated>2010-10-14T07:14:17Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.240: /* Erdős–Stone theorem */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Extremal Graph Theory ==&lt;br /&gt;
&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}} &lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős–Stone theorem ===&lt;br /&gt;
Let &amp;lt;math&amp;gt;K_s^r=K_{\underbrace{s,s,\cdots,s}_{r}}&amp;lt;/math&amp;gt; be the complete &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-partite graph with &amp;lt;math&amp;gt;s&amp;lt;/math&amp;gt; vertices in each class, i.e., the Turán graph &amp;lt;math&amp;gt;T(rs,r)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any integers &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and any &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is sufficiently large then every graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices and with at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph, i.e.,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\mathrm{ex}(n,K_s^r)= \left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:For every nonempty graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\lim_{n\rightarrow\infty}\frac{\mathrm{ex}(n,H)}{{n\choose 2}}=\frac{\chi(H)-2}{\chi(H)-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Cycle Structures ==&lt;br /&gt;
=== Girth ===&lt;br /&gt;
&lt;br /&gt;
=== Hamiltonian cycle ===&lt;/div&gt;</summary>
		<author><name>172.21.1.240</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3517</id>
		<title>Combinatorics (Fall 2010)/Extremal graphs</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3517"/>
		<updated>2010-10-14T07:04:04Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.240: /* Erdős–Stone theorem */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Extremal Graph Theory ==&lt;br /&gt;
&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}} &lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős–Stone theorem ===&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any integers &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and any &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;, if &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is sufficiently large then every graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices and with at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph, i.e.,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\mathrm{ex}(n,K_{r,s})= \left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:For every nonempty graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\lim_{n\rightarrow\infty}\frac{\mathrm{ex}(n,H)}{{n\choose 2}}=\frac{\chi(H)-2}{\chi(H)-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Cycle Structures ==&lt;br /&gt;
=== Girth ===&lt;br /&gt;
&lt;br /&gt;
=== Hamiltonian cycle ===&lt;/div&gt;</summary>
		<author><name>172.21.1.240</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3516</id>
		<title>Combinatorics (Fall 2010)/Extremal graphs</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3516"/>
		<updated>2010-10-14T07:00:05Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.240: /* Erdős–Stone theorem */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Extremal Graph Theory ==&lt;br /&gt;
&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}} &lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős–Stone theorem ===&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any integers &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and any &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;, there exists an &amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt; such that every graph with &amp;lt;math&amp;gt;n\ge N_0&amp;lt;/math&amp;gt; vertices and at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph, i.e.,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\mathrm{ex}(n,K_{r,s})= \left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:For every nonempty graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\lim_{n\rightarrow\infty}\frac{\mathrm{ex}(n,H)}{{n\choose 2}}=\frac{\chi(H)-2}{\chi(H)-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Cycle Structures ==&lt;br /&gt;
=== Girth ===&lt;br /&gt;
&lt;br /&gt;
=== Hamiltonian cycle ===&lt;/div&gt;</summary>
		<author><name>172.21.1.240</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3515</id>
		<title>Combinatorics (Fall 2010)/Extremal graphs</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3515"/>
		<updated>2010-10-14T06:59:33Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.240: /* Erdős–Stone theorem */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Extremal Graph Theory ==&lt;br /&gt;
&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}} &lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős–Stone theorem ===&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any integers &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and any &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;, there exists an &amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt; such that every graph with &amp;lt;math&amp;gt;n\ge N_0&amp;lt;/math&amp;gt; vertices and at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph, i.e.,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\mathrm{ex}(n,K_{s,r})= \left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:For every nonempty graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\lim_{n\rightarrow\infty}\frac{\mathrm{ex}(n,H)}{{n\choose 2}}=\frac{\chi(H)-2}{\chi(H)-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Cycle Structures ==&lt;br /&gt;
=== Girth ===&lt;br /&gt;
&lt;br /&gt;
=== Hamiltonian cycle ===&lt;/div&gt;</summary>
		<author><name>172.21.1.240</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3514</id>
		<title>Combinatorics (Fall 2010)/Extremal graphs</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3514"/>
		<updated>2010-10-14T06:58:37Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.240: /* Erdős–Stone theorem */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Extremal Graph Theory ==&lt;br /&gt;
&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}} &lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős–Stone theorem ===&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any integers &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and any &amp;lt;math&amp;gt;\epsilon&amp;gt;0&amp;lt;/math&amp;gt;, there exists an &amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt; such that every graph with &amp;lt;math&amp;gt;n\ge N_0&amp;lt;/math&amp;gt; vertices and at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph, i.e.,&lt;br /&gt;
:::&amp;lt;math&amp;gt;\mathrm{ex}(n,K_{s,r})\le \left(\frac{r-2}{2(r-1)}+o(1)\right)n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
:For every nonempty graph &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\lim_{n\rightarrow\infty}\frac{\mathrm{ex}(n,H)}{{n\choose 2}}=\frac{\chi(H)-2}{\chi(H)-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Cycle Structures ==&lt;br /&gt;
=== Girth ===&lt;br /&gt;
&lt;br /&gt;
=== Hamiltonian cycle ===&lt;/div&gt;</summary>
		<author><name>172.21.1.240</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3513</id>
		<title>Combinatorics (Fall 2010)/Extremal graphs</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Extremal_graphs&amp;diff=3513"/>
		<updated>2010-10-14T06:50:52Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.240: /* Erdős–Stone theorem */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Extremal Graph Theory ==&lt;br /&gt;
&lt;br /&gt;
=== Mantel&#039;s theorem ===&lt;br /&gt;
We consider a typical extremal problem for graphs: the largest possible number of edges of &#039;&#039;&#039;triangle-free&#039;&#039;&#039; graphs, i.e. graphs contains no &amp;lt;math&amp;gt;K_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem (Mantel 1907)|&lt;br /&gt;
:Suppose &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; is graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertice without triangles. Then &amp;lt;math&amp;gt;|E|\le\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (pigeonhole principle)|&lt;br /&gt;
We prove an equivalent theorem: Any &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;|E|&amp;gt;\frac{n^2}{4}&amp;lt;/math&amp;gt; must have a triangle.&lt;br /&gt;
&lt;br /&gt;
Use induction on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The theorem holds trivially for &amp;lt;math&amp;gt;n\le 3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Induction hypothesis: assume the theorem hold for &amp;lt;math&amp;gt;|V|\le n-1&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices, without loss of generality, assume that &amp;lt;math&amp;gt;|E|=\frac{n^2}{4}+1&amp;lt;/math&amp;gt;, we will show that &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must contain a triangle. Take a &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; be the subgraph of &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; induced by &amp;lt;math&amp;gt;V\setminus \{u,v\}&amp;lt;/math&amp;gt;. Clearly, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt; vertices.&lt;br /&gt;
:&#039;&#039;&#039;Case.1:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;&amp;gt;\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then by the induction hypothesis, &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
:&#039;&#039;&#039;Case.2:&#039;&#039;&#039; If &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;\le\frac{(n-2)^2}{4}&amp;lt;/math&amp;gt; edges, then at least &amp;lt;math&amp;gt;\left(\frac{n^2}{4}+1\right)-\frac{(n-2)^2}{4}-1=n-1&amp;lt;/math&amp;gt; edges are between &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{u,v\}&amp;lt;/math&amp;gt;. By pigeonhole principle, there must be a vertex in &amp;lt;math&amp;gt;H&amp;lt;/math&amp;gt; that is adjacent to both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. Thus, &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has a triangle.&lt;br /&gt;
}} &lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (Cauchy-Schwarz inequality)|(Mantel&#039;s original proof)&lt;br /&gt;
For any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, no vertex can be a neighbor of both &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, or otherwise there will be a triangle. Thus, for any edge &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;d_u+d_v\le n&amp;lt;/math&amp;gt;. It follows that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)\le n|E|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Note that &amp;lt;math&amp;gt;d(v)&amp;lt;/math&amp;gt; appears exactly &amp;lt;math&amp;gt;d_v&amp;lt;/math&amp;gt; times in the sum, so that&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{uv\in E}(d_u+d_v)=\sum_{v\in V}d_v^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Applying Chauchy-Schwarz inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
n|E|\ge\sum_{v\in V}d_v^2\ge\frac{\left(\sum_{v\in V}d_v\right)^2}{n}=\frac{4|E|^2}{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where the last equation is due to Euler&#039;s equality &amp;lt;math&amp;gt;\sum_{v\in V}d_v=2|E|&amp;lt;/math&amp;gt;. The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (inequality of the arithmetic and geometric mean)|&lt;br /&gt;
Assume that &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; has &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt; vertices and is triangle-free.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be the largest independent set in &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\alpha=|A|&amp;lt;/math&amp;gt;. &lt;br /&gt;
Since &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is triangle-free, for very vertex &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;, all its neighbors must form an independent set, thus &amp;lt;math&amp;gt;d(v)\le \alpha&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;v\in V&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=V\setminus A&amp;lt;/math&amp;gt; and let &amp;lt;math&amp;gt;\beta=|B|&amp;lt;/math&amp;gt;.&lt;br /&gt;
Since &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is an independent set, all edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; must have at least one endpoint in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Counting the edges in &amp;lt;math&amp;gt;E&amp;lt;/math&amp;gt; according to their endpoints in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, we obtain &amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v&amp;lt;/math&amp;gt;. By the inequality of the arithmetic and geometric mean,&lt;br /&gt;
:&amp;lt;math&amp;gt;|E|\le\sum_{v\in B}d_v\le\alpha\beta\le\left(\frac{\alpha+\beta}{2}\right)^2=\frac{n^2}{4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Turán&#039;s theorem ===&lt;br /&gt;
{{Theorem|Theorem (Turán 1941)|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a graph with &amp;lt;math&amp;gt;|V|=n&amp;lt;/math&amp;gt;. If &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, &amp;lt;math&amp;gt;k\ge 2&amp;lt;/math&amp;gt;, then&lt;br /&gt;
::&amp;lt;math&amp;gt;|E|\le\frac{r-2}{2(r-1)}n^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|First proof. (induction)|(Turán&#039;s original proof)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Second proof. (weight shifting)|(due to Motzkin and Straus)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Third proof. (the probabilistic method)|(due to Alon and Spencer)&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Prooftitle|Fourth proof.|&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(V,E)&amp;lt;/math&amp;gt; be a &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free graph on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices with a maximum number of edges.&lt;br /&gt;
:&#039;&#039;&#039;Claim:&#039;&#039;&#039; &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; does not contain three vertices &amp;lt;math&amp;gt;u,v,w&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;uv\in E&amp;lt;/math&amp;gt; but &amp;lt;math&amp;gt;uw\not\in E, vw\not\in E&amp;lt;/math&amp;gt;.&lt;br /&gt;
Suppose otherwise. There are two cases.&lt;br /&gt;
* &#039;&#039;&#039;Case.1:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;d(w)&amp;lt;d(v)&amp;lt;/math&amp;gt;. Without loss of generality, suppose that &amp;lt;math&amp;gt;d(w)&amp;lt;d(u)&amp;lt;/math&amp;gt;. We duplicate &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; by creating a new vertex &amp;lt;math&amp;gt;u&#039;&amp;lt;/math&amp;gt; which has exactly the same neighbors as &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; (but &amp;lt;math&amp;gt;uu&#039;&amp;lt;/math&amp;gt; is not an edge). Such duplication will not increase the clique size. We then remove &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt;. The resulting graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is still &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique-free, and has &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices. The number of edges in &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+d(u)-d(w)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
:which contradicts the assumption that &amp;lt;math&amp;gt;|E(G)|&amp;lt;/math&amp;gt; is maximal.&lt;br /&gt;
* &#039;&#039;&#039;Case.2:&#039;&#039;&#039; &amp;lt;math&amp;gt;d(w)\ge d(u)&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;d(w)\ge d(v)&amp;lt;/math&amp;gt;. Duplicate &amp;lt;math&amp;gt;w&amp;lt;/math&amp;gt; twice and delete &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;v&amp;lt;/math&amp;gt;. The new graph &amp;lt;math&amp;gt;G&#039;&amp;lt;/math&amp;gt; has no &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt;-clique, and the number of edges is&lt;br /&gt;
::&amp;lt;math&amp;gt;|E(G&#039;)|=|E(G)|+2d(w)-(d(u)+d(v)+1)&amp;gt;|E(G)|\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:Contradiction again.&lt;br /&gt;
&lt;br /&gt;
The claim implies that &amp;lt;math&amp;gt;uv\not\in E&amp;lt;/math&amp;gt; defines an equivalence relation on vertices (to be more precise, it guarantees the transitivity of the relation, while the reflexivity and symmetry hold directly). Graph &amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; must be a complete multipartite graph &amp;lt;math&amp;gt;K_{n_1,n_2,\ldots,n_{r-1}}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;n_1+n_2+\cdots +n_{r-1}=n&amp;lt;/math&amp;gt;. Optimize the edge number, we have the Turán graph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Erdős–Stone theorem ===&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Fundamental theorem of extremal graph theory (Erdős–Stone 1946)|&lt;br /&gt;
:For any &amp;lt;math&amp;gt;r\ge 2&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;s\ge 1&amp;lt;/math&amp;gt;, and every sufficiently large &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, every graph with &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; vertices and at least &amp;lt;math&amp;gt;\left(\frac{r-2}{2(r-1)}+\epsilon\right)n^2&amp;lt;/math&amp;gt; edges contains &amp;lt;math&amp;gt;K_{r,s}&amp;lt;/math&amp;gt; as a subgraph.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Corollary|&lt;br /&gt;
&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Cycle Structures ==&lt;br /&gt;
=== Girth ===&lt;br /&gt;
&lt;br /&gt;
=== Hamiltonian cycle ===&lt;/div&gt;</summary>
		<author><name>172.21.1.240</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Partitions,_sieve_methods&amp;diff=3002</id>
		<title>Combinatorics (Fall 2010)/Partitions, sieve methods</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Partitions,_sieve_methods&amp;diff=3002"/>
		<updated>2010-09-12T06:33:00Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Reference */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Partitions ==&lt;br /&gt;
We count the ways of partitioning &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;identical&#039;&#039; objects into &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; &#039;&#039;unordered&#039;&#039; groups. This is equivalent to counting the ways partitioning a number &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; unordered parts.&lt;br /&gt;
&lt;br /&gt;
A &#039;&#039;&#039;&amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition&#039;&#039;&#039; of a number &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is a multiset &amp;lt;math&amp;gt;\{x_1,x_2,\ldots,x_k\}&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;x_i\ge 1&amp;lt;/math&amp;gt; for every element &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We define &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt; as the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For example, number 7 has the following partitions:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\{7\}&lt;br /&gt;
&amp;amp; p_1(7)=1\\&lt;br /&gt;
&amp;amp;\{1,6\},\{2,5\},\{3,4\}&lt;br /&gt;
&amp;amp; p_2(7)=3\\&lt;br /&gt;
&amp;amp;\{1,1,5\}, \{1,2,4\}, \{1,3,3\}, \{2,2,3\} &lt;br /&gt;
&amp;amp; p_3(7)=4\\&lt;br /&gt;
&amp;amp;\{1,1,1,4\},\{1,1,2,3\}, \{1,2,2,2\}&lt;br /&gt;
&amp;amp; p_4(7)=3\\&lt;br /&gt;
&amp;amp;\{1,1,1,1,3\},\{1,1,1,2,2\}&lt;br /&gt;
&amp;amp; p_5(7)=2\\&lt;br /&gt;
&amp;amp;\{1,1,1,1,1,2\}&lt;br /&gt;
&amp;amp; p_6(7)=1\\&lt;br /&gt;
&amp;amp;\{1,1,1,1,1,1,1\}&lt;br /&gt;
&amp;amp; p_7(7)=1&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Equivalently, we can also define that A &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of a number &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-tuple &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; with:&lt;br /&gt;
* &amp;lt;math&amp;gt;x_1\ge x_2\ge\cdots\ge x_k\ge 1&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt; the number of integral solutions to the above system.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p(n)=\sum_{k=1}^n p_k(n)&amp;lt;/math&amp;gt; be the total number of partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. The function &amp;lt;math&amp;gt;p(n)&amp;lt;/math&amp;gt; is called the &#039;&#039;&#039;partition number&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
=== Counting &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;===&lt;br /&gt;
We now try to determine &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;. Unlike most problems we learned in the last lecture, &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt; does not have a nice closed form formula. We now give a recurrence for &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:&amp;lt;math&amp;gt;p_k(n)=p_{k-1}(n-1)+p_k(n-k)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Suppose that &amp;lt;math&amp;gt;(x_1,\ldots,x_k)&amp;lt;/math&amp;gt; is a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Note that it must hold that&lt;br /&gt;
:&amp;lt;math&amp;gt;x_1\ge x_2\ge \cdots \ge x_k\ge 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
There are two cases: &amp;lt;math&amp;gt;x_k=1&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;x_k&amp;gt;1&amp;lt;/math&amp;gt;.&lt;br /&gt;
;Case 1.&lt;br /&gt;
:If &amp;lt;math&amp;gt;x_k=1&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;(x_1,\cdots,x_{k-1})&amp;lt;/math&amp;gt; is a distinct &amp;lt;math&amp;gt;(k-1)&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;. And every &amp;lt;math&amp;gt;(k-1)&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt; can be obtained in this way. Thus the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; in this case is &amp;lt;math&amp;gt;p_{k-1}(n-1)&amp;lt;/math&amp;gt;. &lt;br /&gt;
;Case 2.&lt;br /&gt;
:If &amp;lt;math&amp;gt;x_k&amp;gt;1&amp;lt;/math&amp;gt;, then &amp;lt;math&amp;gt;(x_1-1,\cdots,x_{k}-1)&amp;lt;/math&amp;gt; is a distinct &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt;. And every &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; can be obtained in this way. Thus the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; in this case is &amp;lt;math&amp;gt;p_{k}(n-k)&amp;lt;/math&amp;gt;. &lt;br /&gt;
In conclusion, the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;p_{k-1}(n-1)+p_k(n-k)&amp;lt;/math&amp;gt;, i.e.&lt;br /&gt;
:&amp;lt;math&amp;gt;p_k(n)=p_{k-1}(n-1)+p_k(n-k)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Use the above recurrence, we can compute the &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;  for some decent &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; by computer simulation.&lt;br /&gt;
&lt;br /&gt;
If we are not restricted ourselves to the precise estimation of &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;, the next theorem gives an asymptotic estimation of &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;. Note that it only holds for &#039;&#039;&#039;constant&#039;&#039;&#039; &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;, i.e. &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; does not depend on &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
For any fixed &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;p_k(n)\sim\frac{n^{k-1}}{k!(k-1)!}&amp;lt;/math&amp;gt;,&lt;br /&gt;
as &amp;lt;math&amp;gt;n\rightarrow \infty&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Suppose that &amp;lt;math&amp;gt;(x_1,\ldots,x_k)&amp;lt;/math&amp;gt; is a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1\ge x_2\ge \cdots \ge x_k\ge 1&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;k!&amp;lt;/math&amp;gt; permutations of &amp;lt;math&amp;gt;(x_1,\ldots,x_k)&amp;lt;/math&amp;gt; yield at most &amp;lt;math&amp;gt;k!&amp;lt;/math&amp;gt; many &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-compositions (the &#039;&#039;ordered&#039;&#039; sum of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; positive integers). There are &amp;lt;math&amp;gt;{n-1\choose k-1}&amp;lt;/math&amp;gt; many &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, every one of which can be yielded in this way by permuting a partition. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;k!p_k(n)\ge{n-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;y_i=x_i+k-i&amp;lt;/math&amp;gt;. That is, &amp;lt;math&amp;gt;y_k=x_k, y_{k-1}=x_k+1, y_{k-2}=x_k+2,\ldots, y_{1}=x_k+k-1&amp;lt;/math&amp;gt;. Then, it holds that&lt;br /&gt;
* &amp;lt;math&amp;gt;y_1&amp;gt;y_2&amp;gt;\cdots&amp;gt;y_k\ge 1&amp;lt;/math&amp;gt;; and &lt;br /&gt;
* &amp;lt;math&amp;gt;y_1+y_2+\cdots+y_k=n+\frac{k(k-1)}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Each permutation of &amp;lt;math&amp;gt;(y_1,y_2,\ldots,y_k)&amp;lt;/math&amp;gt; yields a &#039;&#039;&#039;distinct&#039;&#039;&#039; &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-composition of &amp;lt;math&amp;gt;n+\frac{k(k-1)}{2}&amp;lt;/math&amp;gt;, because all &amp;lt;math&amp;gt;y_i&amp;lt;/math&amp;gt; are distinct.&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;k!p_k(n)\le {n+\frac{k(k-1)}{2}-1\choose k-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Combining the two inequalities, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{{n-1\choose k-1}}{k!}\le p_k(n)\le \frac{{n+\frac{k(k-1)}{2}-1\choose k-1}}{k!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The theorem follows.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Ferrers diagram ===&lt;br /&gt;
A partition of a number &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be represented as a diagram of dots (or squares), called a &#039;&#039;&#039;Ferrers diagram&#039;&#039;&#039; (the square version of Ferrers diagram is also called a &#039;&#039;&#039;Young diagram&#039;&#039;&#039;, named after a structured called Young tableaux). &lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; with that &amp;lt;math&amp;gt;x_1\ge x_2\ge \cdots x_k\ge 1&amp;lt;/math&amp;gt; be a partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Its Ferrers diagram consists of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; rows, where the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th row contains &amp;lt;math&amp;gt;x_i&amp;lt;/math&amp;gt; dots (or squares).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
{|border=&amp;quot;0&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
{|border=&amp;quot;0&amp;quot;&lt;br /&gt;
|[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess xot45.svg|22px]]||[[File:Chess xot45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess xot45.svg|22px]]&lt;br /&gt;
|}&lt;br /&gt;
|&lt;br /&gt;
[[File:Chess t45.svg|120px]]&lt;br /&gt;
|align=center|&lt;br /&gt;
{|border=&amp;quot;2&amp;quot;  cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]&lt;br /&gt;
|}&lt;br /&gt;
|-&lt;br /&gt;
|align=center|Ferrers diagram (&#039;&#039;dot version&#039;&#039;) of (5,4,2,1)||&lt;br /&gt;
|align=center|Ferrers diagram (&#039;&#039;square version&#039;&#039;) of (5,4,2,1)&lt;br /&gt;
|}&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
;Conjugate partition&lt;br /&gt;
The partition we get by reading the Ferrers diagram by column instead of rows is called the &#039;&#039;&#039;conjugate&#039;&#039;&#039; of the original partition.&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
{|border=&amp;quot;0&amp;quot;&lt;br /&gt;
|align=center|&lt;br /&gt;
{|border=&amp;quot;2&amp;quot;  cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]&lt;br /&gt;
|}&lt;br /&gt;
|&lt;br /&gt;
[[File:Chess t45.svg|120px]]&lt;br /&gt;
|align=center|&lt;br /&gt;
{|border=&amp;quot;2&amp;quot;  cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]||[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]&lt;br /&gt;
|-&lt;br /&gt;
|[[File:Chess t45.svg|22px]]&lt;br /&gt;
|}&lt;br /&gt;
|-&lt;br /&gt;
|align=center|&amp;lt;math&amp;gt;(6,4,4,2,1)&amp;lt;/math&amp;gt;||&lt;br /&gt;
|align=center|conjugate: &amp;lt;math&amp;gt;(5,4,3,3,1,1)&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Clearly, &lt;br /&gt;
* different partitions cannot have the same conjugate, and &lt;br /&gt;
* every partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is the conjugate of some partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;,&lt;br /&gt;
so the conjugation mapping is a permutation on the set of partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This fact is very useful in proving theorems for partitions numbers.&lt;br /&gt;
&lt;br /&gt;
Some theorems of partitions can be easily proved by representing partitions in Ferrers diagrams.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
# The number of partitions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; which have largest summand &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;, is &amp;lt;math&amp;gt;p_k(n)&amp;lt;/math&amp;gt;. &lt;br /&gt;
# The number of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; parts equals the number of partitions of &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; into at most &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; parts. Formally,&lt;br /&gt;
::&amp;lt;math&amp;gt;p_k(n)=\sum_{j=1}^k p_j(n-k)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
# For every &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition, the conjugate partition has largest part &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;. And vice versa.&lt;br /&gt;
# For a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, remove the leftmost cell of every row of the Ferrers diagram. Totally &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; cells are removed and the remaining diagram is a partition of &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; into at most &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; parts. And for a partition of &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; into at most &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; parts, add a cell to each of the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; rows (including the empty ones). This will give us a &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-partition of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. It is easy to see the above mappings are 1-1 correspondences. Thus, the number of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; into &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; parts equals the number of partitions of &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; into at most &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; parts.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Principle of Inclusion-Exclusion ==&lt;br /&gt;
Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; be two finite sets. The cardinality of their union is&lt;br /&gt;
:&amp;lt;math&amp;gt;|A\cup B|=|A|+|B|-{\color{Blue}|A\cap B|}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For three sets &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, the cardinality of the union of these three sets is computed as&lt;br /&gt;
:&amp;lt;math&amp;gt;|A\cup B\cup C|=|A|+|B|+|C|-{\color{Blue}|A\cap B|}-{\color{Blue}|A\cap C|}-{\color{Blue}|B\cap C|}+{\color{Red}|A\cap B\cap C|}&amp;lt;/math&amp;gt;.&lt;br /&gt;
This is illustrated by the following figure.&lt;br /&gt;
::[[Image:Inclusion-exclusion.png|200px|border|center]] &lt;br /&gt;
&lt;br /&gt;
Generally, the &#039;&#039;&#039;Principle of Inclusion-Exclusion&#039;&#039;&#039; states the rule for computing the union of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; finite sets &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt;, such that&lt;br /&gt;
{{Equation|&lt;br /&gt;
&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\left|\bigcup_{i=1}^nA_i\right|&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{I\subseteq\{1,\ldots,n\}}(-1)^{|I|-1}\left|\bigcap_{i\in I}A_i\right|.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
-----&lt;br /&gt;
&lt;br /&gt;
In combinatorial enumeration, the Principle of Inclusion-Exclusion is usually applied in its complement form.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n\subseteq U&amp;lt;/math&amp;gt; be subsets of some finite set &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt;. Here &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; is some universe of combinatorial objects, whose cardinality is easy to calculate (e.g. all strings, tuples, permutations), and each &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; contains the objects with some specific property (e.g. a &amp;quot;pattern&amp;quot;) which we want to avoid. The problem is to count the number of objects without any of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; properties. We write &amp;lt;math&amp;gt;\bar{A_i}=U-A&amp;lt;/math&amp;gt;. The number of objects without any of the properties &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
{{Equation|&lt;br /&gt;
&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\left|\bar{A_1}\cap\bar{A_2}\cap\cdots\cap\bar{A_n}\right|=\left|U-\bigcup_{i=1}^nA_i\right|&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
|U|-\sum_{I\subseteq\{1,\ldots,n\}}(-1)^{|I|}\left|\bigcap_{i\in I}A_i\right|.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
For an &amp;lt;math&amp;gt;I\subseteq\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, we denote&lt;br /&gt;
:&amp;lt;math&amp;gt;A_I=\bigcap_{i\in I}A_i&amp;lt;/math&amp;gt;&lt;br /&gt;
with the convention that &amp;lt;math&amp;gt;A_\emptyset=U&amp;lt;/math&amp;gt;. The above equation is stated as:&lt;br /&gt;
{{Theorem|Principle of Inclusion-Exclusion|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A_1,A_2,\ldots,A_n&amp;lt;/math&amp;gt; be a family of subsets of &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt;. Then the number of elements of &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; which lie in none of the subsets &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; is&lt;br /&gt;
::&amp;lt;math&amp;gt;\sum_{I\subseteq\{1,\ldots, n\}}(-1)^{|I|}|A_I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S_k=\sum_{|I|=k}|A_I|\,&amp;lt;/math&amp;gt;. Conventionally, &amp;lt;math&amp;gt;S_0=|A_\emptyset|=|U|&amp;lt;/math&amp;gt;. The principle of inclusion-exclusion can be expressed as&lt;br /&gt;
{{Equation|&amp;lt;math&amp;gt;&lt;br /&gt;
S_0-S_1+S_2+\cdots+(-1)^nS_n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Surjections ===&lt;br /&gt;
In the twelvefold way, we discuss the counting problems incurred by the mappings &amp;lt;math&amp;gt;f:N\rightarrow M&amp;lt;/math&amp;gt;. The basic case is that elements from both &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;M&amp;lt;/math&amp;gt; are distinguishable. In this case, it is easy to count the number of arbitrary mappings (which is &amp;lt;math&amp;gt;m^n&amp;lt;/math&amp;gt;) and the number of injective (one-to-one) mappings (which is &amp;lt;math&amp;gt;(m)_n&amp;lt;/math&amp;gt;), but the number of surjective is difficult. Here we apply the principle of inclusion-exclusion to count the number of surjective (onto) mappings.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:The number of surjective mappings from an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set to an &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-set is given by&lt;br /&gt;
::&amp;lt;math&amp;gt;\sum_{k=1}^m(-1)^{m-k}{m\choose k}k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;U=\{f:[n]\rightarrow[m]\}&amp;lt;/math&amp;gt; be the set of mappings from &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt; to &amp;lt;math&amp;gt;[m]&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;|U|=m^n&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
For &amp;lt;math&amp;gt;i\in[m]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; be the set of mappings &amp;lt;math&amp;gt;f:[n]\rightarrow[m]&amp;lt;/math&amp;gt; that none of &amp;lt;math&amp;gt;j\in[n]&amp;lt;/math&amp;gt; is mapped to &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;, i.e. &amp;lt;math&amp;gt;A_i=\{f:[n]\rightarrow[m]\setminus\{i\}\}&amp;lt;/math&amp;gt;, thus &amp;lt;math&amp;gt;|A_i|=(m-1)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
More generally, for &amp;lt;math&amp;gt;I\subseteq [m]&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;A_I=\bigcap_{i\in I}A_i&amp;lt;/math&amp;gt; contains the mappings &amp;lt;math&amp;gt;f:[n]\rightarrow[m]\setminus I&amp;lt;/math&amp;gt;. And &amp;lt;math&amp;gt;|A_I|=(m-|I|)^n\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
A mapping &amp;lt;math&amp;gt;f:[n]\rightarrow[m]&amp;lt;/math&amp;gt; is surjective if &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; lies in none of &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt;. By the principle of inclusion-exclusion, the number of surjective &amp;lt;math&amp;gt;f:[n]\rightarrow[m]&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{I\subseteq[m]}(-1)^{|I|}\left|A_I\right|=\sum_{I\subseteq[m]}(-1)^{|I|}(m-|I|)^n=\sum_{j=0}^m(-1)^j{m\choose j}(m-j)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Let &amp;lt;math&amp;gt;k=m-j&amp;lt;/math&amp;gt;. The theorem is proved.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Recall that, in the twelvefold way, we establish a relation between surjections and partitions.&lt;br /&gt;
&lt;br /&gt;
* Surjection to ordered partition:&lt;br /&gt;
:For a surjective &amp;lt;math&amp;gt;f:[n]\rightarrow[m]&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;(f^{-1}(0),f^{-1}(1),\ldots,f^{-1}(m-1))&amp;lt;/math&amp;gt; is an &#039;&#039;&#039;ordered partition&#039;&#039;&#039; of &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Ordered partition to surjection:&lt;br /&gt;
:For an ordered &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-partition &amp;lt;math&amp;gt;(B_0,B_1,\ldots, B_{m-1})&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt;, we can define a function &amp;lt;math&amp;gt;f:[n]\rightarrow[m]&amp;lt;/math&amp;gt; by letting &amp;lt;math&amp;gt;f(i)=j&amp;lt;/math&amp;gt; if and only if &amp;lt;math&amp;gt;i\in B_j&amp;lt;/math&amp;gt;. &amp;lt;math&amp;gt;f&amp;lt;/math&amp;gt; is surjective since as a partition, none of &amp;lt;math&amp;gt;B_i&amp;lt;/math&amp;gt; is empty.&lt;br /&gt;
&lt;br /&gt;
Therefore, we have a one-to-one correspondence between surjective mappings from an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set to an &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-set and the ordered &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-partitions of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set.&lt;br /&gt;
&lt;br /&gt;
The Stirling number of the second kind &amp;lt;math&amp;gt;S(n,m)&amp;lt;/math&amp;gt; is the number of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-partitions of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. There are &amp;lt;math&amp;gt;m!&amp;lt;/math&amp;gt; ways to order an &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;-partition, thus the number of surjective mappings &amp;lt;math&amp;gt;f:[n]\rightarrow[m]&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;m! S(n,m)&amp;lt;/math&amp;gt;. Combining with what we have proved for surjections, we give the following result for the Stirling number of the second kind.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:&amp;lt;math&amp;gt;S(n,m)=\frac{1}{m!}\sum_{k=1}^m(-1)^{m-k}{m\choose k}k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Derangements ===&lt;br /&gt;
We now count the number of bijections from a set to itself with no fixed points. This is the &#039;&#039;&#039;derangement problem&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
For a permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, a &#039;&#039;&#039;fixed point&#039;&#039;&#039; is such an &amp;lt;math&amp;gt;i\in\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;\pi(i)=i&amp;lt;/math&amp;gt;.&lt;br /&gt;
A [http://en.wikipedia.org/wiki/Derangement &#039;&#039;&#039;derangement&#039;&#039;&#039;] of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; is a permutation of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; that has no fixed points.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:The number of derangements of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; given by&lt;br /&gt;
::&amp;lt;math&amp;gt;n!\sum_{k=0}^n\frac{(-1)^k}{k!}\approx \frac{n!}{\mathrm{e}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; be the set of all permutations of &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt;. So &amp;lt;math&amp;gt;|U|=n!&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; be the set of permutations with fixed point &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;; so &amp;lt;math&amp;gt;|A_i|=(n-1)!&amp;lt;/math&amp;gt;. More generally, for any &amp;lt;math&amp;gt;I\subseteq \{1,2,\ldots,n\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;A_I=\bigcap_{i\in I}A_i&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;|A_I|=(n-|I|)!&amp;lt;/math&amp;gt;, since permutations in &amp;lt;math&amp;gt;A_I&amp;lt;/math&amp;gt; fix every point in &amp;lt;math&amp;gt;I&amp;lt;/math&amp;gt; and permute the remaining points arbitrarily. A permutation is a derangement if and only if it lies in none of the sets &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt;. So the number of derangements is&lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{I\subseteq\{1,2,\ldots,n\}}(-1)^{|I|}(n-|I|)!=\sum_{k=0}^n(-1)^k{n\choose k}(n-k)!=n!\sum_{k=0}^n\frac{(-1)^k}{k!}.&amp;lt;/math&amp;gt;&lt;br /&gt;
By Taylor&#039;s series,&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{\mathrm{e}}=\sum_{k=0}^\infty\frac{(-1)^k}{k!}=\sum_{k=0}^n\frac{(-1)^k}{k!}\pm o\left(\frac{1}{n!}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is not hard to see that &amp;lt;math&amp;gt;n!\sum_{k=0}^n\frac{(-1)^k}{k!}&amp;lt;/math&amp;gt; is the closest integer to &amp;lt;math&amp;gt;\frac{n!}{\mathrm{e}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Therefore, there are about &amp;lt;math&amp;gt;\frac{1}{\mathrm{e}}&amp;lt;/math&amp;gt; fraction of all permutations with no fixed points.&lt;br /&gt;
&lt;br /&gt;
=== Permutations with restricted positions ===&lt;br /&gt;
We introduce a general theory of counting permutations with restricted positions. In the derangement problem, we count the number of permutations that &amp;lt;math&amp;gt;\pi(i)\neq i&amp;lt;/math&amp;gt;. We now generalize to the problem of counting permutations which avoid a set of arbitrarily specified positions. &lt;br /&gt;
&lt;br /&gt;
It is traditionally described using terminology from the game of chess. Let &amp;lt;math&amp;gt;B\subseteq \{1,\ldots,n\}\times \{1,\ldots,n\}&amp;lt;/math&amp;gt;, called a &#039;&#039;&#039;board&#039;&#039;&#039;.  As illustrated below, we can think of &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; as a chess board, with the positions in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; marked by &amp;quot;&amp;lt;math&amp;gt;\times&amp;lt;/math&amp;gt;&amp;quot;.&lt;br /&gt;
{{Chess diagram small&lt;br /&gt;
| &lt;br /&gt;
| &lt;br /&gt;
|=&lt;br /&gt;
 8 |__|xx|xx|__|xx|__|__|xx|=&lt;br /&gt;
 7 |xx|__|__|xx|__|__|xx|__|=&lt;br /&gt;
 6 |xx|__|xx|xx|__|xx|xx|__|=&lt;br /&gt;
 5 |__|xx|__|__|xx|__|xx|__|=&lt;br /&gt;
 4 |xx|__|__|__|xx|xx|xx|__|=&lt;br /&gt;
 3 |__|xx|__|xx|__|__|__|xx|=&lt;br /&gt;
 2 |__|__|xx|__|xx|__|__|xx|=&lt;br /&gt;
 1 |xx|__|__|xx|__|xx|__|__|=&lt;br /&gt;
 a b c d e f g h&lt;br /&gt;
|&lt;br /&gt;
}}&lt;br /&gt;
For a permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\{1,\ldots,n\}&amp;lt;/math&amp;gt;, define the &#039;&#039;&#039;graph&#039;&#039;&#039; &amp;lt;math&amp;gt;G_\pi(V,E)&amp;lt;/math&amp;gt; as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G_\pi &amp;amp;= \{(i,\pi(i))\mid i\in \{1,2,\ldots,n\}\}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This can also be viewed as a set of marked positions on a chess board. Each row and each column has only one marked position, because &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; is a permutation. Thus, we can identify each &amp;lt;math&amp;gt;G_\pi&amp;lt;/math&amp;gt; as a placement of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; rooks (“城堡”，规则同中国象棋里的“车”) without attacking each other.&lt;br /&gt;
&lt;br /&gt;
For example, the following is the &amp;lt;math&amp;gt;G_\pi&amp;lt;/math&amp;gt; of such &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; that &amp;lt;math&amp;gt;\pi(i)=i&amp;lt;/math&amp;gt;.&lt;br /&gt;
{{Chess diagram small&lt;br /&gt;
| &lt;br /&gt;
| &lt;br /&gt;
|=&lt;br /&gt;
 8 |rl|__|__|__|__|__|__|__|=&lt;br /&gt;
 7 |__|rl|__|__|__|__|__|__|=&lt;br /&gt;
 6 |__|__|rl|__|__|__|__|__|=&lt;br /&gt;
 5 |__|__|__|rl|__|__|__|__|=&lt;br /&gt;
 4 |__|__|__|__|rl|__|__|__|=&lt;br /&gt;
 3 |__|__|__|__|__|rl|__|__|=&lt;br /&gt;
 2 |__|__|__|__|__|__|rl|__|=&lt;br /&gt;
 1 |__|__|__|__|__|__|__|rl|=&lt;br /&gt;
 a b c d e f g h&lt;br /&gt;
|&lt;br /&gt;
}}&lt;br /&gt;
Now define&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
N_0 &amp;amp;= \left|\left\{\pi\mid B\cap G_\pi=\emptyset\right\}\right|\\&lt;br /&gt;
r_k &amp;amp;= \mbox{number of }k\mbox{-subsets of }B\mbox{ such that no two elements have a common coordinate}\\&lt;br /&gt;
&amp;amp;=\left|\left\{S\in{B\choose k} \,\bigg|\, \forall (i_1,j_1),(i_2,j_2)\in S, i_1\neq i_2, j_1\neq j_2 \right\}\right|&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Interpreted in chess game,&lt;br /&gt;
* &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;: a set of marked positions in an &amp;lt;math&amp;gt;[n]\times [n]&amp;lt;/math&amp;gt; chess board.&lt;br /&gt;
* &amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt;: the number of ways of placing &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; non-attacking rooks on the chess board such that none of these rooks lie in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.&lt;br /&gt;
* &amp;lt;math&amp;gt;r_k&amp;lt;/math&amp;gt;: number of ways of placing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; non-attacking rooks on &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Our goal is to count &amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt; in terms of &amp;lt;math&amp;gt;r_k&amp;lt;/math&amp;gt;. This gives the number of permutations avoid all positions in a &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;N_0=\sum_{k=0}^n(-1)^kr_k(n-k)!&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
For each &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;A_i=\{\pi\mid (i,\pi(i))\in B\}&amp;lt;/math&amp;gt; be the set of permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; whose &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th position is in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt; is the number of permutations avoid all positions in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. Thus, our goal is to count the number of permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; in none of &amp;lt;math&amp;gt;A_i&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;i\in [n]&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For each &amp;lt;math&amp;gt;I\subseteq [n]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;A_I=\bigcap_{i\in I}A_i&amp;lt;/math&amp;gt;, which is the set of permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;(i,\pi(i))\in B&amp;lt;/math&amp;gt; for all &amp;lt;math&amp;gt;i\in I&amp;lt;/math&amp;gt;. Due to the principle of inclusion-exclusion,&lt;br /&gt;
:&amp;lt;math&amp;gt;N_0=\sum_{I\subseteq [n]} (-1)^{|I|}|A_I|=\sum_{k=0}^n(-1)^k\sum_{I\in{[n]\choose k}}|A_I|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The next observation is that &lt;br /&gt;
:&amp;lt;math&amp;gt;\sum_{I\in{[n]\choose k}}|A_I|=r_k(n-k)!&amp;lt;/math&amp;gt;,&lt;br /&gt;
because we can count both sides by first placing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; non-attacking rooks on &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; and placing &amp;lt;math&amp;gt;n-k&amp;lt;/math&amp;gt; additional non-attacking rooks on &amp;lt;math&amp;gt;[n]\times [n]&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;(n-k)!&amp;lt;/math&amp;gt; ways. &lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;N_0=\sum_{k=0}^n(-1)^kr_k(n-k)!&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
====Derangement problem====&lt;br /&gt;
We use the above general method to solve the derange problem again.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=\{(1,1),(2,2),\ldots,(n,n)\}&amp;lt;/math&amp;gt; as the chess board.  A derangement &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; is a placement of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; non-attacking rooks such that none of them is in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. &lt;br /&gt;
{{Chess diagram small&lt;br /&gt;
| &lt;br /&gt;
| &lt;br /&gt;
|=&lt;br /&gt;
 8 |xx|__|__|__|__|__|__|__|=&lt;br /&gt;
 7 |__|xx|__|__|__|__|__|__|=&lt;br /&gt;
 6 |__|__|xx|__|__|__|__|__|=&lt;br /&gt;
 5 |__|__|__|xx|__|__|__|__|=&lt;br /&gt;
 4 |__|__|__|__|xx|__|__|__|=&lt;br /&gt;
 3 |__|__|__|__|__|xx|__|__|=&lt;br /&gt;
 2 |__|__|__|__|__|__|xx|__|=&lt;br /&gt;
 1 |__|__|__|__|__|__|__|xx|=&lt;br /&gt;
 a b c d e f g h&lt;br /&gt;
|&lt;br /&gt;
}}&lt;br /&gt;
Clearly, the number of ways of placing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; non-attacking rooks on &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;r_k={n\choose k}&amp;lt;/math&amp;gt;. We want to count &amp;lt;math&amp;gt;N_0&amp;lt;/math&amp;gt;, which gives the number of ways of placing &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; non-attacking rooks such that none of these rooks lie in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By the above theorem&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
N_0=\sum_{k=0}^n(-1)^kr_k(n-k)!=\sum_{k=0}^n(-1)^k{n\choose k}(n-k)!=\sum_{k=0}^n(-1)^k\frac{n!}{k!}=n!\sum_{k=0}^n(-1)^k\frac{1}{k!}\approx\frac{n!}{e}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
====Problème des ménages====&lt;br /&gt;
Suppose that in a banquet, we want to seat &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; couples at a circular table, satisfying the following constraints:&lt;br /&gt;
* Men and women are in alternate places.&lt;br /&gt;
* No one sits next to his/her spouse.&lt;br /&gt;
&lt;br /&gt;
In how many ways can this be done?&lt;br /&gt;
&lt;br /&gt;
(For convenience, we assume that every seat at the table marked differently so that rotating the seats clockwise or anti-clockwise will end up with a &#039;&#039;&#039;different&#039;&#039;&#039; solution.)&lt;br /&gt;
&lt;br /&gt;
First, let the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; ladies find their seats. They may either sit at the odd numbered seats or even numbered seats, in either case, there are &amp;lt;math&amp;gt;n!&amp;lt;/math&amp;gt; different orders. Thus, there are &amp;lt;math&amp;gt;2(n!)&amp;lt;/math&amp;gt; ways to seat the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; ladies.&lt;br /&gt;
&lt;br /&gt;
After sitting the wives, we label the remaining &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; places clockwise as &amp;lt;math&amp;gt;0,1,\ldots, n-1&amp;lt;/math&amp;gt;. And a seating of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; husbands is given by a permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;[n]&amp;lt;/math&amp;gt; defined as follows. Let &amp;lt;math&amp;gt;\pi(i)&amp;lt;/math&amp;gt; be the seat of the husband of he lady sitting at the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th place.&lt;br /&gt;
&lt;br /&gt;
It is easy to see that &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; satisfies that &amp;lt;math&amp;gt;\pi(i)\neq i&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\pi(i)\not\equiv i+1\pmod n&amp;lt;/math&amp;gt;, and every permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; with these properties gives a feasible seating of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; husbands. Thus, we only need to count the number of permutations &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\pi(i)\not\equiv i, i+1\pmod n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Take &amp;lt;math&amp;gt;B=\{(0,0),(1,1),\ldots,(n-1,n-1), (0,1),(1,2),\ldots,(n-2,n-1),(n-1,0)\}&amp;lt;/math&amp;gt; as the chess board.  A permutation &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; which defines a way of seating the husbands, is a placement of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; non-attacking rooks such that none of them is in &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. &lt;br /&gt;
{{Chess diagram small&lt;br /&gt;
| &lt;br /&gt;
| &lt;br /&gt;
|=&lt;br /&gt;
 8 |xx|xx|__|__|__|__|__|__|=&lt;br /&gt;
 7 |__|xx|xx|__|__|__|__|__|=&lt;br /&gt;
 6 |__|__|xx|xx|__|__|__|__|=&lt;br /&gt;
 5 |__|__|__|xx|xx|__|__|__|=&lt;br /&gt;
 4 |__|__|__|__|xx|xx|__|__|=&lt;br /&gt;
 3 |__|__|__|__|__|xx|xx|__|=&lt;br /&gt;
 2 |__|__|__|__|__|__|xx|xx|=&lt;br /&gt;
 1 |xx|__|__|__|__|__|__|xx|=&lt;br /&gt;
 a b c d e f g h&lt;br /&gt;
|&lt;br /&gt;
}}&lt;br /&gt;
We need to compute &amp;lt;math&amp;gt;r_k&amp;lt;/math&amp;gt;, the number of ways of placing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; non-attacking rooks on &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;. For our choice of &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;r_k&amp;lt;/math&amp;gt; is the number of ways of choosing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; points, no two consecutive, from a collection of &amp;lt;math&amp;gt;2n&amp;lt;/math&amp;gt; points arranged in a circle.&lt;br /&gt;
&lt;br /&gt;
We first see how to do this in a &#039;&#039;line&#039;&#039;.&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:The number of ways of choosing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; &#039;&#039;non-consecutive&#039;&#039; objects from a collection of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; objects arranged in a &#039;&#039;line&#039;&#039;, is &amp;lt;math&amp;gt;{m-k+1\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
We draw a line of &amp;lt;math&amp;gt;m-k&amp;lt;/math&amp;gt; black points, and then insert &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; red points into the &amp;lt;math&amp;gt;m-k+1&amp;lt;/math&amp;gt; spaces between the black points (including the beginning and end).&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
&amp;amp;\sqcup \, \bullet \, \sqcup \, \bullet \, \sqcup \, \bullet \, \sqcup \, \bullet \, \sqcup \, \bullet \, \sqcup \, \bullet \, \sqcup \, \bullet \, \sqcup \\&lt;br /&gt;
&amp;amp;\qquad\qquad\qquad\quad\Downarrow\\&lt;br /&gt;
&amp;amp;\sqcup \, \bullet \,\, {\color{Red}\bullet} \, \bullet \,\, {\color{Red}\bullet} \, \bullet \, \sqcup \, \bullet \,\, {\color{Red}\bullet}\, \, \bullet \, \sqcup \, \bullet \, \sqcup \, \bullet \,\, {\color{Red}\bullet}&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
This gives us a line of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; points, and the red points specifies the chosen objects, which are non-consecutive. The mapping is 1-1 correspondence.&lt;br /&gt;
There are &amp;lt;math&amp;gt;{m-k+1\choose k}&amp;lt;/math&amp;gt; ways of placing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; red points into &amp;lt;math&amp;gt;m-k+1&amp;lt;/math&amp;gt; spaces.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
The problem of choosing non-consecutive objects in a circle can be reduced to the case that the objects are in a line.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Lemma|&lt;br /&gt;
:The number of ways of choosing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; &#039;&#039;non-consecutive&#039;&#039; objects from a collection of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; objects arranged in a &#039;&#039;circle&#039;&#039;, is &amp;lt;math&amp;gt;\frac{m}{m-k}{m-k\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;f(m,k)&amp;lt;/math&amp;gt; be the desired number; and let &amp;lt;math&amp;gt;g(m,k)&amp;lt;/math&amp;gt; be the number of ways of choosing &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; non-consecutive points from &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; points arranged in a circle, next coloring the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; points red, and then coloring one of the uncolored point blue. &lt;br /&gt;
&lt;br /&gt;
Clearly, &amp;lt;math&amp;gt;g(m,k)=(m-k)f(m,k)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
But we can also compute &amp;lt;math&amp;gt;g(m,k)&amp;lt;/math&amp;gt; as follows:&lt;br /&gt;
* Choose one of the &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; points and color it blue. This gives us &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; ways.&lt;br /&gt;
* Cut the circle to make a line of &amp;lt;math&amp;gt;m-1&amp;lt;/math&amp;gt; points by removing the blue point.&lt;br /&gt;
* Choose &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; non-consecutive points from the line of &amp;lt;math&amp;gt;m-1&amp;lt;/math&amp;gt; points and color them red. This gives &amp;lt;math&amp;gt;{m-k\choose k}&amp;lt;/math&amp;gt; ways due to the previous lemma.&lt;br /&gt;
&lt;br /&gt;
Thus, &amp;lt;math&amp;gt;g(m,k)=m{m-k\choose k}&amp;lt;/math&amp;gt;. Therefore we have the desired number &amp;lt;math&amp;gt;f(m,k)=\frac{m}{m-k}{m-k\choose k}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
By the above lemma, we have that &amp;lt;math&amp;gt;r_k=\frac{2n}{2n-k}{2n-k\choose k}&amp;lt;/math&amp;gt;. Then apply the theorem of counting permutations with restricted positions,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
N_0=\sum_{k=0}^n(-1)^kr_k(n-k)!=\sum_{k=0}^n(-1)^k\frac{2n}{2n-k}{2n-k\choose k}(n-k)!.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This gives the number of ways of seating the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; husbands &#039;&#039;after the ladies are seated&#039;&#039;. Recall that there are &amp;lt;math&amp;gt;2n!&amp;lt;/math&amp;gt; ways of seating the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; ladies. Thus, the total number of ways of seating &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; couples as required by problème des ménages is &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
2n!\sum_{k=0}^n(-1)^k\frac{2n}{2n-k}{2n-k\choose k}(n-k)!.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== The Euler totient function === &lt;br /&gt;
Two integers &amp;lt;math&amp;gt;m, n&amp;lt;/math&amp;gt; are said to be &#039;&#039;&#039;relatively prime&#039;&#039;&#039; if their greatest common diviser &amp;lt;math&amp;gt;\mathrm{gcd}(m,n)=1&amp;lt;/math&amp;gt;. For a positive integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;\phi(n)&amp;lt;/math&amp;gt; be the number of positive integers from &amp;lt;math&amp;gt;\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; that are relative prime to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This function, called the Euler &amp;lt;math&amp;gt;\phi&amp;lt;/math&amp;gt; function or &#039;&#039;&#039;the Euler totient function&#039;&#039;&#039;, is fundamental in number theory.&lt;br /&gt;
&lt;br /&gt;
We know derive a formula for this function by using the principle of inclusion-exclusion.&lt;br /&gt;
{{Theorem|Theorem (The Euler totient function)|&lt;br /&gt;
Suppose &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; is divisible by precisely &amp;lt;math&amp;gt;r&amp;lt;/math&amp;gt; different primes, denoted &amp;lt;math&amp;gt;p_1,\ldots,p_r&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;\phi(n)=n\prod_{i=1}^r\left(1-\frac{1}{p_i}\right)&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
{{Proof|&lt;br /&gt;
Let &amp;lt;math&amp;gt;U=\{1,2,\ldots,n\}&amp;lt;/math&amp;gt; be the universe. The number of positive integers from &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; which is divisible by some &amp;lt;math&amp;gt;p_{i_1},p_{i_2},\ldots,p_{i_s}\in\{p_1,\ldots,p_r\}&amp;lt;/math&amp;gt;, is &amp;lt;math&amp;gt;\frac{n}{p_{i_1}p_{i_2}\cdots p_{i_s}}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\phi(n)&amp;lt;/math&amp;gt; is the number of integers from &amp;lt;math&amp;gt;U&amp;lt;/math&amp;gt; which is not divisible by any &amp;lt;math&amp;gt;p_1,\ldots,p_r&amp;lt;/math&amp;gt;.&lt;br /&gt;
By principle of inclusion-exclusion,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\phi(n)&lt;br /&gt;
&amp;amp;=n+\sum_{k=1}^r(-1)^k\sum_{1\le i_1&amp;lt;i_2&amp;lt;\cdots &amp;lt;i_k\le n}\frac{n}{p_{i_1}p_{i_2}\cdots p_{i_k}}\\&lt;br /&gt;
&amp;amp;=n-\sum_{1\le i\le n}\frac{n}{p_i}+\sum_{1\le i&amp;lt;j\le n}\frac{n}{p_i p_j}-\sum_{1\le i&amp;lt;j&amp;lt;k\le n}\frac{n}{p_{i} p_{j} p_{k}}+\cdots + (-1)^r\frac{n}{p_{1}p_{2}\cdots p_{r}}\\&lt;br /&gt;
&amp;amp;=n\left(1-\sum_{1\le i\le n}\frac{1}{p_i}+\sum_{1\le i&amp;lt;j\le n}\frac{1}{p_i p_j}-\sum_{1\le i&amp;lt;j&amp;lt;k\le n}\frac{1}{p_{i} p_{j} p_{k}}+\cdots + (-1)^r\frac{1}{p_{1}p_{2}\cdots p_{r}}\right)\\&lt;br /&gt;
&amp;amp;=n\prod_{i=1}^n\left(1-\frac{1}{p_i}\right).&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Reference ==&lt;br /&gt;
* &#039;&#039;Stanley,&#039;&#039; Enumerative Combinatorics, Volume 1, Chapter 2.&lt;br /&gt;
* &#039;&#039;van Lin and Wilson&#039;&#039;, A course in combinatorics, Chapter 10, 15.&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3154</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3154"/>
		<updated>2010-09-12T06:32:36Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Reference */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
Suppose we wish to enumerate all subsets of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. To construct a subset, we specifies for every element of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set whether the element is chosen or not. Let us denote the choice to omit an element by &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt;, and the choice to include it by &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;. Using &amp;quot;&amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt;&amp;quot; to represent &amp;quot;OR&amp;quot;, and using the multiplication to denote &amp;quot;AND&amp;quot;, the choices of subsets of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set are expressed as&lt;br /&gt;
:&amp;lt;math&amp;gt;\underbrace{(x_0+x_1)(x_0+x_1)\cdots (x_0+x_1)}_{n\mbox{ elements}}=(x_0+x_1)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For example, when &amp;lt;math&amp;gt;n=3&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
(x_0+x_1)^3&lt;br /&gt;
&amp;amp;=x_0x_0x_0+x_0x_0x_1+x_0x_1x_0+x_0x_1x_1\\&lt;br /&gt;
&amp;amp;\quad +x_1x_0x_0+x_1x_0x_1+x_1x_1x_0+x_1x_1x_1&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
So it &amp;quot;generate&amp;quot; all subsets of the 3-set. Writing &amp;lt;math&amp;gt;1&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;, we have &amp;lt;math&amp;gt;(1+x)^3=1+3x+3x^2+x^3&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; is the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subsets of a 3-element set.&lt;br /&gt;
&lt;br /&gt;
In general, &amp;lt;math&amp;gt;(1+x)^n&amp;lt;/math&amp;gt; has the coefficients which are the number of subsets of fixed sizes of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set.&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
&lt;br /&gt;
Suppose that we have twelve balls: &amp;lt;font color=&amp;quot;red&amp;quot;&amp;gt;3 red&amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;4 blue&amp;lt;/font&amp;gt;, and &amp;lt;font color=&amp;quot;green&amp;quot;&amp;gt;5 green&amp;lt;/font&amp;gt;. Balls with the same color are indistinguishable.&lt;br /&gt;
&lt;br /&gt;
We want to determine the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls from these twelve balls, for some &amp;lt;math&amp;gt;0\le k\le 12&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The generating function of this sequence is&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
&amp;amp;\quad {\color{Red}(1+x+x^2+x^3)}{\color{Blue}(1+x+x^2+x^3+x^4)}{\color{OliveGreen}(1+x+x^2+x^3+x^4+x^5)}\\&lt;br /&gt;
&amp;amp;=1+3x+6x^2+10x^3+14x^4+17x^5+18x^6+17x^7+14x^8+10x^9+6x^{10}+3x^{11}+x^{12}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; gives the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls.&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
=== Example: Quicksort ===&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Reference ==&lt;br /&gt;
* &#039;&#039;Graham, Knuth, and Patashnik&#039;&#039;, Concrete Mathematics: A Foundation for Computer Science, Chapter 7.&lt;br /&gt;
* &#039;&#039;Cameron&#039;&#039;, Combinatorics: Topics, Techniques, Algorithms, Chapter 4.&lt;br /&gt;
* &#039;&#039;van Lin and Wilson&#039;&#039;, A course in combinatorics, Chapter 14.&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3153</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3153"/>
		<updated>2010-09-12T06:32:05Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
Suppose we wish to enumerate all subsets of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. To construct a subset, we specifies for every element of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set whether the element is chosen or not. Let us denote the choice to omit an element by &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt;, and the choice to include it by &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;. Using &amp;quot;&amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt;&amp;quot; to represent &amp;quot;OR&amp;quot;, and using the multiplication to denote &amp;quot;AND&amp;quot;, the choices of subsets of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set are expressed as&lt;br /&gt;
:&amp;lt;math&amp;gt;\underbrace{(x_0+x_1)(x_0+x_1)\cdots (x_0+x_1)}_{n\mbox{ elements}}=(x_0+x_1)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For example, when &amp;lt;math&amp;gt;n=3&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
(x_0+x_1)^3&lt;br /&gt;
&amp;amp;=x_0x_0x_0+x_0x_0x_1+x_0x_1x_0+x_0x_1x_1\\&lt;br /&gt;
&amp;amp;\quad +x_1x_0x_0+x_1x_0x_1+x_1x_1x_0+x_1x_1x_1&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
So it &amp;quot;generate&amp;quot; all subsets of the 3-set. Writing &amp;lt;math&amp;gt;1&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;, we have &amp;lt;math&amp;gt;(1+x)^3=1+3x+3x^2+x^3&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; is the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subsets of a 3-element set.&lt;br /&gt;
&lt;br /&gt;
In general, &amp;lt;math&amp;gt;(1+x)^n&amp;lt;/math&amp;gt; has the coefficients which are the number of subsets of fixed sizes of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set.&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
&lt;br /&gt;
Suppose that we have twelve balls: &amp;lt;font color=&amp;quot;red&amp;quot;&amp;gt;3 red&amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;4 blue&amp;lt;/font&amp;gt;, and &amp;lt;font color=&amp;quot;green&amp;quot;&amp;gt;5 green&amp;lt;/font&amp;gt;. Balls with the same color are indistinguishable.&lt;br /&gt;
&lt;br /&gt;
We want to determine the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls from these twelve balls, for some &amp;lt;math&amp;gt;0\le k\le 12&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The generating function of this sequence is&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
&amp;amp;\quad {\color{Red}(1+x+x^2+x^3)}{\color{Blue}(1+x+x^2+x^3+x^4)}{\color{OliveGreen}(1+x+x^2+x^3+x^4+x^5)}\\&lt;br /&gt;
&amp;amp;=1+3x+6x^2+10x^3+14x^4+17x^5+18x^6+17x^7+14x^8+10x^9+6x^{10}+3x^{11}+x^{12}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; gives the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls.&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
=== Example: Quicksort ===&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
== Reference ==&lt;br /&gt;
* &#039;&#039;Graham, Knuth, and Patashnik&#039;&#039;, Concrete Mathematics: A Foundation for Computer Science, Chapter 7.&lt;br /&gt;
* &#039;&#039;Cameron&#039;&#039;, Combinatorics: Topics, Techniques, Algorithms, Chapter 4.&lt;br /&gt;
* &amp;quot;van Lin and Wilson&amp;quot;, A course in combinatorics, Chapter 14.&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3152</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3152"/>
		<updated>2010-09-12T06:24:37Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Combinations */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
Suppose we wish to enumerate all subsets of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. To construct a subset, we specifies for every element of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set whether the element is chosen or not. Let us denote the choice to omit an element by &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt;, and the choice to include it by &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;. Using &amp;quot;&amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt;&amp;quot; to represent &amp;quot;OR&amp;quot;, and using the multiplication to denote &amp;quot;AND&amp;quot;, the choices of subsets of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set are expressed as&lt;br /&gt;
:&amp;lt;math&amp;gt;\underbrace{(x_0+x_1)(x_0+x_1)\cdots (x_0+x_1)}_{n\mbox{ elements}}=(x_0+x_1)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For example, when &amp;lt;math&amp;gt;n=3&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
(x_0+x_1)^3&lt;br /&gt;
&amp;amp;=x_0x_0x_0+x_0x_0x_1+x_0x_1x_0+x_0x_1x_1\\&lt;br /&gt;
&amp;amp;\quad +x_1x_0x_0+x_1x_0x_1+x_1x_1x_0+x_1x_1x_1&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
So it &amp;quot;generate&amp;quot; all subsets of the 3-set. Writing &amp;lt;math&amp;gt;1&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;, we have &amp;lt;math&amp;gt;(1+x)^3=1+3x+3x^2+x^3&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; is the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subsets of a 3-element set.&lt;br /&gt;
&lt;br /&gt;
In general, &amp;lt;math&amp;gt;(1+x)^n&amp;lt;/math&amp;gt; has the coefficients which are the number of subsets of fixed sizes of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set.&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
&lt;br /&gt;
Suppose that we have twelve balls: &amp;lt;font color=&amp;quot;red&amp;quot;&amp;gt;3 red&amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;4 blue&amp;lt;/font&amp;gt;, and &amp;lt;font color=&amp;quot;green&amp;quot;&amp;gt;5 green&amp;lt;/font&amp;gt;. Balls with the same color are indistinguishable.&lt;br /&gt;
&lt;br /&gt;
We want to determine the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls from these twelve balls, for some &amp;lt;math&amp;gt;0\le k\le 12&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The generating function of this sequence is&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
&amp;amp;\quad {\color{Red}(1+x+x^2+x^3)}{\color{Blue}(1+x+x^2+x^3+x^4)}{\color{OliveGreen}(1+x+x^2+x^3+x^4+x^5)}\\&lt;br /&gt;
&amp;amp;=1+3x+6x^2+10x^3+14x^4+17x^5+18x^6+17x^7+14x^8+10x^9+6x^{10}+3x^{11}+x^{12}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; gives the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls.&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
=== Example: Quicksort ===&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3151</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3151"/>
		<updated>2010-09-12T06:19:53Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Combinations */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
Suppose we wish to enumerate all subsets of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. To construct a subset, we specifies for every element of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set whether the element is chosen or not. Let us denote the choice to omit an element by &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt;, and the choice to include it by &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;. Using &amp;quot;&amp;lt;math&amp;gt;+&amp;lt;/math&amp;gt;&amp;quot; to represent &amp;quot;OR&amp;quot;, and using the multiplication to denote &amp;quot;AND&amp;quot;, the choices of subsets of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set are expressed as&lt;br /&gt;
:&amp;lt;math&amp;gt;\underbrace{(x_0+x_1)(x_0+x_1)\cdots (x_0+x_1)}_{n\mbox{ elements}}=(x_0+x_1)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
For example, when &amp;lt;math&amp;gt;n=3&amp;lt;/math&amp;gt;, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
(x_0+x_1)^3&lt;br /&gt;
&amp;amp;=x_0x_0x_0+x_0x_0x_1+x_0x_1x_0+x_0x_1x_1\\&lt;br /&gt;
&amp;amp;\quad +x_1x_0x_0+x_1x_0x_1+x_1x_1x_0+x_1x_1x_1&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
So it &amp;quot;generate&amp;quot; all subsets of the 3-set. Writing &amp;lt;math&amp;gt;1&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;x_1&amp;lt;/math&amp;gt;, we have &amp;lt;math&amp;gt;(1+x)^3=1+3x+3x^2+x^3&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^k&amp;lt;/math&amp;gt; is the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-subsets of a 3-element set.&lt;br /&gt;
&lt;br /&gt;
In general, &amp;lt;math&amp;gt;(1+x)^n&amp;lt;/math&amp;gt; has the coefficients which are the number of subsets of fixed sizes of an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set.&lt;br /&gt;
&lt;br /&gt;
-----&lt;br /&gt;
&lt;br /&gt;
Suppose that we have twelve balls: &amp;lt;font color=&amp;quot;red&amp;quot;&amp;gt;3 red&amp;lt;/font&amp;gt;, &amp;lt;font color=&amp;quot;blue&amp;quot;&amp;gt;4 blue&amp;lt;/font&amp;gt;, and &amp;lt;font color=&amp;quot;green&amp;quot;&amp;gt;5 green&amp;lt;/font&amp;gt;. Balls with the same color are indistinguishable.&lt;br /&gt;
&lt;br /&gt;
We want to determine the number of ways to select &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; balls from these twelve balls, for some &amp;lt;math&amp;gt;0\le k\le 12&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
:&amp;lt;math&amp;gt;{\color{Red}(1+x+x^2+x^3)}{\color{Blue}(1+x+x^2+x^3+x^4)}{\color{OliveGreen}(1+x+x^2+x^3+x^4+x^5)}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
=== Example: Quicksort ===&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3150</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3150"/>
		<updated>2010-09-12T02:52:32Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Example: multisets */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
=== Example: Quicksort ===&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3149</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3149"/>
		<updated>2010-09-12T02:52:17Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Example: multisets */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Example: Quicksort ===&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3148</id>
		<title>Combinatorics (Fall 2010)/Generating functions</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Combinatorics_(Fall_2010)/Generating_functions&amp;diff=3148"/>
		<updated>2010-09-12T02:46:22Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.108: /* Pólya&amp;#039;s problem of changing money */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Generating Functions ==&lt;br /&gt;
In Stanley&#039;s magnificent book &#039;&#039;Enumerative Combinatorics&#039;&#039;, he comments the generating function as &amp;quot;the most useful but most difficult to understand method (for counting)&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
The solution to a counting problem is usually represented as some &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; depending a parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. Sometimes this &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is called a &#039;&#039;counting function&#039;&#039; as it is a function of the parameter &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;.  &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; can also be treated as a infinite series:&lt;br /&gt;
:&amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &#039;&#039;&#039;ordinary generating function (OGF)&#039;&#039;&#039; defined by &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
G(x)=\sum_{n\ge 0} a_nx^n.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
So &amp;lt;math&amp;gt;G(x)=a_0+a_1x+a_2x^2+\cdots&amp;lt;/math&amp;gt;. An expression in this form is called a [http://en.wikipedia.org/wiki/Formal_power_series &#039;&#039;&#039;formal power series&#039;&#039;&#039;], and &amp;lt;math&amp;gt;a_0,a_1,a_2,\ldots&amp;lt;/math&amp;gt; is the sequence of &#039;&#039;&#039;coefficients&#039;&#039;&#039;. &lt;br /&gt;
&lt;br /&gt;
Furthermore, the generating function can be expanded as&lt;br /&gt;
:G(x)=&amp;lt;math&amp;gt;(\underbrace{1+\cdots+1}_{a_0})+(\underbrace{x+\cdots+x}_{a_1})+(\underbrace{x^2+\cdots+x^2}_{a_2})+\cdots+(\underbrace{x^n+\cdots+x^n}_{a_n})+\cdots&amp;lt;/math&amp;gt;&lt;br /&gt;
so it indeed &amp;quot;generates&amp;quot; all the possible instances of the objects we want to count.&lt;br /&gt;
&lt;br /&gt;
Usually, we do not evaluate the generating function &amp;lt;math&amp;gt;GF(x)&amp;lt;/math&amp;gt; on any particular value. &amp;lt;math&amp;gt;x&amp;lt;/math&amp;gt; remains as a &#039;&#039;&#039;formal variable&#039;&#039;&#039; without assuming any value. The numbers that we want to count are the coefficients carried by the terms in the formal power series. So far the generating function is just another way to represent the sequence&lt;br /&gt;
:&amp;lt;math&amp;gt;(a_0,a_1,a_2,\ldots\ldots)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The true power of generating functions comes from the various algebraic operations that we can perform on these generating functions. We use an example to demonstrate this.&lt;br /&gt;
&lt;br /&gt;
=== Combinations ===&lt;br /&gt;
&lt;br /&gt;
=== Fibonacci numbers  ===&lt;br /&gt;
Consider the following counting problems.&lt;br /&gt;
* Count the number of ways that the nonnegative integer &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; can be written as a sum of ones and twos (in order).&lt;br /&gt;
: The problem asks for the number of compositions of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; with summands from &amp;lt;math&amp;gt;\{1,2\}&amp;lt;/math&amp;gt;. Formally, we are counting the number of tuples &amp;lt;math&amp;gt;(x_1,x_2,\ldots,x_k)&amp;lt;/math&amp;gt; for some &amp;lt;math&amp;gt;k\le n&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;x_i\in\{1,2\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;x_1+x_2+\cdots+x_k=n&amp;lt;/math&amp;gt;.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. We observe that a composition either starts with a 1, in which case the rest is a composition of &amp;lt;math&amp;gt;n-1&amp;lt;/math&amp;gt;; or starts with a 2, in which case the rest is a composition of &amp;lt;math&amp;gt;n-2&amp;lt;/math&amp;gt;. So we have the recursion for &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; that&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
* Count the ways to completely cover a &amp;lt;math&amp;gt;2\times n&amp;lt;/math&amp;gt; rectangle with &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; dominos without any overlaps.&lt;br /&gt;
: Dominos are identical &amp;lt;math&amp;gt;2\times 1&amp;lt;/math&amp;gt; rectangles, so that only their orientations --- vertical or horizontal matter.&lt;br /&gt;
: Let &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; be the solution. It also holds that &amp;lt;math&amp;gt;F_n=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt;. The proof is left as an exercise.&lt;br /&gt;
&lt;br /&gt;
In both problems, the solution is given by &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; which satisfies the following recursion.&lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\begin{cases}&lt;br /&gt;
F_{n-1}+F_{n-2} &amp;amp; \mbox{if }n\ge 2,\\&lt;br /&gt;
1 &amp;amp; \mbox{if }n=1\\&lt;br /&gt;
0 &amp;amp; \mbox{if }n=0.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is called the [http://en.wikipedia.org/wiki/Fibonacci_number Fibonacci number].&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
::&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)&amp;lt;/math&amp;gt;,&lt;br /&gt;
:where &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
The quantity &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; is the so-called [http://en.wikipedia.org/wiki/Golden_ratio golden ratio], a constant with some significance in mathematics and aesthetics.&lt;br /&gt;
&lt;br /&gt;
We now prove this theorem by using generating functions.&lt;br /&gt;
The ordinary generating function for the Fibonacci number &amp;lt;math&amp;gt;F_{n}&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}F_n x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
We have that &amp;lt;math&amp;gt;F_{n}=F_{n-1}+F_{n-2}&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;n\ge 2&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
G(x) &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}F_n x^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
x+\sum_{n\ge 2}(F_{n-1}+F_{n-2})x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
For generating functions, there are general ways to generate &amp;lt;math&amp;gt;F_{n-1}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F_{n-2}&amp;lt;/math&amp;gt;, or the coefficients with any smaller indices.&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
xG(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+1}=\sum_{n\ge 1}F_{n-1} x^n=\sum_{n\ge 2}F_{n-1} x^n\\&lt;br /&gt;
x^2G(x)&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}F_n x^{n+2}=\sum_{n\ge 2}F_{n-2} x^n.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we have&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;,&lt;br /&gt;
hence&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
The value of &amp;lt;math&amp;gt;F_n&amp;lt;/math&amp;gt; is the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in the Taylor series for this formular, which is &amp;lt;math&amp;gt;\frac{G^{(n)}(0)}{n!}=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;. Although this expansion works in principle, the detailed calculus is rather painful.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
It is easier to expand the generating function by breaking it into two geometric series.&lt;br /&gt;
{{Theorem|Proposition|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;\phi=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. It holds that&lt;br /&gt;
::&amp;lt;math&amp;gt;\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
It is easy to verify the above equation, but to deduce it, we need some (high school) calculation.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;2&amp;quot; width=&amp;quot;100%&amp;quot; cellspacing=&amp;quot;4&amp;quot; cellpadding=&amp;quot;3&amp;quot; rules=&amp;quot;all&amp;quot; style=&amp;quot;margin:1em 1em 1em 0; border:solid 1px #AAAAAA; border-collapse:collapse;empty-cells:show;&amp;quot;&lt;br /&gt;
|&lt;br /&gt;
:{|&lt;br /&gt;
|&lt;br /&gt;
&amp;lt;math&amp;gt;1-x-x^2&amp;lt;/math&amp;gt; has two roots &amp;lt;math&amp;gt;\frac{-1\pm\sqrt{5}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Denote that &amp;lt;math&amp;gt;\phi=\frac{2}{-1+\sqrt{5}}=\frac{1+\sqrt{5}}{2}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\hat{\phi}=\frac{2}{-1-\sqrt{5}}=\frac{1-\sqrt{5}}{2}&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Then &amp;lt;math&amp;gt;(1-x-x^2)=(1-\phi x)(1-\hat{\phi}x)&amp;lt;/math&amp;gt;, so we can write &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\frac{x}{1-x-x^2}&lt;br /&gt;
&amp;amp;=\frac{x}{(1-\phi x)(1-\hat{\phi} x)}\\&lt;br /&gt;
&amp;amp;=\frac{\alpha}{(1-\phi x)}+\frac{\beta}{(1-\hat{\phi} x)},&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta&amp;lt;/math&amp;gt; satisfying that&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{cases}&lt;br /&gt;
\alpha+\beta=0\\&lt;br /&gt;
\alpha\phi+\beta\hat{\phi}= -1.&lt;br /&gt;
\end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
Solving this we have that &amp;lt;math&amp;gt;\alpha=\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\beta=-\frac{1}{\sqrt{5}}&amp;lt;/math&amp;gt;. Thus,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
:&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
Note that the expression &amp;lt;math&amp;gt;\frac{1}{1-z}&amp;lt;/math&amp;gt; has a well known geometric expansion:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-z}=\sum_{n\ge 0}z^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Therefore, &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; can be expanded as&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\phi x}-\frac{1}{\sqrt{5}}\cdot\frac{1}{1-\hat{\phi} x}\\&lt;br /&gt;
&amp;amp;=\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\phi x)^n-\frac{1}{\sqrt{5}}\sum_{n\ge 0}(\hat{\phi} x)^n\\&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)x^n.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
So the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Fibonacci number is given by &lt;br /&gt;
:&amp;lt;math&amp;gt;F_n=\frac{1}{\sqrt{5}}\left(\phi^n-\hat{\phi}^n\right)=\frac{1}{\sqrt{5}}\left(\frac{1+\sqrt{5}}{2}\right)^n-\frac{1}{\sqrt{5}}\left(\frac{1-\sqrt{5}}{2}\right)^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Solving recurrences ==&lt;br /&gt;
The following steps describe a general methodology of solving recurrences by generating functions.&lt;br /&gt;
:1. Give a recursion that computes &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;. In the case of Fibonacci sequence&lt;br /&gt;
::&amp;lt;math&amp;gt;a_n=a_{n-1}+a_{n-2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:2. Multiply both sides of the equation by &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; and sum over all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;. This gives the generating function&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}a_nx^n=\sum_{n\ge 0}(a_{n-1}+a_{n-2})x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
:: And manipulate the right hand side of the equation so that it becomes some other expression involving &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=x+(x+x^2)G(x)\,&amp;lt;/math&amp;gt;.&lt;br /&gt;
:3. Solve the resulting equation to derive an explicit formula for &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
::&amp;lt;math&amp;gt;G(x)=\frac{x}{1-x-x^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
:4. Expand &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; into a power series and read off the coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt;, which is a closed form for &amp;lt;math&amp;gt;a_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
The first step is usually established by combinatorial observations, or explicitly given by the problem. The third step is trivial.&lt;br /&gt;
&lt;br /&gt;
The second and the forth steps need some non-trivial analytic techniques.&lt;br /&gt;
&lt;br /&gt;
=== Algebraic operations on generating functions ===&lt;br /&gt;
The second step in the above methodology is somehow tricky. It involves first applying the recurrence to the coefficients of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is easy; and then manipulating the resulting formal power series to express it in terms of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt;, which is more difficult (because it works backwards).&lt;br /&gt;
&lt;br /&gt;
We can apply several natural algebraic operations on the formal power series.&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Generating function manipulation|&lt;br /&gt;
:Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}g_nx^n&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;F(x)=\sum_{n\ge 0}f_nx^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
x^k G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge k}g_{n-k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\frac{G(x)-\sum_{i=0}^{k-1}g_iz^i}{x^k}&lt;br /&gt;
&amp;amp;=\sum_{n\ge 0}g_{n+k}x^n, &amp;amp;\qquad (\mbox{integer }k\ge 0)\\&lt;br /&gt;
\alpha F(x)+\beta G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} (\alpha f_n+\beta g_n)x^n\\&lt;br /&gt;
F(x)G(x)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0}\sum_{k=0}^nf_kg_{n-k}x^n\\&lt;br /&gt;
G(cx)&lt;br /&gt;
&amp;amp;= \sum_{n\ge 0} c^ng_n x^n\\&lt;br /&gt;
G&#039;(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{n\ge 0}(n+1)g_{n+1}x^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
When manipulating generating functions, these rules are applied backwards; that is, from the right-hand-side to the left-hand-side.&lt;br /&gt;
&lt;br /&gt;
=== Expanding generating functions ===&lt;br /&gt;
The last step of solving recurrences by generating function is expanding the closed form generating function &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; to evaluate its &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th coefficient. In principle, we can always use the [http://en.wikipedia.org/wiki/Taylor_series Taylor series]&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}\frac{G^{(n)}(0)}{n!}x^n&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;G^{(n)}(0)&amp;lt;/math&amp;gt; is the value of the &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; evaluated at &amp;lt;math&amp;gt;x=0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Some interesting special cases are very useful.&lt;br /&gt;
&lt;br /&gt;
====Geometric sequence====&lt;br /&gt;
In the example of Fibonacci numbers, we use the well known geometric series:&lt;br /&gt;
:&amp;lt;math&amp;gt;\frac{1}{1-x}=\sum_{n\ge 0}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
It is useful when we can express the generating function in the form of &amp;lt;math&amp;gt;G(x)=\frac{a_1}{1-b_1x}+\frac{a_2}{1-b_2x}+\cdots+\frac{a_k}{1-b_kx}&amp;lt;/math&amp;gt;. The coefficient of &amp;lt;math&amp;gt;x^n&amp;lt;/math&amp;gt; in such &amp;lt;math&amp;gt;G(x)&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;a_1b_1^n+a_2b_2^n+\cdots+a_kb_k^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
====Binomial theorem====&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-th derivative of &amp;lt;math&amp;gt;(1+x)^\alpha&amp;lt;/math&amp;gt; for some real &amp;lt;math&amp;gt;\alpha&amp;lt;/math&amp;gt; is &lt;br /&gt;
:&amp;lt;math&amp;gt;\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)(1+x)^{\alpha-n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
By Taylor series, we get a generalized version of the binomial theorem known as [http://en.wikipedia.org/wiki/Binomial_coefficient#Newton.27s_binomial_series &#039;&#039;&#039;Newton&#039;s formula&#039;&#039;&#039;]:&lt;br /&gt;
{{Theorem|Newton&#039;s formular (generalized binomial theorem)|&lt;br /&gt;
If &amp;lt;math&amp;gt;|x|&amp;lt;1&amp;lt;/math&amp;gt;, then&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x)^\alpha=\sum_{n\ge 0}{\alpha\choose n}x^{n}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where &amp;lt;math&amp;gt;{\alpha\choose n}&amp;lt;/math&amp;gt; is the &#039;&#039;&#039;generalized binomial coefficient&#039;&#039;&#039; defined by &lt;br /&gt;
:&amp;lt;math&amp;gt;{\alpha\choose n}=\frac{\alpha(\alpha-1)(\alpha-2)\cdots(\alpha-n+1)}{n!}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
=== Example: multisets ===&lt;br /&gt;
In the last lecture we gave a combinatorial proof of the number of &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;-multisets on an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-set. Now we give a generating function approach to the problem.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;S=\{x_1,x_2,\ldots,x_n\}&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;-element set. We have&lt;br /&gt;
:&amp;lt;math&amp;gt;(1+x_1+x_1^2+\cdots)(1+x_2+x_2^2+\cdots)\cdots(1+x_n+x_n^2+\cdots)=\sum_{m:S\rightarrow\mathbb{N}} \prod_{x_i\in S}x_i^{m(x_i)}&amp;lt;/math&amp;gt;,&lt;br /&gt;
where each &amp;lt;math&amp;gt;m:S\rightarrow\mathbb{N}&amp;lt;/math&amp;gt; species a possible multiset on &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; with multiplicity function &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let all &amp;lt;math&amp;gt;x_i=x&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
(1+x+x^2+\cdots)^n&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{m:S\rightarrow\mathbb{N}}x^{m(x_1)+\cdots+m(x_n)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{\text{multiset }M\text{ on }S}x^{|M|}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\sum_{k\ge 0}\left({n\choose k}\right)x^k.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the the definition of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. Our task is to evaluate &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Due to the geometric sequence and the Newton&#039;s formula&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
(1+x+x^2+\cdots)^n=(1-x)^{-n}=\sum_{k\ge 0}{-n\choose k}(-x)^k.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left({n\choose k}\right)=(-1)^k{-n\choose k}={n+k-1\choose k}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
The last equation is due to the definition of the generalized binomial coefficient. We use an analytic (generating function) proof to get the same result of &amp;lt;math&amp;gt;\left({n\choose k}\right)&amp;lt;/math&amp;gt; as the combinatorial proof.&lt;br /&gt;
&lt;br /&gt;
== Catalan Number ==&lt;br /&gt;
We now introduce a class of counting problems, all with the same solution, called [http://en.wikipedia.org/wiki/Catalan_number &#039;&#039;&#039;Catalan number&#039;&#039;&#039;]. &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt;th Catalan number is denoted as &amp;lt;math&amp;gt;C_n&amp;lt;/math&amp;gt;.&lt;br /&gt;
In Volume 2 of Stanley&#039;s &#039;&#039;Enumerative Combinatorics&#039;&#039;, a set of exercises describe 66 different interpretations of the Catalan numbers. We give a few examples, cited from Wikipedia.&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;Dyck words&#039;&#039;&#039; of length 2&#039;&#039;n&#039;&#039;. A Dyck word is a string consisting of &#039;&#039;n&#039;&#039; X&#039;s and &#039;&#039;n&#039;&#039; Y&#039;s such that no initial segment of the string has more Y&#039;s than X&#039;s (see also [http://en.wikipedia.org/wiki/Dyck_language Dyck language]). For example, the following are the Dyck words of length 6:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; XXXYYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXXYY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XYXYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYYXY &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; XXYXYY.&amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Re-interpreting the symbol X as an open parenthesis and Y as a close parenthesis, &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; counts the number of expressions containing &#039;&#039;n&#039;&#039; pairs of parentheses which are correctly matched:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;big&amp;gt; ((())) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()(()) &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; ()()() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (())() &amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp; (()()) &amp;lt;/big&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 factors can be completely parenthesized (or the number of ways of associating &#039;&#039;n&#039;&#039; applications of a &#039;&#039;&#039;binary operator&#039;&#039;&#039;). For &#039;&#039;n&#039;&#039; = 3, for example, we have the following five different parenthesizations of four factors:&lt;br /&gt;
&amp;lt;div class=&amp;quot;center&amp;quot;&amp;gt;&amp;lt;math&amp;gt;((ab)c)d \quad (a(bc))d \quad(ab)(cd) \quad a((bc)d) \quad a(b(cd))&amp;lt;/math&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Successive applications of a binary operator can be represented in terms of a &#039;&#039;&#039;full binary tree&#039;&#039;&#039;. (A rooted binary tree is &#039;&#039;full&#039;&#039; if every vertex has either two children or no children.) It follows that &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of full binary trees with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;1 leaves:&lt;br /&gt;
[[Image:Catalan number binary tree example.png|center]] &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of &#039;&#039;&#039;monotonic paths&#039;&#039;&#039; along the edges of a grid with &#039;&#039;n&#039;&#039; × &#039;&#039;n&#039;&#039; square cells, which do not pass above the diagonal. A monotonic path is one which starts in the lower left corner, finishes in the upper right corner, and consists entirely of edges pointing rightwards or upwards. Counting such paths is equivalent to counting Dyck words: X stands for &amp;quot;move right&amp;quot; and Y stands for &amp;quot;move up&amp;quot;. The following diagrams show the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan number 4x4 grid example.svg.png|450px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of different ways a [http://en.wikipedia.org/wiki/Convex_polygon &#039;&#039;&#039;convex polygon&#039;&#039;&#039;] with &#039;&#039;n&#039;&#039;&amp;amp;nbsp;+&amp;amp;nbsp;2 sides can be cut into &#039;&#039;&#039;triangles&#039;&#039;&#039; by connecting vertices with straight lines. The following hexagons illustrate the case &#039;&#039;n&#039;&#039; = 4:&lt;br /&gt;
[[Image:Catalan-Hexagons-example.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of [http://en.wikipedia.org/wiki/Stack_(data_structure) &#039;&#039;&#039;stack&#039;&#039;&#039;]-sortable permutations of {1, ..., &#039;&#039;n&#039;&#039;}. A permutation &#039;&#039;w&#039;&#039; is called &#039;&#039;&#039;stack-sortable&#039;&#039;&#039; if &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;(1,&amp;amp;nbsp;...,&amp;amp;nbsp;&#039;&#039;n&#039;&#039;), where &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) is defined recursively as follows: write &#039;&#039;w&#039;&#039; =&amp;amp;nbsp;&#039;&#039;unv&#039;&#039; where &#039;&#039;n&#039;&#039; is the largest element in &#039;&#039;w&#039;&#039; and &#039;&#039;u&#039;&#039; and &#039;&#039;v&#039;&#039; are shorter sequences, and set &#039;&#039;S&#039;&#039;(&#039;&#039;w&#039;&#039;) =&amp;amp;nbsp;&#039;&#039;S&#039;&#039;(&#039;&#039;u&#039;&#039;)&#039;&#039;S&#039;&#039;(&#039;&#039;v&#039;&#039;)&#039;&#039;n&#039;&#039;, with &#039;&#039;S&#039;&#039; being the identity for one-element sequences. &lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;C&#039;&#039;&amp;lt;sub&amp;gt;&#039;&#039;n&#039;&#039;&amp;lt;/sub&amp;gt; is the number of ways to tile a stairstep shape of height &#039;&#039;n&#039;&#039; with &#039;&#039;n&#039;&#039; rectangles. The following figure illustrates the case &#039;&#039;n&#039;&#039;&amp;amp;nbsp;=&amp;amp;nbsp;4:&lt;br /&gt;
[[Image:Catalan stairsteps 4.png|400px|center]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{Theorem|Recurrence relation for Catalan numbers|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;C_1=1&amp;lt;/math&amp;gt;, and for &amp;lt;math&amp;gt;n&amp;gt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;&lt;br /&gt;
C_n=\sum_{i=1}^{n-1}C_iC_{n-i}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n&amp;lt;/math&amp;gt; be the generating function. Apply the product rule,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)^2=\sum_{n\ge 0}\sum_{k=0}^{n}C_kC_{n-k}x^n=\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n&amp;lt;/math&amp;gt;.&lt;br /&gt;
Due to the recurrence,&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\sum_{n\ge 0}C_nx^n=x+\sum_{n\ge 2}\sum_{k=1}^{n-1}C_kC_{n-k}x^n=x+G(x)^2&amp;lt;/math&amp;gt;.&lt;br /&gt;
Solving this, we obtain&lt;br /&gt;
:&amp;lt;math&amp;gt;G(x)=\frac{1\pm(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Because &amp;lt;math&amp;gt;C_0=0&amp;lt;/math&amp;gt;, it must hold that &amp;lt;math&amp;gt;G(x)=\frac{1-(1-4x)^{1/2}}{2}&amp;lt;/math&amp;gt;, or otherwise the constant term is not zero. Expanding &amp;lt;math&amp;gt;(1-4x)^{1/2}&amp;lt;/math&amp;gt; by Newton&#039;s formula, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
G(x)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\frac{1-(1-4x)^{1/2}}{2}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1-\frac{1}{2}\sum_{n\ge 0}{1/2\choose n}(-4x)^n&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
C_n&lt;br /&gt;
&amp;amp;=-\frac{1}{2}{1/2\choose n}(-4)^n\\&lt;br /&gt;
&amp;amp;=-\frac{1}{2}\cdot\frac{1}{2}\cdot\frac{-1}{2}\cdot\frac{-3}{2}\cdots\frac{-(2n-3)}{2}\cdot(-4)^n/n!\\&lt;br /&gt;
&amp;amp;=\frac{(2n-2)!}{(n-1)!n!}\\&lt;br /&gt;
&amp;amp;=\frac{1}{n}{2n-2\choose n-1}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
So we prove the following closed form for Catalan number.&lt;br /&gt;
{{Theorem|Theorem|&lt;br /&gt;
:&amp;lt;math&amp;gt;C_n=\frac{1}{n}{2n-2\choose n-1}&amp;lt;/math&amp;gt;.&lt;br /&gt;
}}&lt;/div&gt;</summary>
		<author><name>172.21.1.108</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Randomized_Algorithms_(Spring_2010)/Tail_inequalities&amp;diff=1328</id>
		<title>Randomized Algorithms (Spring 2010)/Tail inequalities</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Randomized_Algorithms_(Spring_2010)/Tail_inequalities&amp;diff=1328"/>
		<updated>2010-02-23T06:43:30Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.247: /* The Chernoff bound */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Select the Median ==&lt;br /&gt;
&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Selection_algorithm selection problem] is the problem of finding the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;th smallest element in a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. A typical case of selection problem is finding the &#039;&#039;&#039;median&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Definition&#039;&#039;&#039;&lt;br /&gt;
:The median of a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is the &amp;lt;math&amp;gt;(\lceil n/2\rceil)&amp;lt;/math&amp;gt;th element in the sorted order of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
The median can be found in &amp;lt;math&amp;gt;O(n\log n)&amp;lt;/math&amp;gt; time by sorting. There is a linear-time deterministic algorithm, [http://en.wikipedia.org/wiki/Selection_algorithm#Linear_general_selection_algorithm_-_.22Median_of_Medians_algorithm.22 &amp;quot;median of medians&amp;quot; algorithm], which is quite sophisticated. Here we introduce a much simpler randomized algorithm which also runs in linear time.&lt;br /&gt;
&lt;br /&gt;
=== Randomized median algorithm ===&lt;br /&gt;
The idea of this algorithm is random sampling. For a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;m\in S&amp;lt;/math&amp;gt; denote the median. We observe that if we can find two elements &amp;lt;math&amp;gt;d,u\in S&amp;lt;/math&amp;gt; satisfying the following properties:&lt;br /&gt;
# The median is between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in the sorted order, i.e. &amp;lt;math&amp;gt;d\le m\le u&amp;lt;/math&amp;gt;;&lt;br /&gt;
# The total number of elements between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; is small, specially for &amp;lt;math&amp;gt;C=\{x\in S\mid d\le x\le u\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;|C|=o(n/\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Provided &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; with these two properties, within linear time, we can compute the ranks of &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, construct &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, and sort &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;. Therefore, the median &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; can be picked from &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; in linear time.&lt;br /&gt;
&lt;br /&gt;
So how can we select such elements &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;? Certainly sorting &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; would give us the elements, but isn&#039;t that exactly what we want to avoid in the first place?&lt;br /&gt;
&lt;br /&gt;
Observe that &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; are only asked to roughly satisfy some constraints. This hints us maybe we can construct a &#039;&#039;sketch&#039;&#039; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; which is small enough to sort cheaply and roughly represents &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, and then pick &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; from this sketch. We construct the sketch by randomly sampling a relatively small number of elements from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Then the strategy of algorithm is outlined by:&lt;br /&gt;
* Sample a set &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of elements from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. &lt;br /&gt;
* Sort &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; and choose &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; somewhere around the median of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
* If &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; have the desirable properties, we can compute the median in linear time, or otherwise the algorithm fails.&lt;br /&gt;
&lt;br /&gt;
The parameters to be fixed are: the size of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; (small enough to sort in linear time and large enough to contain sufficient information of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;); and the order of &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; (not too close to have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; between them, and not too far away to have &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; sortable in linear time).&lt;br /&gt;
&lt;br /&gt;
We choose the size of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; are within &amp;lt;math&amp;gt;\sqrt{n}&amp;lt;/math&amp;gt; range around the median of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Randomized Median Algorithm:&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|&#039;&#039;&#039;Input:&#039;&#039;&#039; a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements over totally ordered domain.&lt;br /&gt;
# Pick a multi-set &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\left\lceil n^{3/4}\right\rceil&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, chosen independently and uniformly at random with replacement, and sort &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
# Let &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;\left\lfloor\frac{1}{2}n^{3/4}-\sqrt{n}\right\rfloor&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;\left\lceil\frac{1}{2}n^{3/4}+\sqrt{n}\right\rceil&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
# Construct &amp;lt;math&amp;gt;C=\{x\in S\mid d\le x\le u\}&amp;lt;/math&amp;gt; and compute the ranks &amp;lt;math&amp;gt;r_d=|\{x\in S\mid x&amp;lt;d\}|&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r_u=|\{x\in S\mid x&amp;lt;u\}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
# If &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;r_u&amp;lt;\frac{n}{2}&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;|C|&amp;gt;4n^{3/4}&amp;lt;/math&amp;gt; then return FAIL.&lt;br /&gt;
# Sort &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; and return the &amp;lt;math&amp;gt;\left(\left\lfloor\frac{n}{2}\right\rfloor-r_d+1\right)&amp;lt;/math&amp;gt;th element in the sorted order of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&amp;quot;Sample with replacement&amp;quot; (有放回采样) means that after sampling an element, we put the element back to the set. In this way, each sampled element is independently and identically distributed (&#039;&#039;i.i.d&#039;&#039;) (独立同分布). In the above algorithm, this is for our convenience of analysis.&lt;br /&gt;
&lt;br /&gt;
=== Analysis ===&lt;br /&gt;
The algorithm always terminates in linear time because each line of the algorithm costs at most linear time. The last three line guarantees that the algorithm returns the correct median if it does not fail.&lt;br /&gt;
&lt;br /&gt;
We then only need to bound the probability that the algorithm returns a FAIL. Let &amp;lt;math&amp;gt;m\in S&amp;lt;/math&amp;gt; be the median of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. By Line 4, we know that the algorithm returns a FAIL if and only if at least one of the following events occurs:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_1: Y=|\{x\in R\mid x\le m\}|&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_2: Z=|\{x\in R\mid x\ge m\}|&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3: |C|&amp;gt;4n^{3/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt; directly follows the third condition in Line 4. &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt; are a bit tricky. The first condition in Line 4 is that &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt;, which looks not exactly the same as &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt;, but both &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; and that &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt; are equivalent to the same event: the &amp;lt;math&amp;gt;\left\lfloor\frac{1}{2}n^{3/4}-\sqrt{n}\right\rfloor&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; is greater than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;, thus they are actually equivalent. Similarly, &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt; is equivalent to the second condition of Line 4.&lt;br /&gt;
&lt;br /&gt;
We now bound the probabilities of these events one by one.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 1&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_1]\le \frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; Let &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th sampled element in Line 1 of the algorithm. Let &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; be a indicator random variable such that &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1 &amp;amp; \mbox{if }X_i\le m,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
It is obvious that &amp;lt;math&amp;gt;Y=\sum_{i=1}^{n^{3/4}}Y_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is as defined in &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt;. For every &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;\left\lceil\frac{n}{2}\right\rceil&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; that are less than or equal to the median. The probability that &amp;lt;math&amp;gt;Y_i=1&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
p=\Pr[Y_i=1]=\Pr[X_i\le m]=\frac{1}{n}\left\lceil\frac{n}{2}\right\rceil,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is within the range of &amp;lt;math&amp;gt;\left[\frac{1}{2},\frac{1}{2}+\frac{1}{2n}\right]&amp;lt;/math&amp;gt;. Thus&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y]=n^{3/4}p\ge \frac{1}{2}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The event &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;Y&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt;&#039;s are Bernoulli trials, and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is the sum of &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt; Bernoulli trials, which follows binomial distribution with parameters &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. Thus, the variance is &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{Var}[Y]=n^{3/4}p(1-p)\le \frac{1}{4}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Applying Chebyshev&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}_1]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[Y&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\Pr\left[|Y-\mathbf{E}[Y]|&amp;gt;\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[Y]}{n}\\&lt;br /&gt;
&amp;amp;\le\frac{1}{4}n^{-1/4}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By a similar analysis, we can obtain the following bound for the event &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 2&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_2]\le \frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
We now bound the probability of the event &amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 3&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_3]\le \frac{1}{2}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; The event &amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;|C|&amp;gt;4 n^{3/4}&amp;lt;/math&amp;gt;, which by the Pigeonhole Principle, implies that at leas one of the following must be true:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&amp;lt;/math&amp;gt;: at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is greater than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&#039;&amp;lt;/math&amp;gt;: at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is smaller than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the probability that &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&amp;lt;/math&amp;gt; occurs; the second will have the same bound by symmetry.&lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is the region in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;. If there are at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; greater than the median &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, then the rank of &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in the sorted order of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; must be at least &amp;lt;math&amp;gt;\frac{1}{2}n+2n^{3/4}&amp;lt;/math&amp;gt; and thus &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; has at least &amp;lt;math&amp;gt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt; samples among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X_i\in\{0,1\}&amp;lt;/math&amp;gt; indicate whether the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th sample is among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X=\sum_{i=1}^{n^{3/4}}X_i&amp;lt;/math&amp;gt; be the number of samples in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
It holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;p=\Pr[X_i=1]=\frac{\frac{1}{2}n-2n^{3/4}}{n}=\frac{1}{2}-2n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is a binomial random variable with &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[X]=n^{3/4}p=\frac{1}{2}n^{3/4}-2\sqrt{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{Var}[X]=n^{3/4}p(1-p)=\frac{1}{4}n^{3/4}-4n^{1/4}&amp;lt;\frac{1}{4}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying Chebyshev&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}_3&#039;]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[X\ge\frac{1}{2}n^{3/4}-\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\Pr\left[|X-\mathbf{E}[X]|\ge\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[X]}{n}\\&lt;br /&gt;
&amp;amp;\le\frac{1}{4}n^{-1/4}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Symmetrically, we have that &amp;lt;math&amp;gt;\Pr[\mathcal{E}_3&#039;&#039;]\le\frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Applying the union bound&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_3]\le \Pr[\mathcal{E}_3&#039;]+\Pr[\mathcal{E}_3&#039;&#039;]\le\frac{1}{2}n^{-1/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Combining the three bounds. Applying the union bound to them, the probability that the algorithm returns a FAIL is at most &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\mathcal{E}_1]+\Pr[\mathcal{E}_2]+\Pr[\mathcal{E}_3]\le n^{-1/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Therefore the algorithm always terminates in linear time and returns the correct median with high probability.&lt;br /&gt;
&lt;br /&gt;
== Chernoff Bound ==&lt;br /&gt;
Suppose that we have a fair coin. If we toss it once, then the outcome is completely unpredictable. But if we toss it, say for 1000 times, then the number of HEADs is very likely to be around 500. This striking phenomenon is called the &#039;&#039;&#039;concentration&#039;&#039;&#039;. The Chernoff bound captures the concentration of independent trials.&lt;br /&gt;
&lt;br /&gt;
The Chernoff bound is also a tail bound for the sum of independent random variables which may give us &#039;&#039;exponentially&#039;&#039; sharp bounds.&lt;br /&gt;
&lt;br /&gt;
Before proving the Chernoff bound, we should talk about the moment generating functions.&lt;br /&gt;
&lt;br /&gt;
=== Moment generating functions ===&lt;br /&gt;
The more we know about the moments of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, the more information we would have about &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. There is a so-called &#039;&#039;&#039;moment generating function&#039;&#039;&#039;, which &amp;quot;packs&amp;quot; all the information about the moments of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; into one function.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Definition:&#039;&#039;&#039;&lt;br /&gt;
:The moment generating function of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is defined as &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt; is the parameter of the function.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
By Taylor&#039;s expansion and the linearity of expectations,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\sum_{k=0}^\infty\frac{\lambda^k}{k!}X^k\right]\\&lt;br /&gt;
&amp;amp;=\sum_{k=0}^\infty\frac{\lambda^k}{k!}\mathbf{E}\left[X^k\right]&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
&lt;br /&gt;
The moment generating function &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; is a function of &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== The Chernoff bound ===&lt;br /&gt;
The Chernoff bounds are tail inequalities with exponential decays for the sum of independent trials.&lt;br /&gt;
The bounds are obtained by applying Markov&#039;s inequality to the moment generating function of the sum of independent trials, with some  appropriate choice of the parameter &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Chernoff bound (the upper tail):&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; For any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;X\ge (1+\delta)\mu&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
where the last step follows by Markov&#039;s inequality.&lt;br /&gt;
&lt;br /&gt;
Computing the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[e^{\lambda \sum_{i=1}^n X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\prod_{i=1}^n e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right].&lt;br /&gt;
&amp;amp; (\mbox{for independent random variables})&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p_i=\Pr[X_i=1]&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;i=1,2,\ldots,n&amp;lt;/math&amp;gt;. Then,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[X]=\mathbf{E}\left[\sum_{i=1}^n X_i\right]=\sum_{i=1}^n\mathbf{E}[X_i]=\sum_{i=1}^n p_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the moment generating function for each individual &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X_i}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
p_i\cdot e^{\lambda\cdot 1}+(1-p_i)\cdot e^{\lambda\cdot 0}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1+p_i(e^\lambda -1)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
e^{p_i(e^\lambda-1)},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
where in the last step we apply the Taylor&#039;s expansion so that &amp;lt;math&amp;gt;e^y\ge 1+y&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;y=p_i(e^\lambda-1)\ge 0&amp;lt;/math&amp;gt;. (By doing this, we can transform the product to the sum of &amp;lt;math&amp;gt;p_i&amp;lt;/math&amp;gt;, which is &amp;lt;math&amp;gt;\mu&amp;lt;/math&amp;gt;.) &lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\prod_{i=1}^n e^{p_i(e^\lambda-1)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(\sum_{i=1}^n p_i(e^{\lambda}-1)\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
e^{(e^\lambda-1)\mu}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, we have shown that for any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{e^{(e^\lambda-1)\mu}}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}\right)^\mu&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;gt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The idea of the proof is actually quite clear: we apply Markov&#039;s inequality to &amp;lt;math&amp;gt;e^{\lambda X}&amp;lt;/math&amp;gt; and for the rest, we just estimate the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;. To make the bound as tight as possible, we minimized the &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt; by setting &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;lt;/math&amp;gt;, which can be justified by taking derivatives of &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
We then proceed to the lower tail, the probability that the random variable deviates below the mean value:&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Chernoff bound (the lower tail):&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; For any &amp;lt;math&amp;gt;\lambda&amp;lt;0&amp;lt;/math&amp;gt;, by the same analysis as in the upper tail version,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\le (1-\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1-\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1-\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1-\delta)}}\right)^\mu.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
For any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1-\delta)&amp;lt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
Some useful special forms of the bounds can be derived directly from the above general forms of the bounds. We now know better why we say that the bounds are exponentially sharp.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Useful forms of the Chernoff bound&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:1. for &amp;lt;math&amp;gt;0&amp;lt;\delta\le 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{3}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{2}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
:2. for &amp;lt;math&amp;gt;t\ge 2e\mu&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge t]\le 2^{-t}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; To obtain the bounds in (1), we need to show that for &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt; 1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\le e^{-\delta^2/3}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\le e^{-\delta^2/2}&amp;lt;/math&amp;gt;. We can verify both inequalities by standard analysis techniques.&lt;br /&gt;
&lt;br /&gt;
To obtain the bound in (2), let &amp;lt;math&amp;gt;t=(1+\delta)\mu&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\delta=t/\mu-1\ge 2e-1&amp;lt;/math&amp;gt;. Hence,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge(1+\delta)\mu]&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^\delta}{(1+\delta)^{(1+\delta)}}\right)^\mu\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{1+\delta}\right)^{(1+\delta)\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{2e}\right)^{2e\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
2^{-t}&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Applications of Chernoff Bounds ==&lt;br /&gt;
We now introduce some applications of Chernoff bounds in randomized algorithms.&lt;br /&gt;
&lt;br /&gt;
=== Balls into bins ===&lt;br /&gt;
Throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls uniformly and independently to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins, what is the maximum load of all bins? In the last class, by using a counting argument, we proved that for the case that &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt;, the maximum load is &amp;lt;math&amp;gt;O(\ln n\ln\ln n)&amp;lt;/math&amp;gt; with high probability. Now we show that when there are more balls, the loads are more balanced.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;j\in[m]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_{ij}&amp;lt;/math&amp;gt; be the indicator variable for the event that ball &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; is thrown to bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. Obviously&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_{ij}]=\Pr[\mbox{ball }j\mbox{ is thrown to bin }i]=\frac{1}{n}&amp;lt;/math&amp;gt;&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y_i=\sum_{j\in[m]}X_{ij}&amp;lt;/math&amp;gt; be the load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Let us consider the case when &amp;lt;math&amp;gt;m=6n\ln n&amp;lt;/math&amp;gt;. Then the expected load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[Y_i]=\mathbf{E}\left[\sum_{j\in[m]}X_{ij}\right]=\sum_{j\in[m]}\mathbf{E}[X_{ij}]=m/n=6\ln n&amp;lt;/math&amp;gt;. &lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is a sum of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; mutually independent indicator variable. Applying Chernoff bound, for any particular bin &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[Y_i&amp;gt;12\ln n] =\Pr[Y_i&amp;gt;(1+1)\mu]\le e^{-\frac{\mu}{3}} = e^{-2\ln n}= \frac{1}{n^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying the union bound, the probability that there exists a bin with load &amp;lt;math&amp;gt;&amp;gt;12\ln n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;n\cdot \Pr[Y_1&amp;gt;12\ln n]\le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, with probability at least &amp;lt;math&amp;gt;1-\frac{1}{n}&amp;lt;/math&amp;gt;, the maximum load is within &amp;lt;math&amp;gt;12\ln n=O(m/n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Set balancing ===&lt;br /&gt;
Supposed that we have an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; with 0-1 entries. We are looking for a &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt; that minimizes &amp;lt;math&amp;gt;\|Ab\|_\infty&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;\|\cdot\|_\infty&amp;lt;/math&amp;gt; is the infinity norm (also called &amp;lt;math&amp;gt;L_\infty&amp;lt;/math&amp;gt; norm) of a vector, and for the vector &amp;lt;math&amp;gt;c=Ab&amp;lt;/math&amp;gt;, &lt;br /&gt;
:&amp;lt;math&amp;gt;\|Ab\|_\infty=\max_{i=1,2,\ldots,n}|c_i|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We can also describe this problem as an optimization:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mbox{minimize }&lt;br /&gt;
&amp;amp;\quad&lt;br /&gt;
\|Ab\|_\infty\\&lt;br /&gt;
\mbox{subject to: }&lt;br /&gt;
&amp;amp;\quad&lt;br /&gt;
b\in\{-1,+1\}^m.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|The problem arises in designing statistical experiments. Suppose that we have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; &#039;&#039;&#039;subjects&#039;&#039;&#039;, each of which may have up to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;&#039;features&#039;&#039;&#039;. This gives us an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
\mbox{feature 1:}\\&lt;br /&gt;
\mbox{feature 2:}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
\mbox{feature n:}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{cccc}&lt;br /&gt;
a_{11} &amp;amp; a_{12} &amp;amp; \cdots &amp;amp; a_{1m}\\&lt;br /&gt;
a_{21} &amp;amp; a_{22} &amp;amp; \cdots &amp;amp; a_{2m}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; \vdots\\&lt;br /&gt;
a_{n1} &amp;amp; a_{n2} &amp;amp; \cdots &amp;amp; a_{nm}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right],&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where each column represents a subject and each row represent a feature. An entry &amp;lt;math&amp;gt;a_{ij}\in\{0,1\}&amp;lt;/math&amp;gt; indicates whether subject &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; has feature &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By multiplying a vector &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{cccc}&lt;br /&gt;
a_{11} &amp;amp; a_{12} &amp;amp; \cdots &amp;amp; a_{1m}\\&lt;br /&gt;
a_{21} &amp;amp; a_{22} &amp;amp; \cdots &amp;amp; a_{2m}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; \vdots\\&lt;br /&gt;
a_{n1} &amp;amp; a_{n2} &amp;amp; \cdots &amp;amp; a_{nm}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right]&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
b_{1}\\&lt;br /&gt;
b_{2}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
b_{m}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right]&lt;br /&gt;
=&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
c_{1}\\&lt;br /&gt;
c_{2}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
c_{n}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right],&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
the subjects are partitioned into two disjoint groups: one for -1 and other other for +1. Each &amp;lt;math&amp;gt;c_i&amp;lt;/math&amp;gt; gives the difference between the numbers of subjects with feature &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; in the two groups. By minimizing &amp;lt;math&amp;gt;\|Ab\|_\infty=\|c\|_\infty&amp;lt;/math&amp;gt;, we ask for an optimal partition so that each feature is roughly as balanced as possible between the two groups.&lt;br /&gt;
&lt;br /&gt;
In a scientific experiment, one of the group serves as a [http://en.wikipedia.org/wiki/Scientific_control control group] (对照组). Ideally, we want the two groups are statistically identical, which is usually impossible to achieve in practice. The requirement of minimizing &amp;lt;math&amp;gt;\|Ab\|_\infty&amp;lt;/math&amp;gt; actually means the statistical difference between the two groups are minimized.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
We propose an extremely simple &amp;quot;randomized algorithm&amp;quot; for computing a &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt;: for each &amp;lt;math&amp;gt;i=1,2,\ldots, m&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt; be independently chosen from &amp;lt;math&amp;gt;\{-1,+1\}&amp;lt;/math&amp;gt;, such that &lt;br /&gt;
:&amp;lt;math&amp;gt;b_i=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
-1 &amp;amp; \mbox{with probability }\frac{1}{2}\\&lt;br /&gt;
+1 &amp;amp;\mbox{with probability }\frac{1}{2}&lt;br /&gt;
\end{cases}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This procedure can hardly be called as an &amp;quot;algorithm&amp;quot;, because its decision is made disregard of the input &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. We then show that despite of this obliviousness, the algorithm chooses a good enough &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, such that for any &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\|Ab\|_\infty=O(\sqrt{m\ln n})&amp;lt;/math&amp;gt; with high probability.&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Theorem&#039;&#039;&#039;&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix with 0-1 entries. For a random vector &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; entries chosen independently and with equal probability from &amp;lt;math&amp;gt;\{-1,+1\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[\|Ab\|_\infty&amp;gt;2\sqrt{2m\ln n}]\le\frac{2}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039;&lt;br /&gt;
Consider particularly the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th row of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. The entry of &amp;lt;math&amp;gt;Ab&amp;lt;/math&amp;gt; contributed by row &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;c_i=\sum_{j=1}^m a_{ij}b_j&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; be the non-zero entries in the row. If &amp;lt;math&amp;gt;k\le2\sqrt{2m\ln n}&amp;lt;/math&amp;gt;, then clearly &amp;lt;math&amp;gt;|c_i|&amp;lt;/math&amp;gt; is no greater than &amp;lt;math&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt;. On the other hand if &amp;lt;math&amp;gt;k&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; then the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; nonzero terms in the sum&lt;br /&gt;
:&amp;lt;math&amp;gt;c_i=\sum_{j=1}^m a_{ij}b_j&amp;lt;/math&amp;gt;&lt;br /&gt;
are independent, each with probability 1/2 of being either +1 or -1. &lt;br /&gt;
&lt;br /&gt;
Thus, for these &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; nonzero terms, each &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt; is either positive or negative independently with equal probability. There are expectedly &amp;lt;math&amp;gt;\mu=\frac{k}{2}&amp;lt;/math&amp;gt; positive &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s among these &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; terms, and &amp;lt;math&amp;gt;c_i&amp;lt;-2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; only occurs when there are less than &amp;lt;math&amp;gt;\frac{k}{2}-\sqrt{2m\ln n}=\left(1-\delta\right)\mu&amp;lt;/math&amp;gt; positive &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s, where &amp;lt;math&amp;gt;\delta=\frac{2\sqrt{2m\ln n}}{k}&amp;lt;/math&amp;gt;. Applying Chernoff bound, this event occurs with probability at most&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\exp\left(-\frac{\mu\delta^2}{2}\right)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\frac{k}{2}\cdot\frac{8m\ln n}{2k^2}\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\frac{2m\ln n}{k}\right)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\exp\left(-\frac{2m\ln n}{m}\right)\\&lt;br /&gt;
&amp;amp;\le n^{-2}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The same argument can be applied to negative &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s, so that the probability that &amp;lt;math&amp;gt;c_i&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;n^{-2}&amp;lt;/math&amp;gt;. Therefore, by the union bound, &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[|c_i|&amp;gt; 2\sqrt{2m\ln n}]\le\frac{2}{n^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Apply the union bound to all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; rows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\|Ab\|_\infty&amp;gt;2\sqrt{2m\ln n}]\le n\cdot\Pr[|c_i|&amp;gt; 2\sqrt{2m\ln n}]\le\frac{2}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
So how good is this randomized algorithm? In fact when &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt; there exists a matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\|Ab\|_\infty=\Omega(\sqrt{n})&amp;lt;/math&amp;gt; for any choice of &amp;lt;math&amp;gt;b\in\{-1,+1\}^n&amp;lt;/math&amp;gt;.&lt;/div&gt;</summary>
		<author><name>172.21.1.247</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Randomized_Algorithms_(Spring_2010)/Tail_inequalities&amp;diff=1327</id>
		<title>Randomized Algorithms (Spring 2010)/Tail inequalities</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Randomized_Algorithms_(Spring_2010)/Tail_inequalities&amp;diff=1327"/>
		<updated>2010-02-23T06:40:38Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.247: /* The Chernoff bound */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Select the Median ==&lt;br /&gt;
&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Selection_algorithm selection problem] is the problem of finding the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;th smallest element in a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. A typical case of selection problem is finding the &#039;&#039;&#039;median&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Definition&#039;&#039;&#039;&lt;br /&gt;
:The median of a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is the &amp;lt;math&amp;gt;(\lceil n/2\rceil)&amp;lt;/math&amp;gt;th element in the sorted order of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
The median can be found in &amp;lt;math&amp;gt;O(n\log n)&amp;lt;/math&amp;gt; time by sorting. There is a linear-time deterministic algorithm, [http://en.wikipedia.org/wiki/Selection_algorithm#Linear_general_selection_algorithm_-_.22Median_of_Medians_algorithm.22 &amp;quot;median of medians&amp;quot; algorithm], which is quite sophisticated. Here we introduce a much simpler randomized algorithm which also runs in linear time.&lt;br /&gt;
&lt;br /&gt;
=== Randomized median algorithm ===&lt;br /&gt;
The idea of this algorithm is random sampling. For a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;m\in S&amp;lt;/math&amp;gt; denote the median. We observe that if we can find two elements &amp;lt;math&amp;gt;d,u\in S&amp;lt;/math&amp;gt; satisfying the following properties:&lt;br /&gt;
# The median is between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in the sorted order, i.e. &amp;lt;math&amp;gt;d\le m\le u&amp;lt;/math&amp;gt;;&lt;br /&gt;
# The total number of elements between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; is small, specially for &amp;lt;math&amp;gt;C=\{x\in S\mid d\le x\le u\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;|C|=o(n/\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Provided &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; with these two properties, within linear time, we can compute the ranks of &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, construct &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, and sort &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;. Therefore, the median &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; can be picked from &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; in linear time.&lt;br /&gt;
&lt;br /&gt;
So how can we select such elements &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;? Certainly sorting &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; would give us the elements, but isn&#039;t that exactly what we want to avoid in the first place?&lt;br /&gt;
&lt;br /&gt;
Observe that &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; are only asked to roughly satisfy some constraints. This hints us maybe we can construct a &#039;&#039;sketch&#039;&#039; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; which is small enough to sort cheaply and roughly represents &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, and then pick &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; from this sketch. We construct the sketch by randomly sampling a relatively small number of elements from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Then the strategy of algorithm is outlined by:&lt;br /&gt;
* Sample a set &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of elements from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. &lt;br /&gt;
* Sort &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; and choose &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; somewhere around the median of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
* If &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; have the desirable properties, we can compute the median in linear time, or otherwise the algorithm fails.&lt;br /&gt;
&lt;br /&gt;
The parameters to be fixed are: the size of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; (small enough to sort in linear time and large enough to contain sufficient information of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;); and the order of &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; (not too close to have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; between them, and not too far away to have &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; sortable in linear time).&lt;br /&gt;
&lt;br /&gt;
We choose the size of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; are within &amp;lt;math&amp;gt;\sqrt{n}&amp;lt;/math&amp;gt; range around the median of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Randomized Median Algorithm:&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|&#039;&#039;&#039;Input:&#039;&#039;&#039; a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements over totally ordered domain.&lt;br /&gt;
# Pick a multi-set &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\left\lceil n^{3/4}\right\rceil&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, chosen independently and uniformly at random with replacement, and sort &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
# Let &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;\left\lfloor\frac{1}{2}n^{3/4}-\sqrt{n}\right\rfloor&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;\left\lceil\frac{1}{2}n^{3/4}+\sqrt{n}\right\rceil&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
# Construct &amp;lt;math&amp;gt;C=\{x\in S\mid d\le x\le u\}&amp;lt;/math&amp;gt; and compute the ranks &amp;lt;math&amp;gt;r_d=|\{x\in S\mid x&amp;lt;d\}|&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r_u=|\{x\in S\mid x&amp;lt;u\}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
# If &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;r_u&amp;lt;\frac{n}{2}&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;|C|&amp;gt;4n^{3/4}&amp;lt;/math&amp;gt; then return FAIL.&lt;br /&gt;
# Sort &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; and return the &amp;lt;math&amp;gt;\left(\left\lfloor\frac{n}{2}\right\rfloor-r_d+1\right)&amp;lt;/math&amp;gt;th element in the sorted order of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&amp;quot;Sample with replacement&amp;quot; (有放回采样) means that after sampling an element, we put the element back to the set. In this way, each sampled element is independently and identically distributed (&#039;&#039;i.i.d&#039;&#039;) (独立同分布). In the above algorithm, this is for our convenience of analysis.&lt;br /&gt;
&lt;br /&gt;
=== Analysis ===&lt;br /&gt;
The algorithm always terminates in linear time because each line of the algorithm costs at most linear time. The last three line guarantees that the algorithm returns the correct median if it does not fail.&lt;br /&gt;
&lt;br /&gt;
We then only need to bound the probability that the algorithm returns a FAIL. Let &amp;lt;math&amp;gt;m\in S&amp;lt;/math&amp;gt; be the median of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. By Line 4, we know that the algorithm returns a FAIL if and only if at least one of the following events occurs:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_1: Y=|\{x\in R\mid x\le m\}|&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_2: Z=|\{x\in R\mid x\ge m\}|&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3: |C|&amp;gt;4n^{3/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt; directly follows the third condition in Line 4. &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt; are a bit tricky. The first condition in Line 4 is that &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt;, which looks not exactly the same as &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt;, but both &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; and that &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt; are equivalent to the same event: the &amp;lt;math&amp;gt;\left\lfloor\frac{1}{2}n^{3/4}-\sqrt{n}\right\rfloor&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; is greater than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;, thus they are actually equivalent. Similarly, &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt; is equivalent to the second condition of Line 4.&lt;br /&gt;
&lt;br /&gt;
We now bound the probabilities of these events one by one.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 1&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_1]\le \frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; Let &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th sampled element in Line 1 of the algorithm. Let &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; be a indicator random variable such that &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1 &amp;amp; \mbox{if }X_i\le m,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
It is obvious that &amp;lt;math&amp;gt;Y=\sum_{i=1}^{n^{3/4}}Y_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is as defined in &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt;. For every &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;\left\lceil\frac{n}{2}\right\rceil&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; that are less than or equal to the median. The probability that &amp;lt;math&amp;gt;Y_i=1&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
p=\Pr[Y_i=1]=\Pr[X_i\le m]=\frac{1}{n}\left\lceil\frac{n}{2}\right\rceil,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is within the range of &amp;lt;math&amp;gt;\left[\frac{1}{2},\frac{1}{2}+\frac{1}{2n}\right]&amp;lt;/math&amp;gt;. Thus&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y]=n^{3/4}p\ge \frac{1}{2}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The event &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;Y&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt;&#039;s are Bernoulli trials, and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is the sum of &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt; Bernoulli trials, which follows binomial distribution with parameters &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. Thus, the variance is &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{Var}[Y]=n^{3/4}p(1-p)\le \frac{1}{4}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Applying Chebyshev&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}_1]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[Y&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\Pr\left[|Y-\mathbf{E}[Y]|&amp;gt;\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[Y]}{n}\\&lt;br /&gt;
&amp;amp;\le\frac{1}{4}n^{-1/4}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By a similar analysis, we can obtain the following bound for the event &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 2&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_2]\le \frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
We now bound the probability of the event &amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 3&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_3]\le \frac{1}{2}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; The event &amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;|C|&amp;gt;4 n^{3/4}&amp;lt;/math&amp;gt;, which by the Pigeonhole Principle, implies that at leas one of the following must be true:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&amp;lt;/math&amp;gt;: at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is greater than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&#039;&amp;lt;/math&amp;gt;: at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is smaller than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the probability that &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&amp;lt;/math&amp;gt; occurs; the second will have the same bound by symmetry.&lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is the region in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;. If there are at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; greater than the median &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, then the rank of &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in the sorted order of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; must be at least &amp;lt;math&amp;gt;\frac{1}{2}n+2n^{3/4}&amp;lt;/math&amp;gt; and thus &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; has at least &amp;lt;math&amp;gt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt; samples among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X_i\in\{0,1\}&amp;lt;/math&amp;gt; indicate whether the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th sample is among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X=\sum_{i=1}^{n^{3/4}}X_i&amp;lt;/math&amp;gt; be the number of samples in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
It holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;p=\Pr[X_i=1]=\frac{\frac{1}{2}n-2n^{3/4}}{n}=\frac{1}{2}-2n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is a binomial random variable with &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[X]=n^{3/4}p=\frac{1}{2}n^{3/4}-2\sqrt{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{Var}[X]=n^{3/4}p(1-p)=\frac{1}{4}n^{3/4}-4n^{1/4}&amp;lt;\frac{1}{4}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying Chebyshev&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}_3&#039;]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[X\ge\frac{1}{2}n^{3/4}-\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\Pr\left[|X-\mathbf{E}[X]|\ge\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[X]}{n}\\&lt;br /&gt;
&amp;amp;\le\frac{1}{4}n^{-1/4}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Symmetrically, we have that &amp;lt;math&amp;gt;\Pr[\mathcal{E}_3&#039;&#039;]\le\frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Applying the union bound&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_3]\le \Pr[\mathcal{E}_3&#039;]+\Pr[\mathcal{E}_3&#039;&#039;]\le\frac{1}{2}n^{-1/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Combining the three bounds. Applying the union bound to them, the probability that the algorithm returns a FAIL is at most &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\mathcal{E}_1]+\Pr[\mathcal{E}_2]+\Pr[\mathcal{E}_3]\le n^{-1/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Therefore the algorithm always terminates in linear time and returns the correct median with high probability.&lt;br /&gt;
&lt;br /&gt;
== Chernoff Bound ==&lt;br /&gt;
Suppose that we have a fair coin. If we toss it once, then the outcome is completely unpredictable. But if we toss it, say for 1000 times, then the number of HEADs is very likely to be around 500. This striking phenomenon is called the &#039;&#039;&#039;concentration&#039;&#039;&#039;. The Chernoff bound captures the concentration of independent trials.&lt;br /&gt;
&lt;br /&gt;
The Chernoff bound is also a tail bound for the sum of independent random variables which may give us &#039;&#039;exponentially&#039;&#039; sharp bounds.&lt;br /&gt;
&lt;br /&gt;
Before proving the Chernoff bound, we should talk about the moment generating functions.&lt;br /&gt;
&lt;br /&gt;
=== Moment generating functions ===&lt;br /&gt;
The more we know about the moments of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, the more information we would have about &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. There is a so-called &#039;&#039;&#039;moment generating function&#039;&#039;&#039;, which &amp;quot;packs&amp;quot; all the information about the moments of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; into one function.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Definition:&#039;&#039;&#039;&lt;br /&gt;
:The moment generating function of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is defined as &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt; is the parameter of the function.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
By Taylor&#039;s expansion and the linearity of expectations,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\sum_{k=0}^\infty\frac{\lambda^k}{k!}X^k\right]\\&lt;br /&gt;
&amp;amp;=\sum_{k=0}^\infty\frac{\lambda^k}{k!}\mathbf{E}\left[X^k\right]&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
&lt;br /&gt;
The moment generating function &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; is a function of &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== The Chernoff bound ===&lt;br /&gt;
The Chernoff bounds are tail inequalities with exponential decays for the sum of independent trials.&lt;br /&gt;
The bounds are obtained by applying Markov&#039;s inequality to the moment generating function of the sum of independent trials, with some  appropriate choice of the parameter &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Chernoff bound (the upper tail):&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; For any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;X\ge (1+\delta)\mu&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
where the last step follows by Markov&#039;s inequality.&lt;br /&gt;
&lt;br /&gt;
Computing the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[e^{\lambda \sum_{i=1}^n X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\prod_{i=1}^n e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right]. &amp;amp;(X_i\mbox{&#039;s are independent})&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p_i=\Pr[X_i=1]&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;i=1,2,\ldots,n&amp;lt;/math&amp;gt;. Then,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[X]=\mathbf{E}\left[\sum_{i=1}^n X_i\right]=\sum_{i=1}^n\mathbf{E}[X_i]=\sum_{i=1}^n p_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the moment generating function for each individual &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X_i}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
p_i\cdot e^{\lambda\cdot 1}+(1-p_i)\cdot e^{\lambda\cdot 0}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1+p_i(e^\lambda -1)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
e^{p_i(e^\lambda-1)},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
where in the last step we apply the Taylor&#039;s expansion so that &amp;lt;math&amp;gt;e^y\ge 1+y&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;y=p_i(e^\lambda-1)\ge 0&amp;lt;/math&amp;gt;. (By doing this, we can transform the product to the sum of &amp;lt;math&amp;gt;p_i&amp;lt;/math&amp;gt;, which is &amp;lt;math&amp;gt;\mu&amp;lt;/math&amp;gt;.) &lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\prod_{i=1}^n e^{p_i(e^\lambda-1)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(\sum_{i=1}^n p_i(e^{\lambda}-1)\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
e^{(e^\lambda-1)\mu}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, we have shown that for any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{e^{(e^\lambda-1)\mu}}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}\right)^\mu&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;gt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The idea of the proof is actually quite clear: we apply Markov&#039;s inequality to &amp;lt;math&amp;gt;e^{\lambda X}&amp;lt;/math&amp;gt; and for the rest, we just estimate the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;. To make the bound as tight as possible, we minimized the &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt; by setting &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;lt;/math&amp;gt;, which can be justified by taking derivatives of &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
We then proceed to the lower tail, the probability that the random variable deviates below the mean value:&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Chernoff bound (the lower tail):&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; For any &amp;lt;math&amp;gt;\lambda&amp;lt;0&amp;lt;/math&amp;gt;, by the same analysis as in the upper tail version,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\le (1-\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1-\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1-\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1-\delta)}}\right)^\mu.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
For any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1-\delta)&amp;lt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
Some useful special forms of the bounds can be derived directly from the above general forms of the bounds. We now know better why we say that the bounds are exponentially sharp.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Useful forms of the Chernoff bound&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:1. for &amp;lt;math&amp;gt;0&amp;lt;\delta\le 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{3}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{2}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
:2. for &amp;lt;math&amp;gt;t\ge 2e\mu&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge t]\le 2^{-t}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; To obtain the bounds in (1), we need to show that for &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt; 1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\le e^{-\delta^2/3}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\le e^{-\delta^2/2}&amp;lt;/math&amp;gt;. We can verify both inequalities by standard analysis techniques.&lt;br /&gt;
&lt;br /&gt;
To obtain the bound in (2), let &amp;lt;math&amp;gt;t=(1+\delta)\mu&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\delta=t/\mu-1\ge 2e-1&amp;lt;/math&amp;gt;. Hence,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge(1+\delta)\mu]&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^\delta}{(1+\delta)^{(1+\delta)}}\right)^\mu\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{1+\delta}\right)^{(1+\delta)\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{2e}\right)^{2e\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
2^{-t}&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Applications of Chernoff Bounds ==&lt;br /&gt;
We now introduce some applications of Chernoff bounds in randomized algorithms.&lt;br /&gt;
&lt;br /&gt;
=== Balls into bins ===&lt;br /&gt;
Throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls uniformly and independently to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins, what is the maximum load of all bins? In the last class, by using a counting argument, we proved that for the case that &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt;, the maximum load is &amp;lt;math&amp;gt;O(\ln n\ln\ln n)&amp;lt;/math&amp;gt; with high probability. Now we show that when there are more balls, the loads are more balanced.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;j\in[m]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_{ij}&amp;lt;/math&amp;gt; be the indicator variable for the event that ball &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; is thrown to bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. Obviously&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_{ij}]=\Pr[\mbox{ball }j\mbox{ is thrown to bin }i]=\frac{1}{n}&amp;lt;/math&amp;gt;&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y_i=\sum_{j\in[m]}X_{ij}&amp;lt;/math&amp;gt; be the load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Let us consider the case when &amp;lt;math&amp;gt;m=6n\ln n&amp;lt;/math&amp;gt;. Then the expected load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[Y_i]=\mathbf{E}\left[\sum_{j\in[m]}X_{ij}\right]=\sum_{j\in[m]}\mathbf{E}[X_{ij}]=m/n=6\ln n&amp;lt;/math&amp;gt;. &lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is a sum of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; mutually independent indicator variable. Applying Chernoff bound, for any particular bin &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[Y_i&amp;gt;12\ln n] =\Pr[Y_i&amp;gt;(1+1)\mu]\le e^{-\frac{\mu}{3}} = e^{-2\ln n}= \frac{1}{n^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying the union bound, the probability that there exists a bin with load &amp;lt;math&amp;gt;&amp;gt;12\ln n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;n\cdot \Pr[Y_1&amp;gt;12\ln n]\le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, with probability at least &amp;lt;math&amp;gt;1-\frac{1}{n}&amp;lt;/math&amp;gt;, the maximum load is within &amp;lt;math&amp;gt;12\ln n=O(m/n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Set balancing ===&lt;br /&gt;
Supposed that we have an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; with 0-1 entries. We are looking for a &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt; that minimizes &amp;lt;math&amp;gt;\|Ab\|_\infty&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;\|\cdot\|_\infty&amp;lt;/math&amp;gt; is the infinity norm (also called &amp;lt;math&amp;gt;L_\infty&amp;lt;/math&amp;gt; norm) of a vector, and for the vector &amp;lt;math&amp;gt;c=Ab&amp;lt;/math&amp;gt;, &lt;br /&gt;
:&amp;lt;math&amp;gt;\|Ab\|_\infty=\max_{i=1,2,\ldots,n}|c_i|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We can also describe this problem as an optimization:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mbox{minimize }&lt;br /&gt;
&amp;amp;\quad&lt;br /&gt;
\|Ab\|_\infty\\&lt;br /&gt;
\mbox{subject to: }&lt;br /&gt;
&amp;amp;\quad&lt;br /&gt;
b\in\{-1,+1\}^m.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|The problem arises in designing statistical experiments. Suppose that we have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; &#039;&#039;&#039;subjects&#039;&#039;&#039;, each of which may have up to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;&#039;features&#039;&#039;&#039;. This gives us an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
\mbox{feature 1:}\\&lt;br /&gt;
\mbox{feature 2:}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
\mbox{feature n:}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{cccc}&lt;br /&gt;
a_{11} &amp;amp; a_{12} &amp;amp; \cdots &amp;amp; a_{1m}\\&lt;br /&gt;
a_{21} &amp;amp; a_{22} &amp;amp; \cdots &amp;amp; a_{2m}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; \vdots\\&lt;br /&gt;
a_{n1} &amp;amp; a_{n2} &amp;amp; \cdots &amp;amp; a_{nm}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right],&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where each column represents a subject and each row represent a feature. An entry &amp;lt;math&amp;gt;a_{ij}\in\{0,1\}&amp;lt;/math&amp;gt; indicates whether subject &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; has feature &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By multiplying a vector &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{cccc}&lt;br /&gt;
a_{11} &amp;amp; a_{12} &amp;amp; \cdots &amp;amp; a_{1m}\\&lt;br /&gt;
a_{21} &amp;amp; a_{22} &amp;amp; \cdots &amp;amp; a_{2m}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; \vdots\\&lt;br /&gt;
a_{n1} &amp;amp; a_{n2} &amp;amp; \cdots &amp;amp; a_{nm}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right]&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
b_{1}\\&lt;br /&gt;
b_{2}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
b_{m}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right]&lt;br /&gt;
=&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
c_{1}\\&lt;br /&gt;
c_{2}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
c_{n}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right],&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
the subjects are partitioned into two disjoint groups: one for -1 and other other for +1. Each &amp;lt;math&amp;gt;c_i&amp;lt;/math&amp;gt; gives the difference between the numbers of subjects with feature &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; in the two groups. By minimizing &amp;lt;math&amp;gt;\|Ab\|_\infty=\|c\|_\infty&amp;lt;/math&amp;gt;, we ask for an optimal partition so that each feature is roughly as balanced as possible between the two groups.&lt;br /&gt;
&lt;br /&gt;
In a scientific experiment, one of the group serves as a [http://en.wikipedia.org/wiki/Scientific_control control group] (对照组). Ideally, we want the two groups are statistically identical, which is usually impossible to achieve in practice. The requirement of minimizing &amp;lt;math&amp;gt;\|Ab\|_\infty&amp;lt;/math&amp;gt; actually means the statistical difference between the two groups are minimized.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
We propose an extremely simple &amp;quot;randomized algorithm&amp;quot; for computing a &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt;: for each &amp;lt;math&amp;gt;i=1,2,\ldots, m&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt; be independently chosen from &amp;lt;math&amp;gt;\{-1,+1\}&amp;lt;/math&amp;gt;, such that &lt;br /&gt;
:&amp;lt;math&amp;gt;b_i=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
-1 &amp;amp; \mbox{with probability }\frac{1}{2}\\&lt;br /&gt;
+1 &amp;amp;\mbox{with probability }\frac{1}{2}&lt;br /&gt;
\end{cases}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This procedure can hardly be called as an &amp;quot;algorithm&amp;quot;, because its decision is made disregard of the input &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. We then show that despite of this obliviousness, the algorithm chooses a good enough &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, such that for any &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\|Ab\|_\infty=O(\sqrt{m\ln n})&amp;lt;/math&amp;gt; with high probability.&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Theorem&#039;&#039;&#039;&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix with 0-1 entries. For a random vector &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; entries chosen independently and with equal probability from &amp;lt;math&amp;gt;\{-1,+1\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[\|Ab\|_\infty&amp;gt;2\sqrt{2m\ln n}]\le\frac{2}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039;&lt;br /&gt;
Consider particularly the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th row of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. The entry of &amp;lt;math&amp;gt;Ab&amp;lt;/math&amp;gt; contributed by row &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;c_i=\sum_{j=1}^m a_{ij}b_j&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; be the non-zero entries in the row. If &amp;lt;math&amp;gt;k\le2\sqrt{2m\ln n}&amp;lt;/math&amp;gt;, then clearly &amp;lt;math&amp;gt;|c_i|&amp;lt;/math&amp;gt; is no greater than &amp;lt;math&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt;. On the other hand if &amp;lt;math&amp;gt;k&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; then the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; nonzero terms in the sum&lt;br /&gt;
:&amp;lt;math&amp;gt;c_i=\sum_{j=1}^m a_{ij}b_j&amp;lt;/math&amp;gt;&lt;br /&gt;
are independent, each with probability 1/2 of being either +1 or -1. &lt;br /&gt;
&lt;br /&gt;
Thus, for these &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; nonzero terms, each &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt; is either positive or negative independently with equal probability. There are expectedly &amp;lt;math&amp;gt;\mu=\frac{k}{2}&amp;lt;/math&amp;gt; positive &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s among these &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; terms, and &amp;lt;math&amp;gt;c_i&amp;lt;-2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; only occurs when there are less than &amp;lt;math&amp;gt;\frac{k}{2}-\sqrt{2m\ln n}=\left(1-\delta\right)\mu&amp;lt;/math&amp;gt; positive &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s, where &amp;lt;math&amp;gt;\delta=\frac{2\sqrt{2m\ln n}}{k}&amp;lt;/math&amp;gt;. Applying Chernoff bound, this event occurs with probability at most&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\exp\left(-\frac{\mu\delta^2}{2}\right)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\frac{k}{2}\cdot\frac{8m\ln n}{2k^2}\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\frac{2m\ln n}{k}\right)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\exp\left(-\frac{2m\ln n}{m}\right)\\&lt;br /&gt;
&amp;amp;\le n^{-2}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The same argument can be applied to negative &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s, so that the probability that &amp;lt;math&amp;gt;c_i&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;n^{-2}&amp;lt;/math&amp;gt;. Therefore, by the union bound, &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[|c_i|&amp;gt; 2\sqrt{2m\ln n}]\le\frac{2}{n^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Apply the union bound to all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; rows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\|Ab\|_\infty&amp;gt;2\sqrt{2m\ln n}]\le n\cdot\Pr[|c_i|&amp;gt; 2\sqrt{2m\ln n}]\le\frac{2}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
So how good is this randomized algorithm? In fact when &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt; there exists a matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\|Ab\|_\infty=\Omega(\sqrt{n})&amp;lt;/math&amp;gt; for any choice of &amp;lt;math&amp;gt;b\in\{-1,+1\}^n&amp;lt;/math&amp;gt;.&lt;/div&gt;</summary>
		<author><name>172.21.1.247</name></author>
	</entry>
	<entry>
		<id>https://tcs.nju.edu.cn/wiki/index.php?title=Randomized_Algorithms_(Spring_2010)/Tail_inequalities&amp;diff=1326</id>
		<title>Randomized Algorithms (Spring 2010)/Tail inequalities</title>
		<link rel="alternate" type="text/html" href="https://tcs.nju.edu.cn/wiki/index.php?title=Randomized_Algorithms_(Spring_2010)/Tail_inequalities&amp;diff=1326"/>
		<updated>2010-02-23T06:37:53Z</updated>

		<summary type="html">&lt;p&gt;172.21.1.247: /* The Chernoff bound */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Select the Median ==&lt;br /&gt;
&lt;br /&gt;
The [http://en.wikipedia.org/wiki/Selection_algorithm selection problem] is the problem of finding the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt;th smallest element in a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. A typical case of selection problem is finding the &#039;&#039;&#039;median&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Definition&#039;&#039;&#039;&lt;br /&gt;
:The median of a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; is the &amp;lt;math&amp;gt;(\lceil n/2\rceil)&amp;lt;/math&amp;gt;th element in the sorted order of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
The median can be found in &amp;lt;math&amp;gt;O(n\log n)&amp;lt;/math&amp;gt; time by sorting. There is a linear-time deterministic algorithm, [http://en.wikipedia.org/wiki/Selection_algorithm#Linear_general_selection_algorithm_-_.22Median_of_Medians_algorithm.22 &amp;quot;median of medians&amp;quot; algorithm], which is quite sophisticated. Here we introduce a much simpler randomized algorithm which also runs in linear time.&lt;br /&gt;
&lt;br /&gt;
=== Randomized median algorithm ===&lt;br /&gt;
The idea of this algorithm is random sampling. For a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;m\in S&amp;lt;/math&amp;gt; denote the median. We observe that if we can find two elements &amp;lt;math&amp;gt;d,u\in S&amp;lt;/math&amp;gt; satisfying the following properties:&lt;br /&gt;
# The median is between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in the sorted order, i.e. &amp;lt;math&amp;gt;d\le m\le u&amp;lt;/math&amp;gt;;&lt;br /&gt;
# The total number of elements between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; is small, specially for &amp;lt;math&amp;gt;C=\{x\in S\mid d\le x\le u\}&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;|C|=o(n/\log n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Provided &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; with these two properties, within linear time, we can compute the ranks of &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, construct &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;, and sort &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;. Therefore, the median &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; can be picked from &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; in linear time.&lt;br /&gt;
&lt;br /&gt;
So how can we select such elements &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;? Certainly sorting &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; would give us the elements, but isn&#039;t that exactly what we want to avoid in the first place?&lt;br /&gt;
&lt;br /&gt;
Observe that &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; are only asked to roughly satisfy some constraints. This hints us maybe we can construct a &#039;&#039;sketch&#039;&#039; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; which is small enough to sort cheaply and roughly represents &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, and then pick &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; from this sketch. We construct the sketch by randomly sampling a relatively small number of elements from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Then the strategy of algorithm is outlined by:&lt;br /&gt;
* Sample a set &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of elements from &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. &lt;br /&gt;
* Sort &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; and choose &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; somewhere around the median of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
* If &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; have the desirable properties, we can compute the median in linear time, or otherwise the algorithm fails.&lt;br /&gt;
&lt;br /&gt;
The parameters to be fixed are: the size of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; (small enough to sort in linear time and large enough to contain sufficient information of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;); and the order of &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; (not too close to have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; between them, and not too far away to have &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; sortable in linear time).&lt;br /&gt;
&lt;br /&gt;
We choose the size of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; as &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt;, and &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; are within &amp;lt;math&amp;gt;\sqrt{n}&amp;lt;/math&amp;gt; range around the median of &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Randomized Median Algorithm:&#039;&#039;&#039;&lt;br /&gt;
|-&lt;br /&gt;
|&#039;&#039;&#039;Input:&#039;&#039;&#039; a set &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; elements over totally ordered domain.&lt;br /&gt;
# Pick a multi-set &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;\left\lceil n^{3/4}\right\rceil&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, chosen independently and uniformly at random with replacement, and sort &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
# Let &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;\left\lfloor\frac{1}{2}n^{3/4}-\sqrt{n}\right\rfloor&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;, and let &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;\left\lceil\frac{1}{2}n^{3/4}+\sqrt{n}\right\rceil&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt;.&lt;br /&gt;
# Construct &amp;lt;math&amp;gt;C=\{x\in S\mid d\le x\le u\}&amp;lt;/math&amp;gt; and compute the ranks &amp;lt;math&amp;gt;r_d=|\{x\in S\mid x&amp;lt;d\}|&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;r_u=|\{x\in S\mid x&amp;lt;u\}|&amp;lt;/math&amp;gt;.&lt;br /&gt;
# If &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;r_u&amp;lt;\frac{n}{2}&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;|C|&amp;gt;4n^{3/4}&amp;lt;/math&amp;gt; then return FAIL.&lt;br /&gt;
# Sort &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; and return the &amp;lt;math&amp;gt;\left(\left\lfloor\frac{n}{2}\right\rfloor-r_d+1\right)&amp;lt;/math&amp;gt;th element in the sorted order of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&amp;quot;Sample with replacement&amp;quot; (有放回采样) means that after sampling an element, we put the element back to the set. In this way, each sampled element is independently and identically distributed (&#039;&#039;i.i.d&#039;&#039;) (独立同分布). In the above algorithm, this is for our convenience of analysis.&lt;br /&gt;
&lt;br /&gt;
=== Analysis ===&lt;br /&gt;
The algorithm always terminates in linear time because each line of the algorithm costs at most linear time. The last three line guarantees that the algorithm returns the correct median if it does not fail.&lt;br /&gt;
&lt;br /&gt;
We then only need to bound the probability that the algorithm returns a FAIL. Let &amp;lt;math&amp;gt;m\in S&amp;lt;/math&amp;gt; be the median of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. By Line 4, we know that the algorithm returns a FAIL if and only if at least one of the following events occurs:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_1: Y=|\{x\in R\mid x\le m\}|&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_2: Z=|\{x\in R\mid x\ge m\}|&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3: |C|&amp;gt;4n^{3/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt; directly follows the third condition in Line 4. &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt; are a bit tricky. The first condition in Line 4 is that &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt;, which looks not exactly the same as &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt;, but both &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; and that &amp;lt;math&amp;gt;r_d&amp;gt;\frac{n}{2}&amp;lt;/math&amp;gt; are equivalent to the same event: the &amp;lt;math&amp;gt;\left\lfloor\frac{1}{2}n^{3/4}-\sqrt{n}\right\rfloor&amp;lt;/math&amp;gt;-th smallest element in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; is greater than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;, thus they are actually equivalent. Similarly, &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt; is equivalent to the second condition of Line 4.&lt;br /&gt;
&lt;br /&gt;
We now bound the probabilities of these events one by one.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 1&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_1]\le \frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; Let &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; be the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th sampled element in Line 1 of the algorithm. Let &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; be a indicator random variable such that &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
Y_i=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
1 &amp;amp; \mbox{if }X_i\le m,\\&lt;br /&gt;
0 &amp;amp; \mbox{otherwise.}&lt;br /&gt;
\end{cases}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
It is obvious that &amp;lt;math&amp;gt;Y=\sum_{i=1}^{n^{3/4}}Y_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is as defined in &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt;. For every &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt;, there are &amp;lt;math&amp;gt;\left\lceil\frac{n}{2}\right\rceil&amp;lt;/math&amp;gt; elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; that are less than or equal to the median. The probability that &amp;lt;math&amp;gt;Y_i=1&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
p=\Pr[Y_i=1]=\Pr[X_i\le m]=\frac{1}{n}\left\lceil\frac{n}{2}\right\rceil,&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
which is within the range of &amp;lt;math&amp;gt;\left[\frac{1}{2},\frac{1}{2}+\frac{1}{2n}\right]&amp;lt;/math&amp;gt;. Thus&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[Y]=n^{3/4}p\ge \frac{1}{2}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The event &amp;lt;math&amp;gt;\mathcal{E}_1&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;Y&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt;&#039;s are Bernoulli trials, and &amp;lt;math&amp;gt;Y&amp;lt;/math&amp;gt; is the sum of &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt; Bernoulli trials, which follows binomial distribution with parameters &amp;lt;math&amp;gt;n^{3/4}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;p&amp;lt;/math&amp;gt;. Thus, the variance is &lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{Var}[Y]=n^{3/4}p(1-p)\le \frac{1}{4}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Applying Chebyshev&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}_1]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[Y&amp;lt;\frac{1}{2}n^{3/4}-\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\Pr\left[|Y-\mathbf{E}[Y]|&amp;gt;\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[Y]}{n}\\&lt;br /&gt;
&amp;amp;\le\frac{1}{4}n^{-1/4}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
By a similar analysis, we can obtain the following bound for the event &amp;lt;math&amp;gt;\mathcal{E}_2&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 2&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_2]\le \frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
We now bound the probability of the event &amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Lemma 3&#039;&#039;&#039;&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_3]\le \frac{1}{2}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; The event &amp;lt;math&amp;gt;\mathcal{E}_3&amp;lt;/math&amp;gt; is defined as that &amp;lt;math&amp;gt;|C|&amp;gt;4 n^{3/4}&amp;lt;/math&amp;gt;, which by the Pigeonhole Principle, implies that at leas one of the following must be true:&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&amp;lt;/math&amp;gt;: at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is greater than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;;&lt;br /&gt;
* &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&#039;&amp;lt;/math&amp;gt;: at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is smaller than &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the probability that &amp;lt;math&amp;gt;\mathcal{E}_3&#039;&amp;lt;/math&amp;gt; occurs; the second will have the same bound by symmetry.&lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; is the region in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; between &amp;lt;math&amp;gt;d&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt;. If there are at least &amp;lt;math&amp;gt;2n^{3/4}&amp;lt;/math&amp;gt; elements of &amp;lt;math&amp;gt;C&amp;lt;/math&amp;gt; greater than the median &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;, then the rank of &amp;lt;math&amp;gt;u&amp;lt;/math&amp;gt; in the sorted order of &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt; must be at least &amp;lt;math&amp;gt;\frac{1}{2}n+2n^{3/4}&amp;lt;/math&amp;gt; and thus &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; has at least &amp;lt;math&amp;gt;\frac{1}{2}n^{3/4}-\sqrt{n}&amp;lt;/math&amp;gt; samples among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;X_i\in\{0,1\}&amp;lt;/math&amp;gt; indicate whether the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;th sample is among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;. Let &amp;lt;math&amp;gt;X=\sum_{i=1}^{n^{3/4}}X_i&amp;lt;/math&amp;gt; be the number of samples in &amp;lt;math&amp;gt;R&amp;lt;/math&amp;gt; among the &amp;lt;math&amp;gt;\frac{1}{2}n-2n^{3/4}&amp;lt;/math&amp;gt; largest elements in &amp;lt;math&amp;gt;S&amp;lt;/math&amp;gt;.&lt;br /&gt;
It holds that&lt;br /&gt;
:&amp;lt;math&amp;gt;p=\Pr[X_i=1]=\frac{\frac{1}{2}n-2n^{3/4}}{n}=\frac{1}{2}-2n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is a binomial random variable with &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{E}[X]=n^{3/4}p=\frac{1}{2}n^{3/4}-2\sqrt{n},&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
and &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\mathbf{Var}[X]=n^{3/4}p(1-p)=\frac{1}{4}n^{3/4}-4n^{1/4}&amp;lt;\frac{1}{4}n^{3/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying Chebyshev&#039;s inequality,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
\Pr[\mathcal{E}_3&#039;]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[X\ge\frac{1}{2}n^{3/4}-\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\Pr\left[|X-\mathbf{E}[X]|\ge\sqrt{n}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{Var}[X]}{n}\\&lt;br /&gt;
&amp;amp;\le\frac{1}{4}n^{-1/4}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Symmetrically, we have that &amp;lt;math&amp;gt;\Pr[\mathcal{E}_3&#039;&#039;]\le\frac{1}{4}n^{-1/4}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Applying the union bound&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\mathcal{E}_3]\le \Pr[\mathcal{E}_3&#039;]+\Pr[\mathcal{E}_3&#039;&#039;]\le\frac{1}{2}n^{-1/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Combining the three bounds. Applying the union bound to them, the probability that the algorithm returns a FAIL is at most &lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[\mathcal{E}_1]+\Pr[\mathcal{E}_2]+\Pr[\mathcal{E}_3]\le n^{-1/4}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Therefore the algorithm always terminates in linear time and returns the correct median with high probability.&lt;br /&gt;
&lt;br /&gt;
== Chernoff Bound ==&lt;br /&gt;
Suppose that we have a fair coin. If we toss it once, then the outcome is completely unpredictable. But if we toss it, say for 1000 times, then the number of HEADs is very likely to be around 500. This striking phenomenon is called the &#039;&#039;&#039;concentration&#039;&#039;&#039;. The Chernoff bound captures the concentration of independent trials.&lt;br /&gt;
&lt;br /&gt;
The Chernoff bound is also a tail bound for the sum of independent random variables which may give us &#039;&#039;exponentially&#039;&#039; sharp bounds.&lt;br /&gt;
&lt;br /&gt;
Before proving the Chernoff bound, we should talk about the moment generating functions.&lt;br /&gt;
&lt;br /&gt;
=== Moment generating functions ===&lt;br /&gt;
The more we know about the moments of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;, the more information we would have about &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt;. There is a so-called &#039;&#039;&#039;moment generating function&#039;&#039;&#039;, which &amp;quot;packs&amp;quot; all the information about the moments of &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; into one function.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Definition:&#039;&#039;&#039;&lt;br /&gt;
:The moment generating function of a random variable &amp;lt;math&amp;gt;X&amp;lt;/math&amp;gt; is defined as &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt; is the parameter of the function.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
By Taylor&#039;s expansion and the linearity of expectations,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\sum_{k=0}^\infty\frac{\lambda^k}{k!}X^k\right]\\&lt;br /&gt;
&amp;amp;=\sum_{k=0}^\infty\frac{\lambda^k}{k!}\mathbf{E}\left[X^k\right]&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
&lt;br /&gt;
The moment generating function &amp;lt;math&amp;gt;\mathbf{E}\left[\mathrm{e}^{\lambda X}\right]&amp;lt;/math&amp;gt; is a function of &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== The Chernoff bound ===&lt;br /&gt;
The Chernoff bounds are tail inequalities with exponential decays for the sum of independent trials.&lt;br /&gt;
The bounds are obtained by applying Markov&#039;s inequality to the moment generating function of the sum of independent trials, with some  appropriate choice of the parameter &amp;lt;math&amp;gt;\lambda&amp;lt;/math&amp;gt;.&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Chernoff bound (the upper tail):&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; For any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;X\ge (1+\delta)\mu&amp;lt;/math&amp;gt; is equivalent to that &amp;lt;math&amp;gt;e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}&amp;lt;/math&amp;gt;, thus&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1+\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
where the last step follows by Markov&#039;s inequality.&lt;br /&gt;
&lt;br /&gt;
Computing the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[e^{\lambda \sum_{i=1}^n X_i}\right]\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\mathbf{E}\left[\prod_{i=1}^n e^{\lambda X_i}\right].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
For independent random variables, the expectation of the product equals the product of the expectations, therefore, for the last term above, we have&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[\prod_{i=1}^n e^{\lambda X_i}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right].&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;p_i=\Pr[X_i=1]&amp;lt;/math&amp;gt; for &amp;lt;math&amp;gt;i=1,2,\ldots,n&amp;lt;/math&amp;gt;. Then,&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[X]=\mathbf{E}\left[\sum_{i=1}^n X_i\right]=\sum_{i=1}^n\mathbf{E}[X_i]=\sum_{i=1}^n p_i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We bound the moment generating function for each individual &amp;lt;math&amp;gt;X_i&amp;lt;/math&amp;gt; as follows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X_i}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
p_i\cdot e^{\lambda\cdot 1}+(1-p_i)\cdot e^{\lambda\cdot 0}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
1+p_i(e^\lambda -1)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
e^{p_i(e^\lambda-1)},&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
where in the last step we apply the Taylor&#039;s expansion so that &amp;lt;math&amp;gt;e^y\ge 1+y&amp;lt;/math&amp;gt; where &amp;lt;math&amp;gt;y=p_i(e^\lambda-1)\ge 0&amp;lt;/math&amp;gt;. (By doing this, we can transform the product to the sum of &amp;lt;math&amp;gt;p_i&amp;lt;/math&amp;gt;, which is &amp;lt;math&amp;gt;\mu&amp;lt;/math&amp;gt;.) &lt;br /&gt;
&lt;br /&gt;
Therefore,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mathbf{E}\left[e^{\lambda X}\right]&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\prod_{i=1}^n \mathbf{E}\left[e^{\lambda X_i}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\prod_{i=1}^n e^{p_i(e^\lambda-1)}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(\sum_{i=1}^n p_i(e^{\lambda}-1)\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
e^{(e^\lambda-1)\mu}.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
Thus, we have shown that for any &amp;lt;math&amp;gt;\lambda&amp;gt;0&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge (1+\delta)\mu] &lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{e^{(e^\lambda-1)\mu}}{e^{\lambda (1+\delta)\mu}}\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}\right)^\mu&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;.&lt;br /&gt;
For any &amp;lt;math&amp;gt;\delta&amp;gt;0&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;gt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]\le\left(\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The idea of the proof is actually quite clear: we apply Markov&#039;s inequality to &amp;lt;math&amp;gt;e^{\lambda X}&amp;lt;/math&amp;gt; and for the rest, we just estimate the moment generating function &amp;lt;math&amp;gt;\mathbf{E}[e^{\lambda X}]&amp;lt;/math&amp;gt;. To make the bound as tight as possible, we minimized the &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt; by setting &amp;lt;math&amp;gt;\lambda=\ln(1+\delta)&amp;lt;/math&amp;gt;, which can be justified by taking derivatives of &amp;lt;math&amp;gt;\frac{e^{(e^\lambda-1)}}{e^{\lambda (1+\delta)}}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
We then proceed to the lower tail, the probability that the random variable deviates below the mean value:&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Chernoff bound (the lower tail):&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. &lt;br /&gt;
:Then for any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; For any &amp;lt;math&amp;gt;\lambda&amp;lt;0&amp;lt;/math&amp;gt;, by the same analysis as in the upper tail version,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\le (1-\delta)\mu] &lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\Pr\left[e^{\lambda X}\ge e^{\lambda (1-\delta)\mu}\right]\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\frac{\mathbf{E}\left[e^{\lambda X}\right]}{e^{\lambda (1-\delta)\mu}}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^{(e^\lambda-1)}}{e^{\lambda (1-\delta)}}\right)^\mu.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt; &lt;br /&gt;
For any &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt;1&amp;lt;/math&amp;gt;, we can let &amp;lt;math&amp;gt;\lambda=\ln(1-\delta)&amp;lt;0&amp;lt;/math&amp;gt; to get&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[X\ge (1-\delta)\mu]\le\left(\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\right)^{\mu}.&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
Some useful special forms of the bounds can be derived directly from the above general forms of the bounds. We now know better why we say that the bounds are exponentially sharp.&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Useful forms of the Chernoff bound&#039;&#039;&#039;&lt;br /&gt;
:Let  &amp;lt;math&amp;gt;X=\sum_{i=1}^n X_i&amp;lt;/math&amp;gt;, where &amp;lt;math&amp;gt;X_1, X_2, \ldots, X_n&amp;lt;/math&amp;gt; are independent Poisson trials. Let &amp;lt;math&amp;gt;\mu=\mathbf{E}[X]&amp;lt;/math&amp;gt;. Then&lt;br /&gt;
:1. for &amp;lt;math&amp;gt;0&amp;lt;\delta\le 1&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge (1+\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{3}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\le (1-\delta)\mu]&amp;lt;\exp\left(-\frac{\mu\delta^2}{2}\right);&amp;lt;/math&amp;gt;&lt;br /&gt;
:2. for &amp;lt;math&amp;gt;t\ge 2e\mu&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[X\ge t]\le 2^{-t}.&amp;lt;/math&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; To obtain the bounds in (1), we need to show that for &amp;lt;math&amp;gt;0&amp;lt;\delta&amp;lt; 1&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\frac{e^{\delta}}{(1+\delta)^{(1+\delta)}}\le e^{-\delta^2/3}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\frac{e^{-\delta}}{(1-\delta)^{(1-\delta)}}\le e^{-\delta^2/2}&amp;lt;/math&amp;gt;. We can verify both inequalities by standard analysis techniques.&lt;br /&gt;
&lt;br /&gt;
To obtain the bound in (2), let &amp;lt;math&amp;gt;t=(1+\delta)\mu&amp;lt;/math&amp;gt;. Then &amp;lt;math&amp;gt;\delta=t/\mu-1\ge 2e-1&amp;lt;/math&amp;gt;. Hence,&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\Pr[X\ge(1+\delta)\mu]&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e^\delta}{(1+\delta)^{(1+\delta)}}\right)^\mu\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{1+\delta}\right)^{(1+\delta)\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\left(\frac{e}{2e}\right)^{2e\mu}\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
2^{-t}&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Applications of Chernoff Bounds ==&lt;br /&gt;
We now introduce some applications of Chernoff bounds in randomized algorithms.&lt;br /&gt;
&lt;br /&gt;
=== Balls into bins ===&lt;br /&gt;
Throwing &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; balls uniformly and independently to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; bins, what is the maximum load of all bins? In the last class, by using a counting argument, we proved that for the case that &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt;, the maximum load is &amp;lt;math&amp;gt;O(\ln n\ln\ln n)&amp;lt;/math&amp;gt; with high probability. Now we show that when there are more balls, the loads are more balanced.&lt;br /&gt;
&lt;br /&gt;
For any &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;j\in[m]&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;X_{ij}&amp;lt;/math&amp;gt; be the indicator variable for the event that ball &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; is thrown to bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. Obviously&lt;br /&gt;
:&amp;lt;math&amp;gt;\mathbf{E}[X_{ij}]=\Pr[\mbox{ball }j\mbox{ is thrown to bin }i]=\frac{1}{n}&amp;lt;/math&amp;gt;&lt;br /&gt;
Let &amp;lt;math&amp;gt;Y_i=\sum_{j\in[m]}X_{ij}&amp;lt;/math&amp;gt; be the load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Let us consider the case when &amp;lt;math&amp;gt;m=6n\ln n&amp;lt;/math&amp;gt;. Then the expected load of bin &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;\mu=\mathbf{E}[Y_i]=\mathbf{E}\left[\sum_{j\in[m]}X_{ij}\right]=\sum_{j\in[m]}\mathbf{E}[X_{ij}]=m/n=6\ln n&amp;lt;/math&amp;gt;. &lt;br /&gt;
Note that &amp;lt;math&amp;gt;Y_i&amp;lt;/math&amp;gt; is a sum of &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; mutually independent indicator variable. Applying Chernoff bound, for any particular bin &amp;lt;math&amp;gt;i\in[n]&amp;lt;/math&amp;gt;,&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\Pr[Y_i&amp;gt;12\ln n] =\Pr[Y_i&amp;gt;(1+1)\mu]\le e^{-\frac{\mu}{3}} = e^{-2\ln n}= \frac{1}{n^2}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
Applying the union bound, the probability that there exists a bin with load &amp;lt;math&amp;gt;&amp;gt;12\ln n&amp;lt;/math&amp;gt; is&lt;br /&gt;
:&amp;lt;math&amp;gt;n\cdot \Pr[Y_1&amp;gt;12\ln n]\le \frac{1}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Therefore, with probability at least &amp;lt;math&amp;gt;1-\frac{1}{n}&amp;lt;/math&amp;gt;, the maximum load is within &amp;lt;math&amp;gt;12\ln n=O(m/n)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
=== Set balancing ===&lt;br /&gt;
Supposed that we have an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; with 0-1 entries. We are looking for a &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt; that minimizes &amp;lt;math&amp;gt;\|Ab\|_\infty&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Recall that &amp;lt;math&amp;gt;\|\cdot\|_\infty&amp;lt;/math&amp;gt; is the infinity norm (also called &amp;lt;math&amp;gt;L_\infty&amp;lt;/math&amp;gt; norm) of a vector, and for the vector &amp;lt;math&amp;gt;c=Ab&amp;lt;/math&amp;gt;, &lt;br /&gt;
:&amp;lt;math&amp;gt;\|Ab\|_\infty=\max_{i=1,2,\ldots,n}|c_i|&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
We can also describe this problem as an optimization:&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\mbox{minimize }&lt;br /&gt;
&amp;amp;\quad&lt;br /&gt;
\|Ab\|_\infty\\&lt;br /&gt;
\mbox{subject to: }&lt;br /&gt;
&amp;amp;\quad&lt;br /&gt;
b\in\{-1,+1\}^m.&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|The problem arises in designing statistical experiments. Suppose that we have &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; &#039;&#039;&#039;subjects&#039;&#039;&#039;, each of which may have up to &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; &#039;&#039;&#039;features&#039;&#039;&#039;. This gives us an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
\mbox{feature 1:}\\&lt;br /&gt;
\mbox{feature 2:}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
\mbox{feature n:}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{cccc}&lt;br /&gt;
a_{11} &amp;amp; a_{12} &amp;amp; \cdots &amp;amp; a_{1m}\\&lt;br /&gt;
a_{21} &amp;amp; a_{22} &amp;amp; \cdots &amp;amp; a_{2m}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; \vdots\\&lt;br /&gt;
a_{n1} &amp;amp; a_{n2} &amp;amp; \cdots &amp;amp; a_{nm}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right],&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
where each column represents a subject and each row represent a feature. An entry &amp;lt;math&amp;gt;a_{ij}\in\{0,1\}&amp;lt;/math&amp;gt; indicates whether subject &amp;lt;math&amp;gt;j&amp;lt;/math&amp;gt; has feature &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
By multiplying a vector &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt;&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{cccc}&lt;br /&gt;
a_{11} &amp;amp; a_{12} &amp;amp; \cdots &amp;amp; a_{1m}\\&lt;br /&gt;
a_{21} &amp;amp; a_{22} &amp;amp; \cdots &amp;amp; a_{2m}\\&lt;br /&gt;
\vdots &amp;amp; \vdots &amp;amp; \ddots &amp;amp; \vdots\\&lt;br /&gt;
a_{n1} &amp;amp; a_{n2} &amp;amp; \cdots &amp;amp; a_{nm}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right]&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
b_{1}\\&lt;br /&gt;
b_{2}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
b_{m}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right]&lt;br /&gt;
=&lt;br /&gt;
\left[&lt;br /&gt;
\begin{array}{c}&lt;br /&gt;
c_{1}\\&lt;br /&gt;
c_{2}\\&lt;br /&gt;
\vdots\\&lt;br /&gt;
c_{n}\\&lt;br /&gt;
\end{array}&lt;br /&gt;
\right],&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
the subjects are partitioned into two disjoint groups: one for -1 and other other for +1. Each &amp;lt;math&amp;gt;c_i&amp;lt;/math&amp;gt; gives the difference between the numbers of subjects with feature &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; in the two groups. By minimizing &amp;lt;math&amp;gt;\|Ab\|_\infty=\|c\|_\infty&amp;lt;/math&amp;gt;, we ask for an optimal partition so that each feature is roughly as balanced as possible between the two groups.&lt;br /&gt;
&lt;br /&gt;
In a scientific experiment, one of the group serves as a [http://en.wikipedia.org/wiki/Scientific_control control group] (对照组). Ideally, we want the two groups are statistically identical, which is usually impossible to achieve in practice. The requirement of minimizing &amp;lt;math&amp;gt;\|Ab\|_\infty&amp;lt;/math&amp;gt; actually means the statistical difference between the two groups are minimized.&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
We propose an extremely simple &amp;quot;randomized algorithm&amp;quot; for computing a &amp;lt;math&amp;gt;b\in\{-1,+1\}^m&amp;lt;/math&amp;gt;: for each &amp;lt;math&amp;gt;i=1,2,\ldots, m&amp;lt;/math&amp;gt;, let &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt; be independently chosen from &amp;lt;math&amp;gt;\{-1,+1\}&amp;lt;/math&amp;gt;, such that &lt;br /&gt;
:&amp;lt;math&amp;gt;b_i=&lt;br /&gt;
\begin{cases}&lt;br /&gt;
-1 &amp;amp; \mbox{with probability }\frac{1}{2}\\&lt;br /&gt;
+1 &amp;amp;\mbox{with probability }\frac{1}{2}&lt;br /&gt;
\end{cases}.&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This procedure can hardly be called as an &amp;quot;algorithm&amp;quot;, because its decision is made disregard of the input &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. We then show that despite of this obliviousness, the algorithm chooses a good enough &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt;, such that for any &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;, &amp;lt;math&amp;gt;\|Ab\|_\infty=O(\sqrt{m\ln n})&amp;lt;/math&amp;gt; with high probability.&lt;br /&gt;
{|border=&amp;quot;1&amp;quot;&lt;br /&gt;
|&#039;&#039;&#039;Theorem&#039;&#039;&#039;&lt;br /&gt;
:Let &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; be an &amp;lt;math&amp;gt;n\times m&amp;lt;/math&amp;gt; matrix with 0-1 entries. For a random vector &amp;lt;math&amp;gt;b&amp;lt;/math&amp;gt; with &amp;lt;math&amp;gt;m&amp;lt;/math&amp;gt; entries chosen independently and with equal probability from &amp;lt;math&amp;gt;\{-1,+1\}&amp;lt;/math&amp;gt;,&lt;br /&gt;
::&amp;lt;math&amp;gt;\Pr[\|Ab\|_\infty&amp;gt;2\sqrt{2m\ln n}]\le\frac{2}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
|}&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039;&lt;br /&gt;
Consider particularly the &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt;-th row of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;. The entry of &amp;lt;math&amp;gt;Ab&amp;lt;/math&amp;gt; contributed by row &amp;lt;math&amp;gt;i&amp;lt;/math&amp;gt; is &amp;lt;math&amp;gt;c_i=\sum_{j=1}^m a_{ij}b_j&amp;lt;/math&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
Let &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; be the non-zero entries in the row. If &amp;lt;math&amp;gt;k\le2\sqrt{2m\ln n}&amp;lt;/math&amp;gt;, then clearly &amp;lt;math&amp;gt;|c_i|&amp;lt;/math&amp;gt; is no greater than &amp;lt;math&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt;. On the other hand if &amp;lt;math&amp;gt;k&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; then the &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; nonzero terms in the sum&lt;br /&gt;
:&amp;lt;math&amp;gt;c_i=\sum_{j=1}^m a_{ij}b_j&amp;lt;/math&amp;gt;&lt;br /&gt;
are independent, each with probability 1/2 of being either +1 or -1. &lt;br /&gt;
&lt;br /&gt;
Thus, for these &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; nonzero terms, each &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt; is either positive or negative independently with equal probability. There are expectedly &amp;lt;math&amp;gt;\mu=\frac{k}{2}&amp;lt;/math&amp;gt; positive &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s among these &amp;lt;math&amp;gt;k&amp;lt;/math&amp;gt; terms, and &amp;lt;math&amp;gt;c_i&amp;lt;-2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; only occurs when there are less than &amp;lt;math&amp;gt;\frac{k}{2}-\sqrt{2m\ln n}=\left(1-\delta\right)\mu&amp;lt;/math&amp;gt; positive &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s, where &amp;lt;math&amp;gt;\delta=\frac{2\sqrt{2m\ln n}}{k}&amp;lt;/math&amp;gt;. Applying Chernoff bound, this event occurs with probability at most&lt;br /&gt;
:&amp;lt;math&amp;gt;\begin{align}&lt;br /&gt;
\exp\left(-\frac{\mu\delta^2}{2}\right)&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\frac{k}{2}\cdot\frac{8m\ln n}{2k^2}\right)\\&lt;br /&gt;
&amp;amp;=&lt;br /&gt;
\exp\left(-\frac{2m\ln n}{k}\right)\\&lt;br /&gt;
&amp;amp;\le&lt;br /&gt;
\exp\left(-\frac{2m\ln n}{m}\right)\\&lt;br /&gt;
&amp;amp;\le n^{-2}.&lt;br /&gt;
\end{align}&lt;br /&gt;
&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The same argument can be applied to negative &amp;lt;math&amp;gt;b_i&amp;lt;/math&amp;gt;&#039;s, so that the probability that &amp;lt;math&amp;gt;c_i&amp;gt;2\sqrt{2m\ln n}&amp;lt;/math&amp;gt; is at most &amp;lt;math&amp;gt;n^{-2}&amp;lt;/math&amp;gt;. Therefore, by the union bound, &lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[|c_i|&amp;gt; 2\sqrt{2m\ln n}]\le\frac{2}{n^2}&amp;lt;/math&amp;gt;.&lt;br /&gt;
Apply the union bound to all &amp;lt;math&amp;gt;n&amp;lt;/math&amp;gt; rows.&lt;br /&gt;
:&amp;lt;math&amp;gt;\Pr[\|Ab\|_\infty&amp;gt;2\sqrt{2m\ln n}]\le n\cdot\Pr[|c_i|&amp;gt; 2\sqrt{2m\ln n}]\le\frac{2}{n}&amp;lt;/math&amp;gt;.&lt;br /&gt;
&amp;lt;math&amp;gt;\square&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
So how good is this randomized algorithm? In fact when &amp;lt;math&amp;gt;m=n&amp;lt;/math&amp;gt; there exists a matrix &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; such that &amp;lt;math&amp;gt;\|Ab\|_\infty=\Omega(\sqrt{n})&amp;lt;/math&amp;gt; for any choice of &amp;lt;math&amp;gt;b\in\{-1,+1\}^n&amp;lt;/math&amp;gt;.&lt;/div&gt;</summary>
		<author><name>172.21.1.247</name></author>
	</entry>
</feed>