Redundancy Elimination

\( \newcommand{\IN}[1]{\mathit{in}[#1]} \) \( \newcommand{\VAL}[1]{\mathit{value}[#1]} \) \( \newcommand{\OUT}[1]{\mathit{out}[#1]} \) \( \newcommand{\USE}[1]{\mathit{use}[#1]} \) \( \newcommand{\DEF}[1]{\mathit{def}[#1]} \) \( \newcommand{\GEN}[1]{\mathit{gen}[#1]} \) \( \newcommand{\KILL}[1]{\mathit{kill}[#1]} \) \( \newcommand{\EXPRS}[1]{\mathit{exprs}[#1]} \) \( \newcommand\INV{\mathit{INV}} \) \( \newcommand\OP{\mathop{\mathit{OP}}} \)

Redundancy elimination

An optimization aims at redundancy elimination if it prevents the same computation from being done twice. This is an important class of optimizations, although it has become less important as compute has sped up relative to memory. Avoiding recomputation by saving a computed result takes space; in some cases it may be better to recompute from scratch!

Local value numbering

Extended basic block

Local value numbering is a classic redundancy elimination optimization. It is called a “local” optimization because it is performed on a single basic block. It can as easily be applied to an extended basic block (EBB), which is a sequence of nodes in which some nodes are allowed to have exit edges, as shown on the right. More generally, it can be applied to any tree-like subgraph in which all nodes have only one predecessor and there is a single node that dominates all other nodes. Each path through such a subgraph forms an EBB.

The idea of value numbering is to label each computed expression with a distinct identifier (a “number”), such that recomputations of the same expression are assigned the same number. If an expression is labeled with the same number as a variable, the variable can be used in place of the expression. Value numbering overlaps but does not completely subsume constant propagation and common subexpression elimination.

For example, every expression and variable in the following code has been given a number, written as a subscript. Expressions with the same number always compute the same value. (For simplicity, constants have not been given a number, but it would actually simplify the implementation.)

a₂ = (i₁ + 5)₂ j₁ = i₁ b₂ = 5 + j₁ i₂ = i₁ + 5 if c₄ = (i₂ + 1)₃ goto L₁ d₃ = (i₂ + 1)₃

To do this numbering, we start at the beginning of the EBB and assign numbers to expressions as they are encountered. As expressions are assigned numbers, the analysis remembers the mapping between that expression, expressed in terms of the value numbers it uses, and the new value number. If the same expression is seen again, the same number is assigned to it.

In the example, the variable i is given number 1. Then the sum i+1 is given number 2, and the analysis records that the expression \((\VAL{1} + 1)\) is the same as \(\VAL 2\). The assignment to a means that variable has value 2 at that program point. Assignment j=i puts value 1 into j, so j is also given number 1 at that program point. The expression 5 + j computes \((\VAL{1} + 1)\), so it is given number 2. This means we can replace 5+j with a, eliminating a redundant computation.

After doing the analysis, we look for value numbers that appear multiple times. An expression that uses a value computed previously is replaced with a variable with the same value number. If no such variable exists, a new variable is introduced at the first time the value was computed, to be used to replace later occurrences of the value.

For example, the above code can be optimized as follows:

a₂ = (i₁ + 5)₂ j₁ = i₁ b₂ = a₂ i₂ = a₂ t₃ = (a₂ + 1)₃ if c₄ = t₃ goto L₁ d₃ = t₃

The mapping from previously computed expressions to value numbers can be maintained implicitly by generating value numbers with a strong hash function. For example, we could generate the actual representation of value 2 by hashing the string “plus(v1, const(1))”, where \(v_1\) represents the hash value assigned to value 1. Note that to make sure that we get the same value number for the computation 5+j, it will be necessary to order the arguments to commutative operators (e.g., +) in some canonical ordering (e.g., sorting in dictionary order).

There are global versions of value numbering that operate on a general CFG. However, these are awkward unless the CFG is converted to single static assignment (SSA) form.

Common subexpression elimination

Common subexpression elimination

Common subexpression elimination (CSE) is a classic optimization that replaces redundantly computed expressions with a variable containing the value of the expression. It works on a general CFG. An expression computed at a node is a common subexpression at that if it is computed on all paths to the the node, and its operands have the same value at the points where it is computed as if it were computed at the current node. In this case, we can reuse the previously computed value at this node.

For example, in the figure above, the expression a+1 is a common subexpression at the bottom node. Therefore, it can be saved into a new temporary t in the top node, and this temporary can be used in the bottom one.

It is worth noting that CSE can make code slower, because it may increase the number of live variables, causing spilling. If there is a lot of register pressure, the reverse transformation, forward substitution, may improve performance. Forward substitution copies expressions forward when it is cheaper to recompute them than to save them in a variable.

Available expressions analysis

An expression is available if it has been computed in a dominating node and its operands have not been redefined. The available expressions analysis finds the set of such expressions. Implicitly, each such expression is tagged with the location in the CFG that it comes from, to allow the CSE transformation to be done.

Available expressions is a forward analysis. We define \(\OUT n\) to be the set of available expressions on edges leaving node \(n\). An expression is available if it was evaluated at \(n\), or was available on all edges entering \(n\), and it was not killed by \(n\):

\begin{align*} \IN{n} &= \bigcap_{n'≺n} \OUT{n} \\ \OUT{n} &= \IN{n}∪ \EXPRS{n} - \KILL{n} \end{align*}

Therefore, dataflow values are sets of expressions ordered by \(⊆\); the meet operator is set intersection (∩), and the top value is the set of all expressions, usually implemented as a special value that acts as the identity for ∩.

The expressions evaluated and killed by a node, \(\EXPRS n\), are summarized in the following table. Note that we try to include memory operands as expressions subject to CSE, because replacing memory accesses with register accesses is a useful optimization.

\( n \) \(\EXPRS{n}\) \( \KILL{n} \)
\(\texttt{start}\) \(∅\) all expressions
\(x ← e\) \(e\) and all subexpressions of \(e\) all expressions containing \(x\)
\([e_1] ← e_2\) \(e_2\), \([e_1]\), and subexpressions thereof all expressions \([e']\) that might alias \([e_1]\) or that might be affected by computing \(e_1\) or \(e_2\).
\( x ← f(\vec e)\) \(\vec{e}\) and subexpressions expressions containing \(x\) and expressions \([e']\) that could be changed by function call to \(f\)
\(\texttt{if}~e\) \(e\) and subexpressions expressions \([e'] \) that could be changed by evaluation of \(e\).

If a node \(n\) computes an expression \(e\) that is used and available in other nodes, the optimization proceeds as follows:

  1. In the basic block that the available expression came from, add a computation \(t ← e\) before the node that computed \(e\), and replace the use of \(e\) with \(t\).
  2. Replace expression \(e\) in other nodes where it is used and available with \(t\).

CSE works well with copy propagation. For example, the variable b in the above code may become dead after copy propagation. However, CSE plus copy propagation can enable more CSE, because CSE only recognizes syntactically identical expressions as the same. Copy propagation can make semantically identical expressions look the same through its renaming of variables. It is possible to generalize Available Expressions to keep track of equalities more semantically, though this makes the analysis much more complex.

Partial redundancy elimination

CSE eliminates computation of fully redundant expressions: those computed on all paths leading to a node. Partially redundant expressions are those computed at least twice along some path, but not necessarily all paths. Partial redundancy elimination (PRE) eliminates these partially redundant expressions. PRE subsumes CSE and loop-invariant code motion.

Partial redundancy elimination

The figure above shows an example of PRE. The computation b+c is redundant along some paths but not others. To make it fully redundant, we place computation of b+c onto earlier edges so that it has always been computed at each point where it is needed. Then the computation can be replaced with the value computed earlier.

Lazy code motion

The idea of lazy code motion is to eliminate all redundant computations while avoid creating any unnecessary computation: computations are moved earlier in the CFG. Further, we want to make sure that although the computations are moved earlier in the CFG, they are postponed as long as possible, to avoid creating register pressure.

The approach is to first identify candidate locations where the partially redundant expression could have been moved in order to make it fully redundant, without creating extra computations. Then among these candidates, we choose the one that comes latest along each path that needs it.

Anticipated expressions

The anticipated expressions analysis (also known as very busy expressions) finds expressions that are needed along every path leaving a given node. If an expression is needed along every path leaving the node, then there can be no wasted computation if the expression is moved to that node.

This is a backward analysis, in which the dataflow values are sets of expressions and the meet operator is ∩.

Once we know the anticipated expressions at each node, we tentatively place computations of these expressions and use an available expressions analysis to find expressions that are fully redundant under the assumption that the anticipated expressions are computed everywhere anticipated. These fully redundant expressions are the expressions to which we can apply the PRE optimization.

Postponable expressions

At this point we know some set of nodes where the expression can be moved, and we know where it is used. But we have not chosen the best place to move the expression. We need to pick a set of edges that separate these two parts of the CFG, and put the computation of the expression on those edges. We want to postpone the computation as long as possible. The postponable expressions analysis finds expressions \(e\) that are anticipated at program point \(p\) but not yet used: every path from the start to \(p\) contains an anticipation of \(e\) and no use before \(p\). This is a forward analysis with meet operator ∩.

Once postponable expressions have been computed, certain edges form a frontier where the expression transitions from postponable to not postponable. It is on these edges that the new node computing the expressions is placed.

Equality saturation

A common element in redundancy elimination optimizations is determining whether two expressions are equal. Traditionally these optimizations just use syntactic equality, often with some support for recognizing equality when the arguments to commutative operations are swapped.

A more general technique for recognizing equal expressions is equality saturation, an analysis technique that finds all equal expressions relative to a particular set of equations, including expressions that do not even appear in the program. Generating entirely new expressions can be used for algebraic simplification, though sometimes at a significant cost when the number of generated expressions is large.

The key to representing a large number of expressions efficiently is the e-graph data structure. The e-graph is a graph in which the nodes represent computed expressions. Nodes corresponding to operations (e.g., +) have successor nodes representing operands. Nodes representing variables or constants would have no successors. In addition, the e-graph maintains a set of e-classes, which are equivalence classes of nodes representing computations known to compute the same value.

Equality saturation proceeds by repeatedly searching the e-graph for opportunities to apply equations to derive new e-graph nodes and to derive new equalities that expand e-classes. This process is followed until saturation, when no further applications are possible.

For example, suppose that we want to optimize the expression (x + 1) + (-1) using equality saturation, using the following equations:

x + y = y + x
(x + y) + z = x + (y + z)
x + (-x) = 0
x + 0 = x

The following animated diagram shows how these equations can be applied to derive various useful equalities related to (x + 1) + (-1), with the dashed blue lines connecting expressions known to be equivalent. Thus, any set of nodes connected transitively by dashed blue lines belong to the same e-class. The final step shows that the expression x is equivalent to the original expression, because there is a dashed blue path between those two nodes.

E-graph construction for (x + 1) + (-1)

To efficiently represent e-classes, each node can contain a pointer to another node that is known to be equivalent. The union-find algorithm with path compression is used to test in near-linear time whether any two nodes are known to compute the same value.

E-graphs are a remarkably compact way to represent a large set of potential equalities, including infinite sets. For example, the e-graph for the equation \(x = x + 0\) automatically represents an infinite set of related equalities: \(x = x + 0 = (x + 0) + 0 = ((x + 0) + 0) + 0 + \dots \). However, the e-graph also in general produces many nodes that turn out not to be useful, such as, in the above example, the representation of the equation \(1 + x = x + 1\). In general, the e-graph is large and expensive to compute, and it may be prudent for the compiler to use it only on performance-critical code segments.

E-classes can also be used as the basis for program analysis, with dataflow analysis facts attached to e-classes.

The egg library is currently a popular and well-engineered implementation of equality saturation, and a number of compilers use egg internally.

References

Max Willsey, Chandrakana Nandi, Yisu Remy Wang, Oliver Flatt, Zachary Tatlock, Pavel Panchekha. egg: Fast and Extensible Equality Saturation. POPL 2021.