Calling Conventions

\(\newcommand{\MOVE}{\textit{MOVE}} \newcommand{\TEMP}{\textit{TEMP}} \newcommand{\CALL}{\textit{CALL}} \newcommand{\MEM}{\textit{MEM}} \newcommand{\ADD}{\textit{ADD}} \newcommand{\CONST}{\textit{CONST}} \newcommand{\SEQ}{\textit{SEQ}} \newcommand{\SEQ}{\textit{SEQ}} \newcommand{\ESEQ}{\textit{ESEQ}} \newcommand{\JUMP}{\textit{JUMP}} \newcommand{\CJUMP}{\textit{CJUMP}} \newcommand{\LABEL}{\textit{LABEL}} \newcommand{\RETURN}{\textit{RETURN}} \newcommand{\dest}{\textit{dest}} \newcommand{\RV}{\textit{RV}} \newcommand{\NAME}{\textit{NAME}} \newcommand{\OP}{\textit{OP}} \newcommand{\EXP}{\textit{EXP}} \newcommand\T[1]{{\cal T}[\![#1]\!]} \)

Topics

One aspect of instruction selection we haven't gotten to is instruction selection for functions and function calls. Recall that calls show up in the lowered IR as call statements \(\CALL_m(f, e_1, \dots, e_n)\).

On the x86-64 ISA, we want to implement these IR nodes using the instruction call. It pushes the current instruction pointer (register rip, also known as the program counter) onto the stack and then jumps to the specified destination. Thus, the instruction call f is equivalent to sub rsp, 8; mov [rsp], rip; jmp f. However, call is more compact and faster.

Calling convention

A calling convention is a standardized contract about how to invoke functions. Having a calling convention allows code generated by different compilers and languages to interoperate. Since we want to give you Xi libraries to compile against, it's important that we follow the same calling conventions.

Unfortunately, there are multiple calling conventions for the x86-64. We will describe here the System V style calling convention used by Linux. The Microsoft calling convention is similar but differs in a few details. Note that the full calling convention is more complex than described here, in order to support struct arguments that are larger than one word. But Xi does not have these language features, simplifying matters.

Unlike in the 32-bit architecture, arguments are usually passed to functions entirely in registers. The first six word-size arguments are passed in the registers rdi, rsi, rdx, rcx, r8, and r9, in that order. For functions with more than 6 arguments, the remaining arguments are pushed onto the stack, in reverse order. Using reverse order supports functions with a variable number of arguments, though this is not a consideration in Xi. The figure above shows what the stack looks like as a procedure/function call proceeds. Note that stacks grow downward in this picture, so the top of the stack is at the lowest used address, which is at the bottom!

The code for doing a function call with \(n\) arguments looks something like the following, assuming that arguments 7–\(n\) have been placed in temporaries \(t_7, \dots t_n\):

push tn
push tn-1
...
push t7
mov r9, t6
...
mov rsi, t2
mov rdi, t1
call f
mov dest, rax
add rsp, 8*(n-6) // if n > 6

Note that the stack pointer register rsp always points to the top entry of the stack. So the instruction push x is equivalent to sub rsp, 8; mov [rsp], x, and there is a corresponding instruction pop x that does this opposite: mov x, [rsp]; add rsp, 8.

This code sequence also shows how function results are handled. The result of a function, if any, comes back in the rax register. Once the function completes, it should clean up the stack so the stack pointer is back where it started. Thus, the final instruction adds the appropriate offset to reverse the effect on rsp of all the The code above requires \(n-7\) temporaries to be in use (“live”) at the same time, and it does not handle evaluation of the argument expressions. Both of these problems can be addressed by writing a translation rule in the style that we have been using:

\( \begin{align*} \T{CALL_m(\NAME(f), e_1, \dots, e_n)} = ~&\T{t_1, e_1} \\ &\T{t_2, e_2} \\ & \dots \\ &\T{t_6, e_6} \\ &\T{t_7, e_7} \\ &\T{t_8, e_8} \\ &\dots \\ &\T{t_n, e_n} \\ &\texttt{push}~t_n \\ &\texttt{push}~t_{n-1} \\ & \dots \\ &\texttt{push}~t_7 \\ &\texttt{mov r9},~t_6 \\ & \dots \\ &\texttt{mov rsi},~t_2 \\ &\texttt{mov rdi},~t_1 \\ &\texttt{call}~f \\ &\texttt{mov}~\RV_1, \texttt{rax} \\ &\texttt{mov}~\RV_2, \texttt{rdx} & \text{(if $m=2$)} \\ &\texttt{add rsp}, 8*(n-6) \end{align*} \)

Returning multiple results

For functions that return multiple results, the calling convention specifies that the second result comes back in rdx. The System V calling convention does not give a way to return more than two results. We address this lack in Xi by having the caller provide the space into which to return the additional results. For a call \(\CALL_m\) with \(m > 2\), space is reserved on the caller's stack. The register rdi is used to pass the address of this stack segment, and the regular function argument are passed starting from register rsi instead of rdi.

\( \begin{align*} \T{CALL_m(\NAME(f), e_1, \dots, e_n)} = ~&\T{t_1, e_1} \\ &\T{t_2, e_2} \\ & \dots \\ &\T{t_6, e_6} \\ &\T{t_7, e_7} \\ &\T{t_8, e_8} \\ &\dots \\ &\T{t_n, e_n} \\ &\texttt{sub rsp}, 8*(m-2) \\ &\texttt{mov rdi, rsp} \\ &\texttt{push}~t_n \\ &\texttt{push}~t_{n-1} \\ & \dots \\ &\texttt{push}~t_6 \\ &\texttt{mov r9},~t_5 \\ & \dots \\ &\texttt{mov rsi},~t_1 \\ &\texttt{call}~f \\ &\texttt{add rsp}, 8*(n-5) \\ &\texttt{mov}~\RV_1, \texttt{rax} \\ &\texttt{mov}~\RV_2, \texttt{rdx} & \text{(if $m=2$)} \\ &\texttt{pop}~\RV_3 \\ &\dots \\ &\texttt{pop}~\RV_m \end{align*} \)

As shown, once the arguments have been removed from the stack, the result values are on the top of the stack and can be retrieved easily.

Inside the function

Stack frame of a running function

Once the function is entered, it sets up its stack frame to look like the figure above. The region labeled “temporary storage” is used to store local variables and other temporaries that don't fit into registers. Because the stack pointer can move around, it is common to use a different register, the frame pointer (or base pointer) register rbp. For example, a variable located immediately below the frame pointer would be accessed with the memory operand [rbp-8]. But if rbp is used, it is necessary to save the caller's rbp before overwriting it. So rbp is saved on the stack immediately after the program counter rip of the caller.

Recall that we generate code for the body of a function defined as f(x):τ{s} as simply the statement translation \({\mathcal S}[\![s]\!]\). And the translation of return e is \(\textit{MOVE}(\textit{RV}, {\mathcal E}[\![e]\!]); \textit{RETURN}\), where \(\textit{RV}\) is really a name for rax. Therefore the IR code generated for the declaration f(x:int,y:int):int { return x+y } is something like \(\textit{MOVE}(\textit{RV},\textit{ADD}(\textit{TEMP}(x),\textit{TEMP}(y))); \textit{RETURN}\). With appropriate instruction selection, we should get something like:

mov rax, x
add rax, y
...
ret

But this code does nothing to set up the stack frame or to tear it down, or even to move the function arguments into "x" and "y". These can be achieved by adding a function prologue and epilogue:

f: push rbp
   mov rbp, rsp
   sub rsp, 8*l
   mov x, rdi
   mov y, rsi
   mov rax, x
   add rax, y
   mov rsp, rbp
   pop rbp
   ret

Here we are assuming that the stack frame needs to contain \(\ell\) temporary words.

In fact, the ISA has an instruction to accomplish the first three instructions directly: enter 8*l, 0 saves the frame pointer and adjusts the stack pointer. And the two instructions preceding ret can be accomplished using leave:

f: enter 8*l, 0
   mov x, rdi
   mov y, rsi
   mov rax, x
   add rax, y
   leave
   ret

In this case, the function doesn't really need a frame pointer at all, since it doesn't use it. Then we don't need enter and leave. And a smart register allocator can choose x=rdi, y=rsi, making two mov instructions superfluous:

f: mov rax, rdi
   add rax, rsi
   ret

A couple of details to watch out for: first, the stack pointer rsp is required always to be 16-byte aligned when a call instruction is performed. (This is a requirement of various system libraries.) Since the call instruction itself pushes an 8-byte word onto the stack (rip), the stack pointer is misaligned on entry. In the example function above, this is not a problem, because it is a leaf procedure that doesn't call anything else. Note that to force the stack pointer to be 16-byte aligned, the instruction and rsp, -16 will do the trick!

Second, it is generally inadvisable to access memory below the current stack pointer, because interrupt routines are likely to overwrite anything there. However, the System V (Linux) calling conventions introduce a “red zone” of 16 words that interrupt routines will not overwrite (at least when compiling non-kernel code). Consequently, leaf procedures can use a small amount of local storage without the cost of explictly setting up a stack frame.

Caller-save vs. callee-save

In the previous example, we needed to save rbp before changing it, in order to restore it to its old value when returning to the caller. The contract between the caller and callee was that callee would not change this register. In general, the calling conventions define certain registers as callee-save registers, meaning that a called procedure must restore those registers to their original values before returning. In practice, saving and restoring callee-save registers is done in the function prologue and epilogue respectively. Using callee-save registers is not worth it if the overhead of saving them to memory is more than the speedup achieved by using them. In the current calling conventions, the callee-save registers are rbp, rsp, rbx, and r12r15.

Other registers are designated as caller-save registers, meaning that a given function is permitted to change them arbitrarily. These registers include rax, rcx, rdx, rsi, rdi, and r8r11. Since called functions may overwrite these registers, it is responsibility of the caller to save these registers (if necessary) and to restore them after the call. Clearly, performance is better if these registers are used for values that do not need to be saved and restored. In practice, the saving and restoring of caller-save registers takes place as part of the code for a function call, extending the translation given earlier.

Optimizations

Eliminating the frame pointer

The previous example shows that the frame pointer doesn't need to be used in some leaf procedures. We can more generally avoid using a register to keep track of the frame pointer when the offset between the stack pointer and the frame pointer is known at compile time. Suppose that the compiler knows the offset is \(8*\ell\). Then any memory reference of the form [rbp + k] can be equally well written as [rsp + k'] where the constant \(k'\) is equal to \(k + 8\ell\). Once all such references to rbp are removed from the code, there is no need to use rbp, and therefore no need to save it and restore it in the prologue and epilogue. If it is used, it can be used as a general-purpose register!

Although this trick is rather nice, it does mean that the distance between the two stack pointers must be statically known. This rules out using dynamic allocation on the stack, including the declaration of on-stack arrays with dynamically determined length or the alloca() system call familiar to experienced C programmers.

Exploiting the red zone

Leaf procedures whose temporaries fit entirely in the red zone can simply use the red zone instead of allocating a stack frame. In this case, the stack pointer is never changed, and all temporaries are located in the red zone at a positive offset from the stack pointer.

Trivial register allocation

At this point, the translation has produced abstract assembly code, in which temporaries have been mapped to a set of registers of unbounded size. The set of registers used includes not only the real registers supported by the ISA, but also pseudo-registers named like the corresponding temporaries. The job of register allocation is to replace these abstract registers with real registers, rewriting the code as necessary.

Doing a good job of register allocation is challenging. However, it is not hard to generate code if all temporaries (abstract registers) are assigned to stack locations. For example, gcc -O0 uses this approach.

Temporaries are assigned to distinct stack locations (e.g., [rbp-8], [rbp-16], etc.). Any given instruction uses at most 3 register operands, so 3 registers are reserved for the purpose of getting values onto and off of the stack. It makes sense to select three caller-save registers, such as rax, rcx, and rdx. In general, instructions are inserted before each abstract assembly instruction, to read operand registers from the stack, and further instructions are inserted afterward, to write results back to the stack. Essentially, each abstract register is allocated to one of these three reserved registers for the duration of a single assembly instruction.

Suppose that we have allocated temporary t1 to location [rbp-8]. Then the abstract assembly instruction push t1 is converted to concrete assembly that first loads the operand into one of the three reserved registers, and then performs the original instruction with that reserved register serving the role of the temporary:

mov rax, [rbp-8]
push rax

For instructions that update temporaries, instructions are added afterward to move the new values of the temporaries into the affected stack locations. So the transformation of the instruction add t2, [t1+8] might be the following, assuming that t1 is located at [rbp-8] and t2 is located at [rbp-16]:

mov rcx, [rbp-8]
mov rax, [rbp-16]
add rax, [rcx+8]
mov [rbp-16], rax

A couple of simple optimizations can help. On CISC instruction sets like x86-64, many instructions can read operands directly from memory or write results directly to memory, avoiding the need to add some instructions. For example, the first conversion above could have been done just as push [rbp-8]. Also, a few temporaries can be allocated directly to registers, as long as three registers are reserved for the job of shuttling temporaries on and off the stack. For example, the callee-save registers rbx and r12r15 are not serving any special purpose yet; nor are the caller-save registers r10–r11. Of course, using these registers comes with an obligation to save and restore their contents under certain conditions.