Remove colour attributes from body and strip most of the font tags out

This commit is contained in:
James Gregory 2013-12-30 13:57:02 +11:00
commit c1f88ddb41
362 changed files with 1632 additions and 1706 deletions

View file

@ -24,7 +24,7 @@
<!--CHAPTER=63//-->
<!--PAGES=1168-1170//-->
<!--UNASSIGNED1//-->
<!--UNASSIGNED2//--></HEAD><BODY LINK=#0000FF ALINK=#000099 VLINK=#0000FF BGCOLOR=#FFFFFF>
<!--UNASSIGNED2//--></HEAD><body>
<CENTER>
<TABLE BORDER>
@ -55,7 +55,7 @@ FST [temp]
<!-- END CODE SNIP //-->
<P>takes 6 cycles in all.) Again, it&rsquo;s possible to execute integer-unit instructions during the 2 (or 3, for FST) cycles after one of these FP instructions starts. There&rsquo;s a more exciting possibility here, though: Given properly structured code, the FPU is capable of averaging 1 cycle per FADD, FSUB, or FMUL. The secret is pipelining.
</P>
<H4 ALIGN="LEFT"><A NAME="Heading5"></A><FONT COLOR="#000077">Pipelining, Latency, and Throughput</FONT></H4>
<H4 ALIGN="LEFT"><A NAME="Heading5"></A>Pipelining, Latency, and Throughput</H4>
<P>The Pentium&rsquo;s FPU is the first pipelined x86 FPU. <I>Pipelining</I> means that the FPU is capable of starting an instruction every cycle, and can simultaneously handle several instructions in various stages of completion. Only certain x86 FP instructions allow another instruction to start on the next cycle, though: FADD, FSUB, and FMUL are pipelined, but FST and FDIV are not. (FLD executes in a single cycle, so pipelining is not an issue.) Thus, in the code sequence</P>
<!-- CODE SNIP //-->
<PRE>
@ -100,7 +100,7 @@ FSUB ST(0),ST(1)
<!-- END CODE SNIP //-->
<P>where the ST(0) operand to FSUB is calculated by FADD. Here, FSUB can&rsquo;t start until FADD has completed, so there are 2 stall cycles between the two instructions. When dependencies like this occur, the FPU runs at latency rather than throughput speeds, and performance can drop by as much as two-thirds.
</P>
<H4 ALIGN="LEFT"><A NAME="Heading6"></A><FONT COLOR="#000077">FXCH</FONT></H4>
<H4 ALIGN="LEFT"><A NAME="Heading6"></A>FXCH</H4>
<P>One piece of the puzzle is still missing. Clearly, to get maximum throughput, we need to interleave FP instructions, such that at any one time ideally three instructions are in the pipeline at once. Further, these instructions must not depend on one another for operands. But ST(0) must always be one of the operands; worse, FLD can only push into ST(0), and FST can only store from ST(0). How, then, can we keep three independent instructions going?
</P>
<P>The easy answer would be for Intel to change the FP registers from a stack to a set of independent registers. Since they couldn&rsquo;t do that, thanks to compatibility issues, they did the next best thing: They made the FXCH instruction, which swaps ST(0) and any other FP register, virtually free. In general, if FXCH is both preceded and followed by FP instructions, then it takes <I>no</I> cycles to execute. (Application Note 500, &ldquo;Optimizations for Intel&rsquo;s 32-bit Processors,&rdquo; February 1994, available from <A HREF="http://www.intel.com">http://www.intel.com</A>, describes all .the conditions under which FXCH is free.) This allows you to move the target of a pending operation from ST(0) to another register, at the same time bringing another register into ST(0) where it can be used, all at no cost. So, for example, we can start three multiplications, then use FXCH to swap back to start adding the results of the first two multiplications, without incurring any stalls, as shown in Listing 63.1.</P>
@ -118,7 +118,7 @@ FSUB ST(0),ST(1)
faddp st(2),st(0) ;starts on cycle 6
</PRE>
<!-- END CODE //-->
<H3><A NAME="Heading7"></A><FONT COLOR="#000077">The Dot Product</FONT></H3>
<H3><A NAME="Heading7"></A>The Dot Product</H3>
<P>Now we&rsquo;re ready to look at fast FP for common 3-D operations; we&rsquo;ll start by looking at how to speed up the dot product. As discussed in Chapter 30, the dot product is heavily used in 3-D to calculate cosines and to project points along vectors. The dot product is calculated as d = u<SUB>1</SUB>v<SUB>1</SUB> + u<SUB>2</SUB>v<SUB>2</SUB> + u<SUB>3</SUB>v<SUB>3</SUB>; with three loads, three multiplies, two adds, and a store, the theoretical minimum time for this calculation is 10 cycles.</P>
<P>Listing 63.2 shows a straightforward dot product implementation. This version loses 7 cycles to stalls. Listing 63.3 cuts the loss to 5 cycles by doing all three FMULs first, then using FXCH to set the third FXCH aside to complete while the results of the first two FMULs, which have completed, are added. Listing 43.3 still loses 50 percent to stalls, but unless some other code is available to be interleaved with the dot product code, that&rsquo;s all we can do to speed things up. Fortunately, dot products are often used in contexts where there&rsquo;s plenty of interleaving potential, as we&rsquo;ll see when we discuss transformation.</P><P><BR></P>
<CENTER>
@ -133,7 +133,7 @@ FSUB ST(0),ST(1)
<hr width="90%" size="1" noshade>
<div align="center">
<font face="Verdana,sans-serif" size="1">Graphics Programming Black Book &copy; 2001 Michael Abrash</font>
Graphics Programming Black Book &copy; 2001 Michael Abrash
</div>
<!-- all of the reference materials (books) have the footer and subfoot reveresed -->
<!-- reference_subfoot = footer -->