Remove colour attributes from body and strip most of the font tags out

This commit is contained in:
James Gregory 2013-12-30 13:57:02 +11:00
commit c1f88ddb41
362 changed files with 1632 additions and 1706 deletions

View file

@ -24,7 +24,7 @@
<!--CHAPTER=12//-->
<!--PAGES=243-246//-->
<!--UNASSIGNED1//-->
<!--UNASSIGNED2//--></HEAD><BODY LINK=#0000FF ALINK=#000099 VLINK=#0000FF BGCOLOR=#FFFFFF>
<!--UNASSIGNED2//--></HEAD><body>
<CENTER>
<TABLE BORDER>
@ -38,7 +38,7 @@
<P><BR></P>
<P>You don&rsquo;t need to understand every corner of the 486 universe unless you&rsquo;re a diehard ASMhead who does this stuff for fun. Just learn enough to be able to speed up the key portions of your programs, and spend the rest of your time on a fast design and overall implementation.
</P>
<H4 ALIGN="LEFT"><A NAME="Heading10"></A><FONT COLOR="#000077">More Fun with Byte Registers</FONT></H4>
<H4 ALIGN="LEFT"><A NAME="Heading10"></A>More Fun with Byte Registers</H4>
<P>Rule #4: Don&rsquo;t load <I>any</I> byte register exactly 2 cycles before using <I>any</I> register to address memory.</P>
<P>This, the last of this chapter&rsquo;s rules, is the strangest of the lot. If any byte register is loaded, and then two cycles later any register is used to point to memory, one cycle is lost. So, for example, this code</P>
<!-- CODE SNIP //-->
@ -78,7 +78,7 @@ mov ax,[bx]
<P>A more sophisticated programmer would expect to lose one cycle, because BX is loaded two cycles before being used to address memory. In fact, though, this code takes 5 cycles&mdash;2 cycles, or 67 percent, longer than normal. Why? Well, under normal conditions, loading a byte register&mdash;CL in this case&mdash;one cycle before using a register to address memory produces no penalty; loading 2 cycles ahead is the only case that normally incurs a penalty. However, think of Rule #4 as meaning that loading a byte register disrupts the memory addressing pipeline as it starts up. Viewed that way, we can see that <B>MOV BX,OFFSET MemVar</B> interrupts the addressing pipeline, forcing it to start again, and then, presumably, <B>MOV CL,AL</B> interrupts the pipeline again because the pipeline is now on its first cycle: the one that loading a byte register can affect.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>I know&mdash;it seems awfully complicated. It isn&rsquo;t, really. Generally, try not to use byte destinations exactly two cycles before using a register to address memory, and try not to load a register either one or two cycles before using it to address memory, and you&rsquo;ll be fine.</I></SMALL>
</TABLE>
<H4 ALIGN="LEFT"><A NAME="Heading11"></A><FONT COLOR="#000077">Timing Your Own 486 Code</FONT></H4>
<H4 ALIGN="LEFT"><A NAME="Heading11"></A>Timing Your Own 486 Code</H4>
<P>In case you want to do some 486 performance analysis of your own, let me show you how I arrived at one of the above conclusions; at the same time, I can warn you of the timing hazards of the cache. Listings 12.1 and 12.2 show the code I ran through the Zen timer in order to establish the effects of loading a byte register before using a register to address memory. Listing 12.1 ran in 120 &micro;s on a 33 MHz 486, or 4 cycles per repetition (120 &micro;s/1000 repetitions = 120 ns per repetition; 120 ns per repetition/30 ns per cycle = 4 cycles per repetition); Listing 12.2 ran in 90 &micro;s, or 3 cycles, establishing that loading a byte register costs a cycle only when it&rsquo;s performed exactly 2 cycles before addressing memory.
</P>
<P><B>LISTING 12.1 LST12-1.ASM</B></P>
@ -128,7 +128,7 @@ Done:
<P>Note that Listings 12.1 and 12.2 each repeat the timing of the code under test a second time, to make sure that the instructions are in the cache on the second pass, the one for which results are displayed. Also note that the code is less than 8K in size, so that it can all fit in the 486&rsquo;s 8K internal cache. If I double the <B>REPT</B> value in Listing 12.2 to 2,000, making the test code larger than 8K, the execution time more than doubles to 224 &micro;s, or 3.7 cycles per repetition; the extra seven-tenths of a cycle comes from fetching non-cached instruction bytes.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Whenever you see non-integral timing results of this sort, it&rsquo;s a good bet that the test code or data isn&rsquo;t cached.</I></SMALL>
</TABLE>
<H3><A NAME="Heading12"></A><FONT COLOR="#000077">The Story Continues</FONT></H3>
<H3><A NAME="Heading12"></A>The Story Continues</H3>
<P>There&rsquo;s certainly plenty more 486 lore to explore, including the 486&rsquo;s unique prefetch queue, more optimization rules, branching optimizations, performance implications of the cache, the cost of cache misses for reads, and the implications of cache write-through for writes. Nonetheless, we&rsquo;ve covered quite a bit of ground in this chapter, and I trust you&rsquo;ve gotten a feel for the considerable extent to which 486 optimization differs from what you&rsquo;re used to. Odd as 486 optimization is, though, it&rsquo;s well worth mastering, for the 486 is, at its best, so staggeringly fast that carefully crafted 486 code can do more than twice as much per cycle as the best 386 code&mdash;which makes it perhaps 50 times as fast as optimized code for the original PC.
</P>
<P>Sometimes it <I>is</I> hard to believe we&rsquo;re still in Kansas!</P><P><BR></P>
@ -144,7 +144,7 @@ Done:
<hr width="90%" size="1" noshade>
<div align="center">
<font face="Verdana,sans-serif" size="1">Graphics Programming Black Book &copy; 2001 Michael Abrash</font>
Graphics Programming Black Book &copy; 2001 Michael Abrash
</div>
<!-- all of the reference materials (books) have the footer and subfoot reveresed -->
<!-- reference_subfoot = footer -->