From 0f5f617adbf9ecf6d1fa0820b3df7096a7888ccf Mon Sep 17 00:00:00 2001 From: James Gregory Date: Mon, 30 Dec 2013 15:11:04 +1100 Subject: [PATCH] Strip comments and inline ids onto h* --- 01-01.html | 23 +++++----------- 01-02.html | 27 ++++++------------- 01-03.html | 25 +++++------------ 01-04.html | 23 +++++----------- 01-05.html | 25 +++++------------ 01-06.html | 23 +++++----------- 02-01.html | 31 +++++++-------------- 02-02.html | 31 +++++++-------------- 02-03.html | 19 +++---------- 03-01.html | 23 +++++----------- 03-02.html | 19 +++---------- 03-03.html | 26 +++++------------- 03-04.html | 19 +++---------- 03-05.html | 27 ++++++------------- 03-06.html | 29 +++++++------------- 03-07.html | 21 ++++----------- 03-08.html | 21 ++++----------- 03-09.html | 33 ++++++++--------------- 03-10.html | 43 +++++++++++------------------ 04-01.html | 27 ++++++------------- 04-02.html | 53 ++++++++++++++---------------------- 04-03.html | 41 +++++++++++----------------- 04-04.html | 32 +++++++--------------- 04-05.html | 32 +++++++--------------- 04-06.html | 26 +++++------------- 04-07.html | 29 +++++++------------- 04-08.html | 32 +++++++--------------- 04-09.html | 25 +++++------------ 04-10.html | 21 ++++----------- 05-01.html | 23 +++++----------- 05-02.html | 31 +++++++-------------- 05-03.html | 19 +++---------- 05-04.html | 19 +++---------- 05-05.html | 17 +++--------- 06-01.html | 29 +++++++------------- 06-02.html | 67 +++++++++++++++++++-------------------------- 07-01.html | 21 ++++----------- 07-02.html | 29 +++++++------------- 07-03.html | 21 ++++----------- 07-04.html | 19 +++---------- 07-05.html | 49 +++++++++++++-------------------- 08-01.html | 23 +++++----------- 08-02.html | 34 ++++++++--------------- 08-03.html | 28 ++++++------------- 08-04.html | 27 ++++++------------- 08-05.html | 30 +++++++-------------- 09-01.html | 55 +++++++++++++++---------------------- 09-02.html | 40 ++++++++++----------------- 09-03.html | 24 +++++------------ 09-04.html | 23 +++++----------- 09-05.html | 28 ++++++------------- 09-06.html | 25 +++++------------ 09-07.html | 50 +++++++++++++--------------------- 10-01.html | 23 +++++----------- 10-02.html | 30 +++++++-------------- 10-03.html | 37 +++++++++---------------- 10-04.html | 19 +++---------- 11-01.html | 33 ++++++++--------------- 11-02.html | 21 ++++----------- 11-03.html | 32 +++++++--------------- 11-04.html | 45 +++++++++++-------------------- 11-05.html | 19 +++---------- 11-06.html | 23 +++++----------- 11-07.html | 21 ++++----------- 11-08.html | 46 +++++++++++-------------------- 12-01.html | 33 ++++++++--------------- 12-02.html | 53 ++++++++++++++---------------------- 12-03.html | 47 +++++++++++++------------------- 12-04.html | 45 ++++++++++++------------------- 13-01.html | 34 ++++++++--------------- 13-02.html | 38 +++++++++----------------- 13-03.html | 35 +++++++++--------------- 13-04.html | 29 +++++++------------- 14-01.html | 21 ++++----------- 14-02.html | 34 +++++++---------------- 14-03.html | 15 ++--------- 14-04.html | 23 +++++----------- 14-05.html | 19 +++---------- 14-06.html | 23 +++++----------- 15-01.html | 26 +++++------------- 15-02.html | 43 +++++++++++------------------ 15-03.html | 35 ++++++++---------------- 15-04.html | 29 +++++++------------- 16-01.html | 25 +++++------------ 16-02.html | 25 +++++------------ 16-03.html | 21 ++++----------- 16-04.html | 25 +++++------------ 16-05.html | 19 +++---------- 16-06.html | 19 +++---------- 16-07.html | 21 ++++----------- 16-08.html | 37 +++++++++---------------- 17-01.html | 23 +++++----------- 17-02.html | 23 +++++----------- 17-03.html | 19 +++---------- 17-04.html | 29 ++++++-------------- 17-05.html | 21 ++++----------- 17-06.html | 26 +++++------------- 17-07.html | 19 +++---------- 17-08.html | 17 +++--------- 18-01.html | 21 ++++----------- 18-02.html | 17 +++--------- 18-03.html | 23 +++++----------- 18-04.html | 34 ++++++++--------------- 18-05.html | 25 +++++------------ 19-01.html | 23 +++++----------- 19-02.html | 27 ++++++------------- 19-03.html | 21 ++++----------- 19-04.html | 23 +++++----------- 20-01.html | 26 +++++------------- 20-02.html | 49 ++++++++++++--------------------- 20-03.html | 43 +++++++++++------------------ 20-04.html | 42 ++++++++++------------------- 21-01.html | 39 +++++++++------------------ 21-02.html | 69 ++++++++++++++++++++--------------------------- 21-03.html | 21 ++++----------- 21-04.html | 23 +++++----------- 21-05.html | 25 +++++------------ 22-01.html | 25 +++++------------ 22-02.html | 27 ++++++------------- 22-03.html | 27 ++++++------------- 23-01.html | 23 +++++----------- 23-02.html | 29 +++++++------------- 23-03.html | 26 +++++------------- 23-04.html | 22 ++++----------- 23-05.html | 19 +++---------- 23-06.html | 21 ++++----------- 24-01.html | 26 +++++------------- 24-02.html | 19 +++---------- 24-03.html | 25 +++++------------ 25-01.html | 33 +++++++---------------- 25-02.html | 19 +++---------- 25-03.html | 22 ++++----------- 25-04.html | 21 ++++----------- 25-05.html | 25 +++++------------ 25-06.html | 19 +++---------- 26-01.html | 21 ++++----------- 26-02.html | 19 +++---------- 26-03.html | 21 ++++----------- 27-01.html | 28 ++++++------------- 27-02.html | 21 ++++----------- 27-03.html | 21 ++++----------- 27-04.html | 21 ++++----------- 27-05.html | 19 +++---------- 28-01.html | 21 ++++----------- 28-02.html | 19 +++---------- 28-03.html | 17 +++--------- 28-04.html | 21 ++++----------- 28-05.html | 19 +++---------- 29-01.html | 30 +++++++-------------- 29-02.html | 19 +++---------- 29-03.html | 26 +++++------------- 29-04.html | 24 +++++------------ 29-05.html | 23 +++++----------- 29-06.html | 26 +++++------------- 30-01.html | 28 ++++++------------- 30-02.html | 19 +++---------- 30-03.html | 21 ++++----------- 30-04.html | 19 +++---------- 30-05.html | 19 +++---------- 30-06.html | 17 +++--------- 30-07.html | 19 +++---------- 31-01.html | 25 +++++------------ 31-02.html | 22 ++++----------- 31-03.html | 19 +++---------- 31-04.html | 17 +++--------- 31-05.html | 21 ++++----------- 32-01.html | 23 +++++----------- 32-02.html | 19 +++---------- 32-03.html | 19 +++---------- 32-04.html | 21 ++++----------- 32-05.html | 22 ++++----------- 33-01.html | 30 +++++++-------------- 33-02.html | 21 ++++----------- 33-03.html | 21 ++++----------- 33-04.html | 19 +++---------- 34-01.html | 23 +++++----------- 34-02.html | 19 +++---------- 34-03.html | 21 ++++----------- 34-04.html | 17 +++--------- 34-05.html | 23 +++++----------- 35-01.html | 28 ++++++------------- 35-02.html | 27 +++++-------------- 35-03.html | 21 ++++----------- 35-04.html | 26 +++++------------- 35-05.html | 24 +++++------------ 35-06.html | 19 +++---------- 35-07.html | 19 +++---------- 36-01.html | 34 ++++++++--------------- 36-02.html | 32 +++++++--------------- 36-03.html | 26 +++++------------- 36-04.html | 19 +++---------- 37-01.html | 25 +++++------------ 37-02.html | 19 +++---------- 38-01.html | 33 +++++++---------------- 38-02.html | 27 ++++++------------- 38-03.html | 23 +++++----------- 38-04.html | 21 ++++----------- 39-01.html | 23 +++++----------- 39-02.html | 21 ++++----------- 39-03.html | 21 ++++----------- 39-04.html | 23 +++++----------- 39-05.html | 19 +++---------- 40-01.html | 33 +++++++---------------- 40-02.html | 24 +++++------------ 40-03.html | 29 +++++++------------- 40-04.html | 23 +++++----------- 40-05.html | 19 +++---------- 41-01.html | 25 +++++------------ 41-02.html | 29 ++++++-------------- 41-03.html | 19 +++---------- 41-04.html | 23 +++++----------- 42-01.html | 26 +++++------------- 42-02.html | 26 +++++------------- 42-03.html | 25 +++++------------ 42-04.html | 23 +++++----------- 42-05.html | 21 ++++----------- 43-01.html | 36 ++++++++----------------- 43-02.html | 24 +++++------------ 43-03.html | 19 +++---------- 43-04.html | 17 +++--------- 43-05.html | 22 ++++----------- 43-06.html | 17 +++--------- 44-01.html | 23 +++++----------- 44-02.html | 19 +++---------- 44-03.html | 19 +++---------- 44-04.html | 31 +++++++-------------- 44-05.html | 22 ++++----------- 44-06.html | 24 +++++------------ 45-01.html | 23 +++++----------- 45-02.html | 29 ++++++-------------- 45-03.html | 21 ++++----------- 45-04.html | 21 ++++----------- 45-05.html | 23 +++++----------- 45-06.html | 17 +++--------- 46-01.html | 21 ++++----------- 46-02.html | 23 +++++----------- 46-03.html | 23 +++++----------- 47-01.html | 21 ++++----------- 47-02.html | 21 ++++----------- 47-03.html | 28 ++++++------------- 47-04.html | 21 ++++----------- 47-05.html | 19 +++---------- 47-06.html | 22 ++++----------- 47-07.html | 23 +++++----------- 48-01.html | 33 +++++++---------------- 48-02.html | 19 +++---------- 48-03.html | 24 +++++------------ 48-04.html | 19 +++---------- 48-05.html | 23 +++++----------- 49-01.html | 25 +++++------------ 49-02.html | 21 ++++----------- 49-03.html | 25 +++++------------ 49-04.html | 28 ++++++------------- 49-05.html | 21 ++++----------- 50-01.html | 23 +++++----------- 50-02.html | 48 +++++++++++---------------------- 50-03.html | 19 +++---------- 50-04.html | 23 +++++----------- 50-05.html | 19 +++---------- 50-06.html | 19 +++---------- 50-07.html | 19 +++---------- 51-01.html | 21 ++++----------- 51-02.html | 27 +++++-------------- 51-03.html | 19 +++---------- 51-04.html | 33 +++++++---------------- 51-05.html | 21 ++++----------- 51-06.html | 23 +++++----------- 52-01.html | 21 ++++----------- 52-02.html | 23 +++++----------- 52-03.html | 27 ++++++------------- 52-04.html | 23 +++++----------- 52-05.html | 19 +++---------- 52-06.html | 23 +++++----------- 52-07.html | 19 +++---------- 52-08.html | 19 +++---------- 53-01.html | 21 ++++----------- 53-02.html | 19 +++---------- 53-03.html | 24 +++++------------ 53-04.html | 23 +++++----------- 54-01.html | 21 ++++----------- 54-02.html | 19 +++---------- 54-03.html | 31 +++++++-------------- 54-04.html | 19 +++---------- 54-05.html | 27 +++++-------------- 55-01.html | 30 +++++++-------------- 55-02.html | 19 +++---------- 55-03.html | 26 +++++------------- 55-04.html | 24 +++++------------ 56-01.html | 33 +++++++---------------- 56-02.html | 32 +++++++--------------- 56-03.html | 25 +++++------------ 57-01.html | 23 +++++----------- 57-02.html | 30 +++++++-------------- 57-03.html | 19 +++---------- 57-04.html | 15 ++--------- 58-01.html | 28 ++++++------------- 58-02.html | 30 +++++++-------------- 58-03.html | 45 +++++++++++-------------------- 58-04.html | 28 ++++++------------- 58-05.html | 25 +++++------------ 59-01.html | 23 +++++----------- 59-02.html | 29 ++++++-------------- 59-03.html | 51 ++++++++++++----------------------- 59-04.html | 36 +++++++++---------------- 59-05.html | 21 ++++----------- 59-06.html | 19 +++---------- 60-01.html | 21 ++++----------- 60-02.html | 36 ++++++++----------------- 60-03.html | 19 +++---------- 60-04.html | 19 +++---------- 61-01.html | 23 +++++----------- 61-02.html | 31 +++++++-------------- 61-03.html | 32 +++++++--------------- 61-04.html | 42 ++++++++++------------------- 62-01.html | 26 +++++------------- 62-02.html | 19 +++---------- 62-03.html | 35 ++++++++---------------- 62-04.html | 24 +++++------------ 63-01.html | 23 +++++----------- 63-02.html | 53 +++++++++++++++--------------------- 63-03.html | 43 +++++++++++------------------ 63-04.html | 25 +++++------------ 64-01.html | 21 ++++----------- 64-02.html | 38 +++++++++----------------- 64-03.html | 41 ++++++++++------------------ 64-04.html | 21 ++++----------- 65-01.html | 21 ++++----------- 65-02.html | 33 +++++++---------------- 65-03.html | 30 +++++++-------------- 65-04.html | 21 ++++----------- 66-01.html | 27 ++++++------------- 66-02.html | 29 ++++++-------------- 66-03.html | 27 +++++-------------- 66-04.html | 21 ++++----------- 67-01.html | 21 ++++----------- 67-02.html | 31 +++++++-------------- 67-03.html | 23 +++++----------- 67-04.html | 15 ++--------- 67-05.html | 17 +++--------- 68-01.html | 25 +++++------------ 68-02.html | 31 +++++++-------------- 68-03.html | 24 +++++------------ 68-04.html | 24 +++++------------ 69-01.html | 21 ++++----------- 69-02.html | 23 +++++----------- 69-03.html | 26 +++++------------- 69-04.html | 26 +++++------------- 70-01.html | 19 +++---------- 70-02.html | 17 +++--------- 70-03.html | 19 +++---------- 70-04.html | 21 ++++----------- 70-05.html | 21 ++++----------- 70-06.html | 25 +++++------------ 70-07.html | 21 ++++----------- 70-08.html | 17 +++--------- 70-09.html | 17 +++--------- about.html | 19 +++---------- about_author.html | 18 +++---------- appendix-a.html | 18 +++---------- book-index.html | 18 +++---------- index.html | 15 ++--------- intro.html | 19 +++---------- 362 files changed, 2519 insertions(+), 6694 deletions(-) diff --git a/01-01.html b/01-01.html index a121ea0..7a32bb0 100644 --- a/01-01.html +++ b/01-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - @@ -39,10 +32,10 @@

Part I

-

Chapter 1
+

Chapter 1
The Best Optimizer Is between Your Ears

-

The Human Element of Code Optimization

+

The Human Element of Code Optimization

This book is devoted to a topic near and dear to my heart: writing software that pushes PCs to the limit. Given run-of-the-mill software, PCs run like the 97-pound-weakling minicomputers they are. Give them the proper care, however, and those ugly boxes are capable of miracles. The key is this: Only on microcomputers do you have the run of the whole machine, without layers of operating systems, drivers, and the like getting in the way. You can do anything you want, and you can understand everything that’s going on, if you so wish.

@@ -56,7 +49,7 @@

...now.

-

Understanding High Performance

+

Understanding High Performance

Before we can create high-performance code, we must understand what high performance is. The objective (not always attained) in creating high-performance software is to make the software able to carry out its appointed tasks so rapidly that it responds instantaneously, as far as the user is concerned. In other words, high-performance code should ideally run so fast that any further improvement in the code would be pointless.

@@ -76,7 +69,7 @@

“What’s a fast slow program?” you ask. That’s a good question, and a brief (true) story is perhaps the best answer.

-

When Fast Isn’t Fast

+

When Fast Isn’t Fast

In the early 1970s, as the first hand-held calculators were hitting the market, I knew a fellow named Irwin. He was a good student, and was planning to be an engineer. Being an engineer back then meant knowing how to use a slide rule, and Irwin could jockey a slipstick with the best of them. In fact, he was so good that he challenged a fellow with a calculator to a duel—and won, becoming a local legend in the process.

@@ -101,10 +94,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/01-02.html b/01-02.html index 52faaae..37ca79a 100644 --- a/01-02.html +++ b/01-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - @@ -37,7 +30,7 @@


-

Rules for Building High-Performance Code

+

Rules for Building High-Performance Code

We’ve got the following rules for creating high-performance software:

@@ -59,15 +52,15 @@

Making rules is easy; the hard part is figuring out how to apply them in the real world. For my money, examining some actual working code is always a good way to get a handle on programming concepts, so let’s look at some of the performance rules in action.

-

Know Where You’re Going

+

Know Where You’re Going

If we’re going to create high-performance code, first we have to know what that code is going to do. As an example, let’s write a program that generates a 16-bit checksum of the bytes in a file. In other words, the program will add each byte in a specified file in turn into a 16-bit value. This checksum value might be used to make sure that a file hasn’t been corrupted, as might occur during transmission over a modem or if a Trojan horse virus rears its ugly head. We’re not going to do anything with the checksum value other than print it out, however; right now we’re only interested in generating that checksum value as rapidly as possible.

-

Make a Big Map

+

Make a Big Map

How are we going to generate a checksum value for a specified file? The logical approach is to get the file name, open the file, read the bytes out of the file, add them together, and print the result. Most of those actions are straightforward; the only tricky part lies in reading the bytes and adding them together.

-

Make Lots of Little Maps

+

Make Lots of Little Maps

Actually, we’re only going to make one little map, because we only have one program section that requires much thought—the section that reads the bytes and adds them up. What’s the best way to do this?

@@ -79,7 +72,7 @@

It’s slow.

-

LISTING 1.1 L1-1.C

+

LISTING 1.1 L1-1.C

 /*
 * Program to calculate the 16-bit checksum of all bytes in the
@@ -121,7 +114,7 @@ main(int argc, char *argv[]) {
      printf(“The checksum is: %u\n”, Checksum);
      exit(0);
 }
-
+

Table 1.1 shows the time taken for Listing 1.1 to generate a checksum of the WordPerfect version 4.2 thesaurus file, TH.WP (362,293 bytes in size), on a 10 MHz AT machine of no special parentage. Execution times are given for Listing 1.1 compiled with Borland and Microsoft compilers, with optimization both on and off; all four times are pretty much the same, however, and all are much too slow to be acceptable. Listing 1.1 requires over two and one-half minutes to checksum one file!

@@ -152,10 +145,6 @@ main(int argc, char *argv[]) {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/01-03.html b/01-03.html index 69df28d..f03fb50 100644 --- a/01-03.html +++ b/01-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - @@ -165,7 +158,7 @@ -

LISTING 1.2 L1-2.C

+

LISTING 1.2 L1-2.C

 /*
 * Program to calculate the 16-bit checksum of the stream of bytes
@@ -199,9 +192,9 @@ main(int argc, char *argv[]) {
       printf(“The checksum is: %u\n”, Checksum);
       exit(0);
 }
-
+ -

LISTING 1.3 L1-3.ASM

+

LISTING 1.3 L1-3.ASM

 ; Assembler subroutine to perform a 16-bit checksum on the file
 ; opened on the passed-in handle. Stores the result in the
@@ -267,13 +260,13 @@ Done:
                  ret
 _ChecksumFileendp
                  end
-
+

The lesson is clear: Optimization makes code faster, but without proper design, optimization just creates fast slow code.

Well, then, how are we going to improve our design? Before we can do that, we have to understand what’s wrong with the current design.

-

Know the Territory

+

Know the Territory

Just why is Listing 1.1 so slow? In a word: overhead. The C library implements the read() function by calling DOS to read the desired number of bytes. (I figured this out by watching the code execute with a debugger, but you can buy library source code from both Microsoft and Borland.) That means that Listing 1.1 (and Listing 1.3 as well) executes one DOS function per byte processed—and DOS functions, especially this one, come with a lot of overhead.

@@ -302,10 +295,6 @@ _ChecksumFileendp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/01-04.html b/01-04.html index 782e3a8..5807c2a 100644 --- a/01-04.html +++ b/01-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - @@ -47,7 +40,7 @@

In this case that means knowing how DOS and the C/C++ file-access libraries do their work. In other words, know the territory!

-

LISTING 1.4 L1-4.C

+

LISTING 1.4 L1-4.C

 /*
 * Program to calculate the 16-bit checksum of the stream of bytes
@@ -82,9 +75,9 @@ main(int argc, char *argv[]) {
       printf(“The checksum is: %u\n”, Checksum);
       exit(0);
 }
-
+ -

Know When It Matters

+

Know When It Matters

The last section contained a particularly interesting phrase: the time-critical portions of your code. Time-critical portions of your code are those portions in which the speed of the code makes a significant difference in the overall performance of your program—and by “significant,” I don’t mean that it makes the code 100 percent faster, or 200 percent, or any particular amount at all, but rather that it makes the program more responsive and/or usable from the user’s perspective.

@@ -100,7 +93,7 @@ main(int argc, char *argv[]) {

Besides, we don’t want to optimize until the design is refined to our satisfaction, and that won’t be the case until we’ve thought about other approaches.

-

Always Consider the Alternatives

+

Always Consider the Alternatives

Listing 1.4 is good, but let’s see if there are other—perhaps less obvious—ways to get the same results faster. Let’s start by considering why Listing 1.4 is so much better than Listing 1.1. Like read(), getc() calls DOS to read from the file; the speed improvement of Listing 1.4 over Listing 1.1 occurs because getc() eads many bytes at once via DOS, then manages those bytes for us. That’s faster than reading them one at a time using read()—but there’s no reason to think that it’s faster than having our program read and manage blocks itself. Easier, yes, but not faster.

@@ -139,10 +132,6 @@ main(int argc, char *argv[]) {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/01-05.html b/01-05.html index 284bcfb..c127789 100644 --- a/01-05.html +++ b/01-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - @@ -47,7 +40,7 @@ -

LISTING 1.5 L1-5.C

+

LISTING 1.5 L1-5.C

 /*
 * Program to calculate the 16-bit checksum of the stream of bytes
@@ -105,7 +98,7 @@ main(int argc, char *argv[]) {
       printf(“The checksum is: %u\n”, Checksum);
       exit(0);
 }
-
+

That brings us to the fourth reason: avoiding an internal-buffered implementation like Listing 1.5 because of the difficulty of coding such an approach. True, it is easier to let a C library function do the work, but it’s not all that hard to do the buffering internally. The key is the concept of handling data in restartable blocks; that is, reading a chunk of data, operating on the data until it runs out, suspending the operation while more data is read in, and then continuing as though nothing had happened.

@@ -113,11 +106,11 @@ main(int argc, char *argv[]) {

At any rate, Listing 1.5 isn’t much more complicated than Listing 1.4—and it’s a lot faster. Always consider the alternatives; a bit of clever thinking and program redesign can go a long way.

-

Know How to Turn On the Juice

+

Know How to Turn On the Juice

I have said time and again that optimization is pointless until the design is settled. When that time comes, however, optimization can indeed make a significant difference. Table 1.1 indicates that the optimized version of Listing 1.5 produced by Microsoft C outperforms an unoptimized version of the same code by more than 60 percent. What’s more, a mostly-assembly version of Listing 1.5, shown in Listings 1.6 and 1.7, outperforms even the best-optimized C version of List1.5 by 26 percent. These are considerable improvements, well worth pursuing—once the design has been maxed out.

-

LISTING 1.6 L1-6.C

+

LISTING 1.6 L1-6.C

 /*
 * Program to calculate the 16-bit checksum of the stream of bytes
@@ -172,7 +165,7 @@ main(int argc, char *argv[]) {
             printf(“The checksum is: %u\n”, Checksum);
             exit(0);
 }
-
+


@@ -191,10 +184,6 @@ main(int argc, char *argv[]) {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/01-06.html b/01-06.html index 597aa6c..614c06f 100644 --- a/01-06.html +++ b/01-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Best Optimizer Is between Your Ears - - @@ -37,7 +30,7 @@


-

LISTING 1.7 L1-7.ASM

+

LISTING 1.7 L1-7.ASM

 ; Assembler subroutine to perform a 16-bit checksum on a block of
 ; bytes 1 to 64K in size. Adds checksum for block into passed-in
@@ -88,7 +81,7 @@ ChecksumLoop:
       ret
 _ChecksumChunkendp
       end
-
+

Note that in Table 1.1, optimization makes little difference except in the case of Listing 1.5, where the design has been refined considerably. Execution time in the other cases is dominated by time spent in DOS and/or the C library, so optimization of the code you write is pretty much irrelevant. What’s more, while the approximately two-times improvement we got by optimizing is not to be sneezed at, it pales against the up-to-50-times improvement we got by redesigning.

@@ -104,13 +97,13 @@ _ChecksumChunkendp

All this is basically a way of saying: Know where you’re going, know the territory, and know when it matters.

-

Where We’ve Been, What We’ve Seen

+

Where We’ve Been, What We’ve Seen

What have we learned? Don’t let other people’s code—even DOS—do the work for you when speed matters, at least not without knowing what that code does and how well it performs.

Optimization only matters after you’ve done your part on the program design end. Consider the ratios on the vertical axis of Table 1.1, which show that optimization is almost totally wasted in the checksumming application without an efficient design. Optimization is no panacea. Table 1.1 shows a two-times improvement from optimization—and a 50-times-plus improvement from redesign. The longstanding debate about which C compiler optimizes code best doesn’t matter quite so much in light of Table 1.1, does it? Your organic optimizer matters much more than your compiler’s optimizer, and there’s always assembly for those usually small sections of code where performance really matters.

-

Where We’re Going

+

Where We’re Going

This chapter has presented a quick step-by-step overview of the design process. I’m not claiming that this is the only way to create high-performance code; it’s just an approach that works for me. Create code however you want, but never forget that design matters more than detailed optimization. Never stop looking for inventive ways to boost performance—and never waste time speeding up code that doesn’t need to be sped up.

@@ -133,10 +126,6 @@ _ChecksumChunkendp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/02-01.html b/02-01.html index ff10ba3..9807c63 100644 --- a/02-01.html +++ b/02-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: A World Apart - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: A World Apart - - @@ -37,10 +30,10 @@


-

Chapter 2
+

Chapter 2
A World Apart

-

The Unique Nature of Assembly Language Optimization

+

The Unique Nature of Assembly Language Optimization

As I showed in the previous chapter, optimization is by no means always a matter of “dropping into assembly.” In fact, in performance tuning high-level language code, assembly should be used rarely, and then only after you’ve made sure a badly chosen or clumsily implemented algorithm isn’t eating you alive. Certainly if you use assembly at all, make absolutely sure you use it right. The potential of assembly code to run slowly is poorly understood by a lot of people, but that potential is great, especially in the hands of the ignorant.

@@ -48,9 +41,9 @@

As usual, the best way to wade in is to present a real-world example.

-

Instructions: The Individual versus the Collective

+

Instructions: The Individual versus the Collective

-

Some time ago, I was asked to work over a critical assembly subroutine in order to make it run as fast as possible. The task of the subroutine was to construct a nibble out of four bits read from different bytes, rotating and combining the bits so that they ultimately ended up neatly aligned in bits 3-0 of a single byte. (In case you’re curious, the object was to construct a 16-color pixel from bits scattered over 4 bytes.) I examined the subroutine line by line, saving a cycle here and a cycle there, until the code truly seemed to be optimized. When I was done, the key part of the code looked something like this:

+

Some time ago, I was asked to work over a critical assembly subroutine in order to make it run as fast as possible. The task of the subroutine was to construct a nibble out of four bits read from different bytes, rotating and combining the bits so that they ultimately ended up neatly aligned in bits 3-0 of a single byte. (In case you’re curious, the object was to construct a 16-color pixel from bits scattered over 4 bytes.) I examined the subroutine line by line, saving a cycle here and a cycle there, until the code truly seemed to be optimized. When I was done, the key part of the code looked something like this:

 LoopTop:
       lodsb            ;get the next byte to extract a bit from
@@ -60,7 +53,7 @@ LoopTop:
       dec   cx         ;the next bit goes 1 place to the right
       dec   dx         ;count down the number of bits
       jnz   LoopTop    ;process the next bit, if any
-
+

Now, it’s hard to write code that’s much faster than seven instructions, only one of which accesses memory, and most programmers would have called it a day at this point. Still, something bothered me, so I spent a bit of time going over the code again. Suddenly, the answer struck me—the code was rotating each bit into place separately, so that a multibit rotation was being performed every time through the loop, for a total of four separate time-consuming multibit rotations!

@@ -72,7 +65,7 @@ LoopTop: -

I changed the code to the following:

+

I changed the code to the following:

 LoopTop:
       lodsb            ;get the next byte to extract a bit from
@@ -83,13 +76,13 @@ LoopTop:
       jnz   LoopTop    ;process the next bit, if any
       rol   bl,cl      ;rotate all four bits into their final
                        ; positions at the same time
-
+

This moved the costly multibit rotation out of the loop so that it was performed just once, rather than four times. While the code may not look much different from the original, and in fact still contains exactly the same number of instructions, the performance of the entire subroutine improved by about 10 percent from just this one change. (Incidentally, that wasn’t the end of the optimization; I eliminated the DEC and JNJ instructions by expanding the four iterations of the loop—but that’s a tale for another chapter.)

The point is this: To write truly superior assembly programs, you need to know what the various instructions do and which instructions execute fastest...and more. You must also learn to look at your programming problems from a variety of perspectives so that you can put those fast instructions to work in the most effective ways.

-

Assembly Is Fundamentally Different

+

Assembly Is Fundamentally Different

Is it really so hard as all that to write good assembly code for the PC? Yes! Thanks to the decidedly quirky nature of the x86 family CPUs, assembly language differs fundamentally from other languages, and is undeniably harder to work with. On the other hand, the potential of assembly code is much greater than that of other languages, as well.

@@ -112,10 +105,6 @@ LoopTop:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/02-02.html b/02-02.html index 17b71cc..0cf5f7d 100644 --- a/02-02.html +++ b/02-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: A World Apart - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: A World Apart - - @@ -37,7 +30,7 @@


-

Transformation Inefficiencies

+

Transformation Inefficiencies

No matter how well an implementation is derived from the corresponding design, however, high-level languages like C/C++ and Pascal inevitably introduce additional transformation inefficiencies, as shown in Figure 2.1.

@@ -45,25 +38,23 @@

High-level languages provide artificial environments that lend themselves relatively well to human programming skills, in order to ease the transition from design to implementation. The price for this ease of implementation is a considerable loss of efficiency in transforming source code into machine language. This is particularly true given that the x86 family in real and 16-bit protected mode, with its specialized memory-addressing instructions and segmented memory architecture, does not lend itself particularly well to compiler design. Even the 32-bit mode of the 386 and its successors, with their more powerful addressing modes, offer fewer registers than compilers would like.

-


- Figure 2.1
  The high-level language transformation inefficiencies.

+


+ Figure 2.1
  The high-level language transformation inefficiencies.

Assembly, on the other hand, is simply a human-oriented representation of machine language. As a result, assembly provides a difficult programming environment—the bare hardware and systems software of the computer—but properly constructed assembly programs suffer no transformation loss, as shown in Figure 2.2.

Only one transformation is required when creating an assembler program, and that single transformation is completely under the programmer’s control. Assemblers perform no transformation from source code to machine language; instead, they merely map assembler instructions to machine language instructions on a one-to-one basis. As a result, the programmer is able to produce machine language code that’s precisely tailored to the needs of each task a given application requires.

-


- Figure 2.2
  Properly constructed assembly programs suffer no transformation loss.

+


+ Figure 2.2
  Properly constructed assembly programs suffer no transformation loss.

The key, of course, is the programmer, since in assembly the programmer must essentially perform the transformation from the application specification to machine language entirely on his or her own. (The assembler merely handles the direct translation from assembly to machine language.)

-

Self-Reliance

+

Self-Reliance

The first part of assembly language optimization, then, is self. An assembler is nothing more than a tool to let you design machine-language programs without having to think in hexadecimal codes. So assembly language programmers—unlike all other programmers—must take full responsibility for the quality of their code. Since assemblers provide little help at any level higher than the generation of machine language, the assembly programmer must be capable both of coding any programming construct directly and of controlling the PC at the lowest practical level—the operating system, the BIOS, even the hardware where necessary. High-level languages handle most of this transparently to the programmer, but in assembly everything is fair—and necessary—game, which brings us to another aspect of assembly optimization: knowledge.

-

Knowledge

+

Knowledge

In the PC world, you can never have enough knowledge, and every item you add to your store will make your programs better. Thorough familiarity with both the operating system APIs and BIOS interfaces is important; since those interfaces are well-documented and reasonably straightforward, my advice is to get a good book or two and bring yourself up to speed. Similarly, familiarity with the PC hardware is required. While that topic covers a lot of ground—display adapters, keyboards, serial ports, printer ports, timer and DMA channels, memory organization, and more—most of the hardware is well-documented, and articles about programming major hardware components appear frequently in the literature, so this sort of knowledge can be acquired readily enough.

@@ -94,10 +85,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/02-03.html b/02-03.html index 4bcedbd..65eeb90 100644 --- a/02-03.html +++ b/02-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: A World Apart - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: A World Apart - - @@ -37,7 +30,7 @@


-

The Flexible Mind

+

The Flexible Mind

Is the never-ending collection of information all there is to the assembly optimization, then? Hardly. Knowledge is simply a necessary base on which to build. Let’s take a moment to examine the objectives of good assembly programming, and the remainder of the forces that act on assembly optimization will fall into place.

@@ -65,7 +58,7 @@

The gist of all this is simply that good assembly programming is done in the context of a solid overall framework unique to each program, and the flexible mind is the key to creating that framework and holding it together.

-

Where to Begin?

+

Where to Begin?

To summarize, the skill of assembly language optimization is a combination of knowledge, perspective, and a way of thought that makes possible the genesis of absolutely the fastest or the smallest code. With that in mind, what should the first step be? Development of the flexible mind is an obvious step. Still, the flexible mind is no better than the knowledge at its disposal. The first step in the journey toward mastering optimization at that exalted level, then, would seem to be learning how to learn.

@@ -86,10 +79,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-01.html b/03-01.html index 5ebd8c3..cd68bce 100644 --- a/03-01.html +++ b/03-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -37,10 +30,10 @@


-

Chapter 3
+

Chapter 3
Assume Nothing

-

Understanding and Using the Zen Timer

+

Understanding and Using the Zen Timer

When you’re pushing the envelope in writing optimized PC code, you’re likely to become more than a little compulsive about finding approaches that let you wring more speed from your computer. In the process, you’re bound to make mistakes, which is fine—as long as you watch for those mistakes and learn from them.

@@ -50,7 +43,7 @@

It ran slower than the original version!

-

The Costs of Ignorance

+

The Costs of Ignorance

As diligent as the author had been, he had nonetheless committed a cardinal sin of x86 assembly language programming: He had assumed that the information available to him was both correct and complete. While the execution times provided by Intel for its processors are indeed correct, they are incomplete; the other—and often more important—part of code performance is instruction fetch time, a topic to which I will return in later chapters.

@@ -70,7 +63,7 @@

Ignorance can also be responsible for considerable wasted effort. I recall a debate in the letters column of one computer magazine about exactly how quickly text can be drawn on a Color/Graphics Adapter (CGA) screen without causing snow. The letter-writers counted every cycle in their timing loops, just as the author in the story that started this chapter had. Like that author, the letter-writers had failed to take the prefetch queue into account. In fact, they had neglected the effects of video wait states as well, so the code they discussed was actually much slower than their estimates. The proper test would, of course, have been to run the code to see if snow resulted, since the only true measure of code performance is observing it in action.

-

The Zen Timer

+

The Zen Timer

Clearly, one key to mastering Zen-class optimization is a tool with which to measure code performance. The most accurate way to measure performance is with expensive hardware, but reasonable measurements at no cost can be made with the PC’s 8253 timer chip, which counts at a rate of slightly over 1,000,000 times per second. The 8253 can be started at the beginning of a block of code of interest and stopped at the end of that code, with the resulting count indicating how long the code took to execute with an accuracy of about 1 microsecond. (A microsecond is one millionth of a second, and is abbreviated µs). To be precise, the 8253 counts once every 838.1 nanoseconds. (A nanosecond is one billionth of a second, and is abbreviated ns.)

@@ -93,10 +86,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-02.html b/03-02.html index af7124f..85baa5b 100644 --- a/03-02.html +++ b/03-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -37,7 +30,7 @@


-

LISTING 3.1 PZTIMER.ASM

+

LISTING 3.1 PZTIMER.ASM

 ; The precision Zen timer (PZTIMER.ASM)
 ;
@@ -478,7 +471,7 @@ ZTimerReport  endp
 
 Code   ends
        end
-
+


@@ -497,10 +490,6 @@ Code ends
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-03.html b/03-03.html index 81db7f1..17e2445 100644 --- a/03-03.html +++ b/03-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -37,19 +30,19 @@


-

The Zen Timer Is a Means, Not an End

+

The Zen Timer Is a Means, Not an End

We’re going to spend the rest of this chapter seeing what the Zen timer can do, examining how it works, and learning how to use it. I’ll be using the Zen timer again and again over the course of this book, so it’s essential that you learn what the Zen timer can do and how to use it. On the other hand, it is by no means essential that you understand exactly how the Zen timer works. (Interesting, yes; essential, no.)

In other words, the Zen timer isn’t really part of the knowledge we seek; rather, it’s one tool with which we’ll acquire that knowledge. Consequently, you shouldn’t worry if you don’t fully grasp the inner workings of the Zen timer. Instead, focus on learning how to use it, and you’ll be on the right road.

-

Starting the Zen Timer

+

Starting the Zen Timer

ZTimerOn is called at the start of a segment of code to be timed. ZTimerOn saves the context of the calling code, disables interrupts, sets timer 0 of the 8253 to mode 2 (divide-by-N mode), sets the initial timer count to 0, restores the context of the calling code, and returns. (I’d like to note that while Intel’s documentation for the 8253 seems to indicate that a timer won’t reset to 0 until it finishes counting down, in actual practice, timers seem to reset to 0 as soon as they’re loaded.)

Two aspects of ZTimerOn are worth discussing further. One point of interest is that ZTimerOn disables interrupts. (ZTimerOff later restores interrupts to the state they were in when ZTimerOn was called.) Were interrupts not disabled by ZTimerOn, keyboard, mouse, timer, and other interrupts could occur during the timing interval, and the time required to service those interrupts would incorrectly and erratically appear to be part of the execution time of the code being measured. As a result, code timed with the Zen timer should not expect any hardware interrupts to occur during the interval between any call to ZTimerOn and the corresponding call to ZTimerOff, and should not enable interrupts during that time.

-

Time and the PC

+

Time and the PC

A second interesting point about ZTimerOn is that it may introduce some small inaccuracy into the system clock time whenever it is called. To understand why this is so, we need to examine the way in which both the 8253 and the PC’s system clock (which keeps the current time) work.

@@ -59,9 +52,8 @@

Timer 1 is dedicated to providing dynamic RAM refresh, and should not be tampered with lest system crashes result.

-


- Figure 3.1
  The configuration of the 8253 timer chip in the PC.

+


+ Figure 3.1
  The configuration of the 8253 timer chip in the PC.

Finally, timer 0 is used to drive the system clock. As programmed by the BIOS at power-up, every 65,536 (64K) counts, or 54.925 milliseconds, timer 0 generates a rising edge on its output line. (A millisecond is one-thousandth of a second, and is abbreviated ms.) This line is connected to the hardware interrupt 0 (IRQ0) line on the system board, so every 54.925 ms, timer 0 causes hardware interrupt 0 to occur.

@@ -92,10 +84,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-04.html b/03-04.html index 33c83fc..feb39cf 100644 --- a/03-04.html +++ b/03-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -49,7 +42,7 @@

Nonetheless, it’s a good idea to reboot your computer at the end of each session with the Zen timer in order to make sure that the system clock is correct.

-

Stopping the Zen Timer

+

Stopping the Zen Timer

At some point after ZTimerOn is called, ZTimerOff must always be called to mark the end of the timing interval. ZTimerOff saves the context of the calling program, latches and reads the timer 0 count, converts that count from the countdown value that the timer maintains to the number of counts elapsed since ZTimerOn was called, and stores the result. Immediately after latching the timer 0 count—and before enabling interrupts—ZTimerOff checks the 8259 interrupt controller to see if there is a pending timer interrupt, setting a flag to mark that the timer overflowed if there is indeed a pending timer interrupt.

@@ -59,7 +52,7 @@

One interesting aspect of ZTimerOff is the manner in which timer 0 is stopped in order to read the timer count. We don’t actually have to stop timer 0 to read the count; the 8253 provides a special latched read feature for the specific purpose of reading the count while a time is running. (That’s a good thing, too; we’ve no documented way to stop timer 0 if we wanted to, since its gate input isn’t connected. Later in this chapter, though, we’ll see that timer 0 can be stopped after all.) We simply tell the 8253 to latch the current count, and the 8253 does so without breaking stride.

-

Reporting Timing Results

+

Reporting Timing Results

ZTimerReport may be called to display timing results at any time after both ZTimerOn and ZTimerOff have been called. ZTimerReport first checks to see whether the timer overflowed (counted down to 0 and turned over) before ZTimerOff was called; if overflow did occur, ZTimerOff prints a message to that effect and returns. Otherwise, ZTimerReport subtracts the reference count (representing the overhead of the Zen timer) from the count measured between the calls to ZTimerOn and ZTimerOff, converts the result from timer counts to microseconds, and prints the resulting time in microseconds to the standard output.

@@ -94,10 +87,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-05.html b/03-05.html index 181d0f5..f6402d8 100644 --- a/03-05.html +++ b/03-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -37,7 +30,7 @@


-

Notes on the Zen Timer

+

Notes on the Zen Timer

The Zen timer subroutines are designed to be near-called from assembly language code running in the public segment Code. The Zen timer subroutines can, however, be called from any assembly or high-level language code that generates OBJ files that are compatible with the Microsoft linker, simply by modifying the segment that the timer code runs in to match the segment used by the code being timed, or by changing the Zen timer routines to far procedures and making far calls to the Zen timer code from the code being timed, as discussed at the end of this chapter. All three subroutines preserve all registers and all flags except the interrupt flag, so calls to these routines are transparent to the calling code.

@@ -51,11 +44,11 @@

Not that this variation between models makes the Zen timer one whit less useful—quite the contrary. The Zen timer is an excellent tool for evaluating code performance over the entire spectrum of PC-compatible computers.

-

A Sample Use of the Zen Timer

+

A Sample Use of the Zen Timer

Listing 3.2 shows a test-bed program for measuring code performance with the Zen timer. This program sets DS equal to CS (for reasons we’ll discuss shortly), includes the code to be measured from the file TESTCODE, and calls ZTimerReport to display the timing results. Consequently, the code being measured should be in the file TESTCODE, and should contain calls to ZTimerOn and ZTimerOff .

-

LISTING 3.2 PZTEST.ASM

+

LISTING 3.2 PZTEST.ASM

 ; Program to measure performance of code that takes less than
 ; 54 ms to execute. (PZTEST.ASM)
@@ -94,11 +87,11 @@ Start proc near
 Start endp
 Code  ends
       end  Start
-
+

Listing 3.3 shows some sample code to be timed. This listing measures the time required to execute 1,000 loads of AL from the memory variable MemVar . Note that Listing 3.3 calls ZTimerOn to start timing, performs 1,000 MOV instructions in a row, and calls ZTimerOff to end timing. When Listing 3.2 is named TESTCODE and included by Listing 3.3, Listing 3.2 calls ZTimerReport to display the execution time after the code in Listing 3.3 has been run.

-

LISTING 3.3 LST3-3.ASM

+

LISTING 3.3 LST3-3.ASM

 ; Test file;
 ; Measures the performance of 1,000 loads of AL from
@@ -124,7 +117,7 @@ Skip:
 ; Stop timing.
 ;
     call  ZTimerOff
-
+

It’s worth noting that Listing 3.3 begins by jumping around the memory variable MemVar. This approach lets us avoid reproducing Listing 3.2 in its entirety for each code fragment we want to measure; by defining any needed data right in the code segment and jumping around that data, each listing becomes self-contained and can be plugged directly into Listing 3.2 as TESTCODE. Listing 3.2 sets DS equal to CS before doing anything else precisely so that data can be embedded in code fragments being timed. Note that only after the initial jump is performed in Listing 3.3 is the Zen timer started, since we don’t want to include the execution time of start-up code in the timing interval. That’s why the calls to ZTimerOn and ZTimerOff are in TESTCODE, not in PZTEST.ASM; this way, we have full control over which portion of TESTCODE is timed, and we can keep set-up code and the like out of the timing interval.

@@ -145,10 +138,6 @@ Skip:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-06.html b/03-06.html index 542eaeb..396ecaf 100644 --- a/03-06.html +++ b/03-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -39,7 +32,7 @@

Listing 3.3 is used by naming it TESTCODE, assembling both Listing 3.2 (which includes TESTCODE) and Listing 3.1 with TASM or MASM, and linking the two resulting OBJ files together by way of the Borland orMicrosoft linker. Listing 3.4 shows a batch file, PZTIME.BAT, which does all that; when run, this batch file generates and runs the executable file PZTEST.EXE. PZTIME.BAT (Listing 3.4) assumes that the file PZTIMER.ASM contains Listing 3.1, and the file PZTEST.ASM contains Listing 3.2. The command-line parameter to PZTIME.BAT is the name of the file to be copied to TESTCODE and included into PZTEST.ASM. (Note that Turbo Assembler can be substituted for MASM by replacing “masm” with “tasm” and “link” with “tlink” in Listing 3.4. The same is true of Listing 3.7.)

-

LISTING 3.4 PZTIME.BAT

+

LISTING 3.4 PZTIME.BAT

 echo off
 rem
@@ -102,25 +95,25 @@ echo ***************************************************************
 echo * An error occurred while building the precision Zen timer.   *
 echo ***************************************************************
 :end
-
+ -

Assuming that Listing 3.3 is named LST3-3.ASM and Listing 3.4 is named PZTIME.BAT, the code in Listing 3.3 would be timed with the command:

+

Assuming that Listing 3.3 is named LST3-3.ASM and Listing 3.4 is named PZTIME.BAT, the code in Listing 3.3 would be timed with the command:

 pztime LST3-3.ASM
-
+

which performs all assembly and linking, and reports the execution time of the code in Listing 3.3.

When the above command is executed on an original 4.77 MHz IBM PC, the time reported by the Zen timer is 3619 µs, or about 3.62 µs per load of AL from memory. (While the exact number is 3.619 µs per load of AL, I’m going to round off that last digit from now on. No matter how many repetitions of a given instruction are timed, there’s just too much noise in the timing process—between dynamic RAM refresh, the prefetch queue, and the internal state of the processor at the start of timing—for that last digit to have any significance.) Given the test PC’s 4.77 MHz clock, this works out to about 17 cycles per MOV, which is actually a good bit longer than Intel’s specified 10-cycle execution time for this instruction. (See the MASM or TASM documentation, or Intel’s processor reference manuals, for official execution times.) Fear not, the Zen timer is right—MOV AL,[MEMVAR] really does take 17 cycles as used in Listing 3.3. Exactly why that is so is just what this book is all about.

-

In order to perform any of the timing tests in this book, enter Listing 3.1 and name it PZTIMER.ASM, enter Listing 3.2 and name it PZTEST.ASM, and enter Listing 3.4 and name it PZTIME.BAT. Then simply enter the listing you wish to run into the file filename and enter the command:

+

In order to perform any of the timing tests in this book, enter Listing 3.1 and name it PZTIMER.ASM, enter Listing 3.2 and name it PZTEST.ASM, and enter Listing 3.4 and name it PZTIME.BAT. Then simply enter the listing you wish to run into the file filename and enter the command:

 pztime <filename>
-
+

In fact, that’s exactly how I timed each of the listings in this book. Code fragments you write yourself can be timed in just the same way. If you wish to time code directly in place in your programs, rather than in the test-bed program of Listing 3.2, simply insert calls to ZTimerOn, ZTimerOff, and ZTimerReport in the appropriate places and link PZTIMER to your program.

-

The Long-Period Zen Timer

+

The Long-Period Zen Timer

With a few exceptions, the Zen timer presented above will serve us well for the remainder of this book since we’ll be focusing on relatively short code sequences that generally take much less than 54 ms to execute. Occasionally, however, we will need to time longer intervals. What’s more, it is very likely that you will want to time code sequences longer than 54 ms at some point in your programming career. Accordingly, I’ve also developed a Zen timer for periods longer than 54 ms. The long-period Zen timer (so named by contrast with the precision Zen timer just presented) shown in Listing 3.5 can measure periods up to one hour in length.

@@ -147,10 +140,6 @@ pztime <filename>
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-07.html b/03-07.html index 3ca4750..e4cbaf3 100644 --- a/03-07.html +++ b/03-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -41,7 +34,7 @@

The long-period Zen timer has some of the same effects on the system time as does the precision Zen timer, so it’s a good idea to reboot the system after a session with the long-period Zen timer. The long-period Zen timer does not, however, have the same potential for introducing major inaccuracy into the system clock time during a single timing run since it leaves interrupts enabled and therefore allows the system clock to update normally.

-

Stopping the Clock

+

Stopping the Clock

There’s a potential problem with the long-period Zen timer. The problem is this: In order to measure times longer than 54 ms, we must maintain not one but two timing components, the timer 0 count and the BIOS time-of-day count. The time-of-day count measures the passage of 54.9 ms intervals, while the timer 0 count measures time within those 54.9 ms intervals. We need to read the two time components simultaneously in order to get a clean reading. Otherwise, we may read the timer count just before it turns over and generates an interrupt, then read the BIOS time-of-day count just after the interrupt has occurred and caused the time-of-day count to turn over, with a resulting 54 ms measurement inaccuracy. (The opposite sequence—reading the time-of-day count and then the timer count—can result in a 54 ms inaccuracy in the other direction.)

@@ -53,7 +46,7 @@

I’ve set up Listing 3.5 so that it can assemble to either use or not use the undocumented timer-stopping feature, as you please. The PS2 equate selects between the two modes of operation. If PS2 is 1 (as it is in Listing 3.5), then the latch-and-read method is used; if PS2 is 0, then the undocumented timer-stop approach is used. The latch-and-read method will work on all PC-compatible computers, but may occasionally produce results that are incorrect by 54 ms. The timer-stop approach avoids synchronization problems, but doesn’t work on all computers.

-

LISTING 3.5 LZTIMER.ASM

+

LISTING 3.5 LZTIMER.ASM

 ;
 ; The long-period Zen timer. (LZTIMER.ASM)
@@ -689,7 +682,7 @@ ZTimerReport    endp
 
 Code   ends
        end
-
+


@@ -708,10 +701,6 @@ Code ends
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-08.html b/03-08.html index 6245f36..9f42ba1 100644 --- a/03-08.html +++ b/03-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -47,13 +40,13 @@

Finally, please note that the precision Zen timer works perfectly well on both PS/2 and non-PS/2 computers. The PS/2 and 8253 considerations we’ve just discussed apply only to the longZen timer.

-

Example Use of the Long-Period Zen Timer

+

Example Use of the Long-Period Zen Timer

The long-period Zen timer has exactly the same calling interface as the precision Zen timer, and can be used in place of the precision Zen timer simply by linking it to the code to be timed in place of linking the precision timer code. Whenever the precision Zen timer informs you that the code being timed takes too long for the precision timer to handle, all you have to do is link in the long-period timer instead.

Listing 3.6 shows a test-bed program for the long-period Zen timer. While this program is similar to Listing 3.2, it’s worth noting that Listing 3.6 waits for a few seconds before calling ZTimerOn, thereby allowing any pending keyboard interrupts to be processed. Since interrupts must be left on in order to time periods longer than 54 ms, the interrupts generated by keystrokes (including the upstroke of the Enter key press that starts the program)—or any other interrupts, for that matter—could incorrectly inflate the time recorded by the long-period Zen timer. In light of this, resist the temptation to type ahead, move the mouse, or the like while the long-period Zen timer is timing.

-

LISTING 3.6 LZTEST.ASM

+

LISTING 3.6 LZTEST.ASM

 ; Program to measure performance of code that takes longer than
 ; 54 ms to execute. (LZTEST.ASM)
@@ -113,7 +106,7 @@ CheckDelayTime:
 Start endp
 Code  ends
       end     Start
-
+

As with the precision Zen timer, the program in Listing 3.6 is used by naming the file containing the code to be timed TESTCODE, then assembling both Listing 3.6 and Listing 3.5 with MASM or TASM and linking the two files together by way of the Microsoft or Borland linker. Listing 3.7 shows a batch file, named LZTIME.BAT, which does all of the above, generating and running the executable file LZTEST.EXE. LZTIME.BAT assumes that the file LZTIMER.ASM contains Listing 3.5 and the file LZTEST.ASM contains Listing 3.6.

@@ -134,10 +127,6 @@ Code ends
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-09.html b/03-09.html index 2d606a0..5b1626d 100644 --- a/03-09.html +++ b/03-09.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -37,7 +30,7 @@


-

LISTING 3.7 LZTIME.BAT

+

LISTING 3.7 LZTIME.BAT

 echo off
 rem
@@ -100,11 +93,11 @@ echo ***************************************************************
 echo * An error occurred while building the long-period Zen timer. *
 echo ***************************************************************
 :end
-
+

Listing 3.8 shows sample code that can be timed with the test-bed program of Listing 3.6. Listing 3.8 measures the time required to execute 20,000 loads of AL from memory, a length of time too long for the precision Zen timer to handle on the 8088.

-

LISTING 3.8 LST3-8.ASM

+

LISTING 3.8 LST3-8.ASM

 ;
 ; Measures the performance of 20,000 loads of AL from
@@ -133,25 +126,25 @@ endm
 ; Stop timing.
 ;
 callZTimerOff
-
+ -

When LZTIME.BAT is run on a PC with the following command line (assuming the code in Listing 3.8 is the file LST3-8.ASM)

+

When LZTIME.BAT is run on a PC with the following command line (assuming the code in Listing 3.8 is the file LST3-8.ASM)

 lztime lst3-8.asm
-
+

the result is 72,544 µs, or about 3.63 µs per load of AL from memory. This is just slightly longer than the time per load of AL measured by the precision Zen timer, as we would expect given that interrupts are left enabled by the long-period Zen timer. The extra fraction of a microsecond measured per MOV reflects the time required to execute the BIOS code that handles the 18.2 timer interrupts that occur each second.

Note that the command can take as much as 10 minutes to finish on a slow PC if you are using MASM, with most of that time spent assembling Listing 3.8. Why? Because MASM is notoriously slow at assembling REPT blocks, and the block in Listing 3.8 is repeated 20,000 times.

-

Using the Zen Timer from C

+

Using the Zen Timer from C

The Zen timer can be used to measure code performance when programming in C—but not right out of the box. As presented earlier, the timer is designed to be called from assembly language; some relatively minor modifications are required before the ZTimerOn (start timer), ZTimerOff (stop timer), and ZTimerReport (display timing results) routines can be called from C. There are two separate cases to be dealt with here: small code model and large; I’ll tackle the simpler one, the small code model, first.

-

Altering the Zen timer for linking to a small code model C program involves the following steps: C hange ZTimerOn to _ZTimerOn, change ZTimerOff to _ZTimerOff, change ZTimerReport to _ZTimerReport, and change Code to _TEXT . Figure 3.2 shows the line numbers and new states of all lines from Listing 3.1 that must be changed. These changes convert the code to use C-style external label names and the small model C code segment. (In C++, use the “C” specifier, as in

+

Altering the Zen timer for linking to a small code model C program involves the following steps: C hange ZTimerOn to _ZTimerOn, change ZTimerOff to _ZTimerOff, change ZTimerReport to _ZTimerReport, and change Code to _TEXT . Figure 3.2 shows the line numbers and new states of all lines from Listing 3.1 that must be changed. These changes convert the code to use C-style external label names and the small model C code segment. (In C++, use the “C” specifier, as in

 extern “C” ZTimerOn(void);
-
+


@@ -170,10 +163,6 @@ extern “C” ZTimerOn(void);
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/03-10.html b/03-10.html index d850de0..854f663 100644 --- a/03-10.html +++ b/03-10.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Assume Nothing - - @@ -39,43 +32,41 @@

when declaring the timer routines extern, so that name-mangling doesn’t occur, and the linker can find the routines’ C-style names.)

-

That’s all it takes; after doing this, you’ll be able to use the Zen timer from C, as, for example, in:

+

That’s all it takes; after doing this, you’ll be able to use the Zen timer from C, as, for example, in:

 ZTimerOn():
 for (i=0, x=0; i<100; i++)
      x += i;
 ZTimerOff();
 ZTimerReport();
-
+

(I’m talking about the precision timer here. The long-period timer—Listing 3.5—requires the same modifications, but to different lines.)

-


- Figure 3.2
  Changes for use with small code model C.

+


+ Figure 3.2
  Changes for use with small code model C.

Altering the Zen timer for use in C’s large code model is a tad more complex, because in addition to the above changes, all functions, including the internal reference timing routines that are used to calculate overhead so it can be subtracted out, must be converted to far. Figure 3.3 shows the line numbers and new states of all lines from Listing 3.1 that must be changed in order to call the Zen timer from large code model C. Again, the line numbers are specific to the precision timer, but the long-period timer is very similar.

The full listings for the C-callable Zen timers are presented in Chapter K on the companion CD-ROM.

-

Watch Out for Optimizing Assemblers!

+

Watch Out for Optimizing Assemblers!

-

One important safety tip when modifying the Zen timer for use with large code model C code: Watch out for optimizing assemblers! TASM actually replaces

+

One important safety tip when modifying the Zen timer for use with large code model C code: Watch out for optimizing assemblers! TASM actually replaces

 call     far ptr ReferenceZTimerOn
-
+ -

with

+

with

 push     cs
 call     near ptr ReferenceZTimerOn
-
+

(and likewise for ReferenceZTimerOff ), which works because ReferenceZTimerOn is in the same segment as the calling code. This is normally a great optimization, being both smaller and faster than a far call. However, it’s not so great for the Zen

-


- Figure 3.3
  Changes for use with large code model C.

+


+ Figure 3.3
  Changes for use with large code model C.

timer, because our purpose in calling the reference timing code is to determine exactly how much time is taken by overhead code—including the far calls to ZTimerOn and ZTimerOff! By converting the far calls to push/near call pairs within the Zen timer module, TASM makes it impossible to emulate exactly the overhead of the Zen timer, and makes timings slightly (about 16 cycles on a 386) less accurate.

@@ -85,13 +76,13 @@ call near ptr ReferenceZTimerOn

I’ve tested the changes shown in Figures 3.2 and 3.3 with TASM and Borland C++ 4.0, and also with the latest MASM and Microsoft C/C++ compiler.

-

Further Reading

+

Further Reading

For those of you who wish to pursue the mechanics of code measurement further, one good article about measuring code performance with the 8253 timer is “Programming Insight: High-Performance Software Analysis on the IBM PC,” by Byron Sheppard, which appeared in the January, 1987 issue of Byte. For complete if somewhat cryptic information on the 8253 timer itself, I refer you to Intel’s Microsystem Components Handbook, which is also a useful reference for a number of other PC components, including the 8259 Programmable Interrupt Controller and the 8237 DMA Controller. For details about the way the 8253 is used in the PC, as well as a great deal of additional information about the PC’s hardware and BIOS resources, I suggest you consult IBM’s series of technical reference manuals for the PC, XT, AT, Model 30, and microchannel computers, such as the Models 50, 60, and 80.

For our purposes, however, it’s not critical that you understand exactly how the Zen timer works. All you really need to know is what the Zen timer can do and how to use it, and we’ve accomplished that in this chapter.

-

Armed with the Zen Timer, Onward and Upward

+

Armed with the Zen Timer, Onward and Upward

The Zen timer is not perfect. For one thing, the finest resolution to which it can measure an interval is at best about 1µs, a period of time in which a 66 MHz Pentium computer can execute as many as 132 instructions (although an 8088-based PC would be hard-pressed to manage two instructions in a microsecond). Another problem is that the timing code itself interferes with the state of the prefetch queue and processor cache at the start of the code being timed, because the timing code is not necessarily fetched and does not necessarily access memory in exactly the same time sequence as the code immediately preceding the code under measurement normally does. This prefetch effect can introduce as much as 3 to 4 µ of inaccuracy. Similarly, the state of the prefetch queue at the end of the code being timed affects how long the code that stops the timer takes to execute. Consequently, the Zen timer tends to be more accurate for longer code sequences, since the relative magnitude of the inaccuracy introduced by the Zen timer becomes less over longer periods.

@@ -114,10 +105,6 @@ call near ptr ReferenceZTimerOn
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-01.html b/04-01.html index fcdb998..23a9aa2 100644 --- a/04-01.html +++ b/04-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -37,10 +30,10 @@


-

Chapter 4
+

Chapter 4
In the Lair of the Cycle-Eaters

-

How the PC Hardware Devours Code Performance

+

How the PC Hardware Devours Code Performance

This chapter, adapted from my earlier book, Zen of Assembly Language located on the companion CD-ROM, goes right to the heart of my philosophy of optimization: Understand where the time really goes when your code runs. That may sound ridiculously simple, but, as this chapter makes clear, it turns out to be a challenging task indeed, one that at times verges on black magic. This chapter is a long-time favorite of mine because it was the first—and to a large extent only—work that I know of that discussed this material, thereby introducing a generation of PC programmers to pedal-to-the-metal optimization.

@@ -48,7 +41,7 @@

So, don’t take either the absolute or the relative execution times presented in this chapter as gospel for newer processors, and read on to later chapters to see how the cycle-eaters and optimization rules have changed over time, but do take the time to at least skim through this chapter to give yourself a good start on the material in the rest of this book.

-

Cycle-Eaters

+

Cycle-Eaters

Programming has many levels, ranging from the familiar (high-level languages, DOS calls, and the like) down to the esoteric things that lie on the shadowy edge of hardware-land. I call these cycle-eaters because, like the monsters in a bad 50s horror movie, they lurk in those shadows, taking their share of your program’s performance without regard to the forces of goodness or the U.S. Army. In this chapter, we’re going to jump right in at the lowest level by examining the cycle-eaters that live beneath the programming interface; that is, beneath your application, DOS, and BIOS—in fact, beneath the instruction set itself.

@@ -58,13 +51,13 @@

Which brings us to cycle-eaters.

-

The Nature of Cycle-Eaters

+

The Nature of Cycle-Eaters

Cycle-eaters are gremlins that live on the bus or in peripherals (and sometimes within the CPU itself), slowing the performance of PC code so that it doesn’t execute at full speed. Most cycle-eaters (and all of those haunting the older Intel processors) live outside the CPU’s Execution Unit, where they can only affect the CPU when the CPU performs a bus access (a memory or I/O read or write). Once your code and data are already inside the CPU, those cycle-eaters can no longer be a problem. Only on the 486 and Pentium CPUs will you find cycle-eaters inside the chip, as we’ll see in later chapters.

The nature and severity of the cycle-eaters vary enormously from processor to processor, and (especially) from memory architecture to memory architecture. In order to understand them all, we need first to understand the simplest among them, those that haunted the original 8088-based IBM PC. Later on in this book, I’ll be better able to explain the newer generation of cycle-eaters in terms of those ancestral cycle-eaters—but we have to get the groundwork down first.

-

The 8088’s Ancestral Cycle-Eaters

+

The 8088’s Ancestral Cycle-Eaters

Internally, the 8088 is a 16-bit processor, capable of running at full speed at all times—unless external data is required. External data must traverse the 8088’s external data bus and the PC’s data bus one byte at a time to and from peripherals, with cycle-eaters lurking along every step of the way. What’s more, external data includes not only memory operands but also instruction bytes, so even instructions with no memory operands can suffer from cycle-eaters. Since some of the 8088’s fastest instructions are register-only instructions, that’s important indeed.

@@ -82,7 +75,7 @@

The locations of these cycle-eaters in the primordial 8088-based PC are shown in Figure 4.1. We’ll cover each of the cycle-eaters in turn in this chapter. The material won’t be easy since cycle-eaters are among the most subtle aspects of assembly programming. By the same token, however, this will be one of the most important and rewarding chapters in this book. Don’t worry if you don’t catch everything in this chapter, but do read it all even if the going gets a bit tough. Cycle-eaters play a key role in later chapters, so some familiarity with them is highly desirable.

-

The 8-Bit Bus Cycle-Eater

+

The 8-Bit Bus Cycle-Eater

Look! Down on the motherboard! It’s a 16-bit processor! It’s an 8-bit processor! It’s...

@@ -109,10 +102,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-02.html b/04-02.html index 8ae0cc7..3bc1f86 100644 --- a/04-02.html +++ b/04-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -37,13 +30,11 @@


-


- Figure 4.1
  The location of the major cycle-eaters in the IBM PC.

+


+ Figure 4.1
  The location of the major cycle-eaters in the IBM PC.

-


- Figure 4.2
  Internal data bus widths of the 8088.

+


+ Figure 4.2
  Internal data bus widths of the 8088.

As shown in Figure 4.1, the 8-bit bus cycle-eater lies squarely on the 8088’s external data bus. Technically, it might be more accurate to place this cycle-eater in the Bus Interface Unit, which breaks 16-bit memory accesses into paired 8-bit accesses, but it is really the limited width of the external data bus that constricts data flow into and out of the 8088. True, the original PC’s bus is also only 8 bits wide, but that’s just to match the 8088’s 8-bit bus; even if the PC’s bus were 16 bits wide, data could still pass into and out of the 8088 chip itself only 1 byte at a time.

@@ -51,39 +42,39 @@

A related cycle-eater lurks beneath the 386SX chip, which is a 32-bit processor internally with only a 16-bit path to system memory. The numbers are different, but the way the cycle-eater operates is exactly the same. AT-compatible systems have 16-bit data buses, which can access a full 16-bit word at a time. The 386SX can process 32 bits (a doubleword) at a time, however, and loses a lot of time fetching that doubleword from memory in two halves.

-

The Impact of the 8-Bit Bus Cycle-Eater

+

The Impact of the 8-Bit Bus Cycle-Eater

-

One obvious effect of the 8-bit bus cycle-eater is that word-sized accesses to memory operands on the 8088 take 4 cycles longer than byte-sized accesses. That’s why the official instruction timings indicate that for code running on an 8088 an additional 4 cycles are required for every word-sized access to a memory operand. For instance,

+

One obvious effect of the 8-bit bus cycle-eater is that word-sized accesses to memory operands on the 8088 take 4 cycles longer than byte-sized accesses. That’s why the official instruction timings indicate that for code running on an 8088 an additional 4 cycles are required for every word-sized access to a memory operand. For instance,

 mov  ax,word ptr [MemVar]
-
+ -

takes 4 cycles longer to read the word at address MemVar than

+

takes 4 cycles longer to read the word at address MemVar than

 mov  al,byte ptr [MemVar]
-
+

takes to read the byte at address MemVar. (Actually, the difference between the two isn’t very likely to be exactly 4 cycles, for reasons that will become clear once we discuss the prefetch queue and dynamic RAM refresh cycle-eaters later in this chapter.)

-

What’s more, in some cases one instruction can perform multiple word-sized accesses, incurring that 4-cycle penalty on each access. For example, adding a value to a word-sized memory variable requires two word-sized accesses—one to read the destination operand from memory prior to adding to it, and one to write the result of the addition back to the destination operand—and thus incurs not one but two 4-cycle penalties. As a result

+

What’s more, in some cases one instruction can perform multiple word-sized accesses, incurring that 4-cycle penalty on each access. For example, adding a value to a word-sized memory variable requires two word-sized accesses—one to read the destination operand from memory prior to adding to it, and one to write the result of the addition back to the destination operand—and thus incurs not one but two 4-cycle penalties. As a result

 add  word ptr [MemVar],ax
-
+ -

takes about 8 cycles longer to execute than:

+

takes about 8 cycles longer to execute than:

 add  byte ptr [MemVar],al
-
+

String instructions can suffer from the 8-bit bus cycle-eater to a greater extent than other instructions. Believe it or not, a single REP MOVSW instruction can lose as much as 131,070 word-sized memory accesses x 4 cycles, or 524,280 cycles to the 8-bit bus cycle-eater! In other words, one 8088 instruction (admittedly, an instruction that does a great deal) can take over one-tenth of a second longer on an 8088 than on an 8086, simply because of the 8-bit bus. One-tenth of a second! That’s a phenomenally long time in computer terms; in one-tenth of a second, the 8088 can perform more than 50,000 additions and subtractions.

The upshot of all this is simply that the 8088 can transfer word-sized data to and from memory at only half the speed of the 8086, which inevitably causes performance problems when coupled with an Execution Unit that can process word-sized data every bit as quickly as an 8086. These problems show up with any code that uses word-sized memory operands. More ominously, as we will see shortly, the 8-bit bus cycle-eater can cause performance problems with other sorts of code as well.

-

What to Do about the 8-Bit Bus Cycle-Eater?

+

What to Do about the 8-Bit Bus Cycle-Eater?

The obvious implication of the 8-bit bus cycle-eater is that byte-sized memory variables should be used whenever possible. After all, the 8088 performs byte-sized memory accesses just as quickly as the 8086. For instance, Listing 4.1, which uses a byte-sized memory variable as a loop counter, runs in 10.03 s per loop. That’s 20 percent faster than the 12.05 µs per loop execution time of Listing 4.2, which uses a word-sized counter. Why the difference in execution times? Simply because each word-sized DEC performs 4 byte-sized memory accesses (two to read the word-sized operand and two to write the result back to memory), while each byte-sized DEC performs only 2 byte-sized memory accesses in all.

-

LISTING 4.1 LST4-1.ASM

+

LISTING 4.1 LST4-1.ASM

 ; Measures the performance of a loop which uses a
 ; byte-sized memory variable as the loop counter.
@@ -98,9 +89,9 @@ LoopTop:
       dec  [Counter]
       jnz  LoopTop
       call ZTimerOff
-
+ -

LISTING 4.2 LST4-2.ASM

+

LISTING 4.2 LST4-2.ASM

 ; Measures the performance of a loop which uses a
 ; word-sized memory variable as the loop counter.
@@ -115,7 +106,7 @@ LoopTop:
       dec   [Counter]
       jnz   LoopTop
       call  ZTimerOff
-
+

I’d like to make a brief aside concerning code optimization in the listings in this book. Throughout this book I’ve modeled the sample code after working code so that the timing results are applicable to real-world programming. In Listings 4.1 and 4.2, for example, I could have shown a still greater advantage for byte-sized operands simply by performing 1,000 DEC instructions in a row, with no branching at all. However, DEC instructions don’t exist in a vacuum, so in the listings I used code that both decremented the counter and tested the result. The difference is that between decrementing a memory location (simply an instruction) and using a loop counter (a functional instruction sequence). If you come across code in this book that seems less than optimal, it’s simply due to my desire to provide code that’s relevant to real programming problems. On the other hand, optimal code is an elusive thing indeed; by no means should you assume that the code in this book is ideal! Examine it, question it, and improve upon it, for an inquisitive, skeptical mind is an important part of the Zen of assembly optimization.

@@ -136,10 +127,6 @@ LoopTop:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-03.html b/04-03.html index fce71be..c79e9b6 100644 --- a/04-03.html +++ b/04-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -37,22 +30,22 @@


-

Back to the 8-bit bus cycle-eater. As I’ve said, in 8088 work you should strive to use byte-sized memory variables whenever possible. That does not mean that you should use 2 byte-sized memory accesses to manipulate a word-sized memory variable in preference to 1 word-sized memory access, as, for instance,

+

Back to the 8-bit bus cycle-eater. As I’ve said, in 8088 work you should strive to use byte-sized memory variables whenever possible. That does not mean that you should use 2 byte-sized memory accesses to manipulate a word-sized memory variable in preference to 1 word-sized memory access, as, for instance,

 mov  dl,byte ptr [MemVar]
 mov  dh,byte ptr [MemVar+1]
-
+ -

versus:

+

versus:

 mov  dx,word ptr [MemVar]
-
+

Recall that every access to a memory byte takes at least 4 cycles; that limitation is built right into the 8088. The 8088 is also built so that the second byte-sized memory access to a 16-bit memory variable takes just those 4 cycles and no more. There’s no way you can manipulate the second byte of a word-sized memory variable faster with a second separate byte-sized instruction in less than 4 cycles. As a matter of fact, you’re bound to access that second byte much more slowly with a separate instruction, thanks to the overhead of instruction fetching and execution, address calculation, and the like.

For example, consider Listing 4.3, which performs 1,000 word-sized reads from memory. This code runs in 3.77 µs per word read on a 4.77 MHz 8088. That’s 45 percent faster than the 5.49 µs per word read of Listing 4.4, which reads the same 1,000 words as Listing 4.3 but does so with 2,000 byte-sized reads. Both listings perform exactly the same number of memory accesses—2,000 accesses, each byte-sized, as all 8088 memory accesses must be. (Remember that the Bus Interface Unit must perform two byte-sized memory accesses in order to handle a word-sized memory operand.) However, Listing 4.3 is considerably faster because it expends only 4 additional cycles to read the second byte of each word, while Listing 4.4 performs a second LODSB, requiring 13 cycles, to read the second byte of each word.

-

LISTING 4.3 LST4-3.ASM

+

LISTING 4.3 LST4-3.ASM

 ; Measures the performance of reading 1,000 words
 ; from memory with 1,000 word-sized accesses.
@@ -62,9 +55,9 @@ mov  dx,word ptr [MemVar]
      call ZTimerOn
      rep  lodsw
      call ZTimerOff
-
+ -

LISTING 4.4 LST4-4.ASM

+

LISTING 4.4 LST4-4.ASM

 ; Measures the performance of reading 1000 words
 ; from memory with 2,000 byte-sized accesses.
@@ -74,7 +67,7 @@ mov  dx,word ptr [MemVar]
      call ZTimerOn
      rep  lodsb
      call ZTimerOff
-
+

In short, if you must perform a 16-bit memory access, let the 8088 break the access into two byte-sized accesses for you. The 8088 is more efficient at that task than your code can possibly be.

@@ -86,7 +79,7 @@ mov dx,word ptr [MemVar]

Yes and no. It’s true that in general we know approximately how much longer a given instruction will take to execute with a word-sized memory operand than with a byte-sized operand, although the dynamic RAM refresh and wait state cycle-eaters (which I’ll cover a little later) can raise the cost of the 8-bit bus cycle-eater considerably. However, all word-sized memory accesses lose 4 cycles to the 8-bit bus cycle-eater, and there’s one sort of word-sized memory access we haven’t discussed yet: instruction fetching. The ugliest manifestation of the 8-bit bus cycle-eater is in fact the prefetch queue cycle-eater.

-

The Prefetch Queue Cycle-Eater

+

The Prefetch Queue Cycle-Eater

In an 8088 context, here’s the prefetch queue cycle-eater in a nutshell: The 8088’s 8-bit external data bus keeps the Bus Interface Unit from fetching instruction bytes as fast as the 16-bit Execution Unit can execute them, so the Execution Unit often lies idle while waiting for the next instruction byte to be fetched.

@@ -96,19 +89,19 @@ mov dx,word ptr [MemVar]

Clearly, then, the prefetch queue cycle-eater is nothing more than one aspect of the 8-bit bus cycle-eater. 8088 code often runs at less than the Execution Unit’s maximum speed because the 8-bit data bus can’t keep up with the demand for instruction bytes. That’s straightforward enough—so why all the fuss about the prefetch queue cycle-eater?

-

What makes the prefetch queue cycle-eater tricky is that it’s undocumented and unpredictable. That is, with a word-sized memory access, such as

+

What makes the prefetch queue cycle-eater tricky is that it’s undocumented and unpredictable. That is, with a word-sized memory access, such as

 mov  [bx],ax
-
+ -

it’s well-documented that an extra 4 cycles will always be required to write the upper byte of AX to memory. Not so with the prefetch queue cycle-eater lurking nearby. For instance, the instructions

+

it’s well-documented that an extra 4 cycles will always be required to write the upper byte of AX to memory. Not so with the prefetch queue cycle-eater lurking nearby. For instance, the instructions

 shr  ax,1
 shr  ax,1
 shr  ax,1
 shr  ax,1
 shr  ax,1
-
+

should execute in 10 cycles, since each SHR takes 2 cycles to execute, according to Intel’s specifications. Those specifications contain Intel’s official instruction execution times, but in this case—and in many others—the specifications are drastically wrong. Why? Because they describe execution time once an instruction reaches the prefetch queue. They say nothing about whether a given instruction will be in the prefetch queue when it’s time for that instruction to run, or how long it will take that instruction to reach the prefetch queue if it’s not there already. Thanks to the low performance of the 8088’s external data bus, that’s a glaring omission—but, alas, an unavoidable one. Let’s look at why the official execution times are wrong, and why that can’t be helped.

@@ -129,10 +122,6 @@ shr ax,1
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-04.html b/04-04.html index f1d875c..2435cf4 100644 --- a/04-04.html +++ b/04-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -37,7 +30,7 @@


-

Official Execution Times Are Only Part of the Story

+

Official Execution Times Are Only Part of the Story

The sequence of 5 SHR instructions in the last example is 10 bytes long. That means that it can never execute in less than 24 cycles even if the 4-byte prefetch queue is full when it starts, since 6 instruction bytes would still remain to be fetched, at 4 cycles per fetch. If the prefetch queue is empty at the start, the sequence could take 40 cycles. In short, thanks to instruction fetching, the code won’t run at its documented speed, and could take up to four times longer than it is supposed to.

@@ -55,7 +48,7 @@

So now you know why the official instruction execution times are often wrong, and why Intel can’t provide better specifications. You also know now why it is that you must time your code if you want to know how fast it really is.

-

There Is No Such Beast as a True Instruction Execution Time

+

There Is No Such Beast as a True Instruction Execution Time

The effect of the code preceding an instruction on the execution time of that instruction makes the Zen timer trickier to use than you might expect, and complicates the interpretation of the results reported by the Zen timer. For one thing, the Zen timer is best used to time code sequences that are more than a few instructions long; below 10µs or so, prefetch queue effects and the limited resolution of the clock driving the timer can cause problems.

@@ -65,7 +58,7 @@

For example, consider the code in Listings 4.5 and 4.6. Listing 4.5 shows our familiar SHR case. Here, because the prefetch queue is always empty, execution time should work out to about 4 cycles per byte, or 8 cycles per SHR, as shown in Figure 4.3. (Figure 4.3 illustrates the relationship between instruction fetching and execution in a simplified way, and is not intended to show the exact timings of 8088 operations.) That’s quite a contrast to the official 2-cycle execution time of SHR. In fact, the Zen timer reports that Listing 4.5 executes in 1.81µs per byte, or slightly more than 4 cycles per byte. (The extra time is the result of the dynamic RAM refresh cycle-eater, which we’ll discuss shortly.) Going by Listing 4.5, we would conclude that the “true” execution time of SHR is 8.64 cycles.

-

LISTING 4.5 LST4-5.ASM

+

LISTING 4.5 LST4-5.ASM

 ; Measures the performance of 1,000 SHR instructions
 ; in a row. Since SHR executes in 2 cycles but is
@@ -78,9 +71,9 @@
       shr   ax,1
       endm
       call  ZTimerOff
-
+ -

LISTING 4.6 LST4-6.ASM

+

LISTING 4.6 LST4-6.ASM

 ; Measures the performance of 1,000 MUL/SHR instruction
 ; pairs in a row. The lengthy execution time of MUL
@@ -94,11 +87,10 @@
       shr   ax,1
       endm
       call  ZTimerOff
-
+ -


- Figure 4.3
  Execution and instruction prefetching sequence for Listing 4.5.

+


+ Figure 4.3
  Execution and instruction prefetching sequence for Listing 4.5.

Now let’s examine Listing 4.6. Here each SHR follows a MUL instruction. Since MUL instructions take so long to execute that the prefetch queue is always full when they finish, each SHR should be ready and waiting in the prefetch queue when the preceding MUL ends. As a result, we’d expect that each SHR would execute in 2 cycles; together with the 118-cycle execution time of multiplying 0 times 0, the total execution time should come to 120 cycles per SHR/MUL pair, as shown in Figure 4.4. And, by God, when we run Listing 4.6 we get an execution time of 25.14 µs per SHR/MUL pair, or exactly 120 cycles! According to these results, the “true” execution time of SHR would seem to be 2 cycles, quite a change from the conclusion we drew from Listing 4.5.

@@ -121,10 +113,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-05.html b/04-05.html index 9cc0267..9c2b7a4 100644 --- a/04-05.html +++ b/04-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -41,11 +34,10 @@

The truth is that it never hurts performance to reduce either the cycle count or the byte count of a given bit of code, but there’s no guarantee that one or the other will improve performance either. For example, consider Listing 4.7, which consists of a series of 4-cycle, 2-byte MOV AL,0 instructions, and which executes at the rate of 1.81 µs per instruction. Now consider Listing 4.8, which replaces the 4-cycle MOV AL,0 with the 3-cycle (but still 2-byte) SUB AL,AL, Despite its 1-cycle-per-instruction advantage, Listing 4.8 runs at exactly the same speed as Listing 4.7. The reason: Both instructions are 2 bytes long, and in both cases it is the 8-cycle instruction fetch time, not the 3 or 4-cycle Execution Unit execution time, that limits performance.

-


- Figure 4.4
  Execution and instruction prefetching sequence for Listing 4.6.

+


+ Figure 4.4
  Execution and instruction prefetching sequence for Listing 4.6.

-

LISTING 4.7 LST4-7.ASM

+

LISTING 4.7 LST4-7.ASM

 ; Measures the performance of repeated MOV AL,0 instructions,
 ; which take 4 cycles each according to Intel's official
@@ -57,9 +49,9 @@
      mov  al,0
      endm
      call ZTimerOff
-
+ -

LISTING 4.8 LST4-8.ASM

+

LISTING 4.8 LST4-8.ASM

 ; Measures the performance of repeated SUB AL,AL instructions,
 ; which take 3 cycles each according to Intel's official
@@ -71,7 +63,7 @@
      sub  al,al
      endm
      call ZTimerOff
-
+

As you can see, it’s easy to be drawn into thinking you’re saving cycles when you’re not. You can only improve the performance of a specific bit of code by reducing the factor—either instruction fetch time or execution time, or sometimes a mix of the two—that’s limiting the performance of that code.

@@ -87,7 +79,7 @@

What we really want is to know how long useful working code takes to run, not how long a single instruction takes, and the Zen timer gives us the tool we need to gather that information. Granted, it would be easier if we could just add up neatly documented instruction execution times—but that’s not going to happen. Without actually measuring the performance of a given code sequence, you simply don’t know how fast it is. For crying out loud, even the people who designed the 8088 at Intel couldn’t tell you exactly how quickly a given 8088 code sequence executes on the PC just by looking at it! Get used to the idea that execution times are only meaningful in context, learn the rules of thumb in this book, and use the Zen timer to measure your code.

-

Approximating Overall Execution Times

+

Approximating Overall Execution Times

Don’t think that because overall instruction execution time is determined by both instruction fetch time and Execution Unit execution time, the two times should be added together when estimating performance. For example, practically speaking, each SHR in Listing 4.5 does not take 8 cycles of instruction fetch time plus 2 cycles of Execution Unit execution time to execute. Figure 4.3 shows that while a given SHR is executing, the fetch of the next SHR is starting, and since the two operations are overlapped for 2 cycles, there’s no sense in charging the time to both instructions. You could think of the extra instruction fetch time for SHR in Listing 4.5 as being 6 cycles, which yields an overall execution time of 8 cycles when added to the 2 cycles of Execution Unit execution time.

@@ -95,7 +87,7 @@

As a working definition, we’ll consider the execution time of a given instruction in a particular context to start when the first byte of the instruction is sent to the Execution Unit and end when the first byte of the next instruction is sent to the EU.

-

What to Do about the Prefetch Queue Cycle-Eater?

+

What to Do about the Prefetch Queue Cycle-Eater?

Reducing the impact of the prefetch queue cycle-eater is one of the overriding principles of high-performance assembly code. How can you do this? One effective technique is to minimize access to memory operands, since such accesses compete with instruction fetching for precious memory accesses. You can also greatly reduce instruction fetch time simply by your choice of instructions: Keep your instructions short. Less time is required to fetch instructions that are 1 or 2 bytes long than instructions that are 5 or 6 bytes long. Reduced instruction fetching lowers minimum execution time (minimum execution time is 4 cycles times the number of instruction bytes) and often leads to faster overall execution.

@@ -118,10 +110,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-06.html b/04-06.html index 7f8c1c9..c3cacf3 100644 --- a/04-06.html +++ b/04-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -41,7 +34,7 @@

All in all, writing good assembler code is as much an art as a science. As a result, you should follow the rules of thumb described here—and then time your code to see how fast it really is. You should experiment freely, but always remember that actual, measured performance is the bottom line.

-

Holding Up the 8088

+

Holding Up the 8088

In this chapter I’ve taken you further and further into the depths of the PC, telling you again and again that you must understand the computer at the lowest possible level in order to write good code. At this point, you may well wonder, “Have we gotten low enough?”

@@ -53,7 +46,7 @@

Let’s start with DRAM refresh, which affects the performance of every program that runs on the PC.

-

Dynamic RAM Refresh: The Invisible Hand

+

Dynamic RAM Refresh: The Invisible Hand

Dynamic RAM (DRAM) refresh is sort of an act of God. By that I mean that DRAM refresh invisibly and inexorably steals a certain fraction of all available memory access time from your programs, when they are accessing memory for code and data. (When they are accessing cache on more recent processors, theoretically the DRAM refresh cycle-eater doesn’t come into play, but there are other cycle-eaters waiting to prey on cache-bound programs.) While you could stop DRAM refresh, you wouldn’t want to since that would be a sure prescription for crashing your computer. In the end, thanks to DRAM refresh, almost all code runs a bit slower on the PC than it otherwise would, and that’s that.

@@ -61,7 +54,7 @@

All of the PC’s system memory consists of DRAM chips. Each DRAM chip in the PC must be completely refreshed about once every four milliseconds in order to ensure the integrity of the data it stores. Obviously, it’s highly desirable that the memory in the PC retain the correct data indefinitely, so each DRAM chip in the PC must always be refreshed within 4 µs of the last refresh. Since there’s no guarantee that a given program will access each and every DRAM block once every 4 µs, the PC contains special circuitry and programming for providing DRAM refresh.

-

How DRAM Refresh Works in the PC

+

How DRAM Refresh Works in the PC

On the original 8088-based IBM PC, timer 1 of the 8253 timer chip is programmed at power-up to generate a signal once every 72 cycles, or once every 15.08µs. That signal goes to channel 0 of the 8237 DMA controller, which requests the bus from the 8088 upon receiving the signal. (DMA stands for direct memory access, the ability of a device other than the 8088 to control the bus and access memory directly, without any help from the 8088.) As soon as the 8088 is between memory accesses, it gives control of the bus to the 8237, which in conjunction with special circuitry on the PC’s motherboard then performs a single 4-cycle read access to 1 of 256 possible addresses, advancing to the next address on each successive access. (The read access is only for the purpose of refreshing the DRAM; the data that is read isn’t used.)

@@ -69,9 +62,8 @@

Don’t sweat the details here. The important point is this: For at least 4 out of every 72 cycles, the original PC’s bus is given over to DRAM refresh and is not available to the 8088, as shown in Figure 4.5. That means that as much as 5.56 percent of the PC’s already inadequate bus capacity is lost. However, DRAM refresh doesn’t necessarily stop the 8088 in its tracks for 4 cycles. The Execution Unit of the 8088 can keep processing while DRAM refresh is occurring, unless the EU needs to access memory. Consequently, DRAM refresh can slow code performance anywhere from 0 percent to 5.56 percent (and actually a bit more, as we'll see shortly), depending on the extent to which DRAM refresh occupies cycles during which the 8088 would otherwise be accessing memory.

-


- Figure 4.5
  The PC bus dynamic RAM (DRAM) refresh.

+


+ Figure 4.5
  The PC bus dynamic RAM (DRAM) refresh.


@@ -90,10 +82,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-07.html b/04-07.html index 5d3056d..c310a15 100644 --- a/04-07.html +++ b/04-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -37,11 +30,11 @@


-

The Impact of DRAM Refresh

+

The Impact of DRAM Refresh

Let’s look at examples from opposite ends of the spectrum in terms of the impact of DRAM refresh on code performance. First, consider the series of MUL instructions in Listing 4.9. Since a 16-bit MUL on the 8088 executes in between 118 and 133 cycles and is only 2 bytes long, there should be plenty of time for the prefetch queue to fill after each instruction, even after DRAM refresh has taken its slice of memory access time. Consequently, the prefetch queue should be able to keep the Execution Unit well-supplied with instruction bytes at all times. Since Listing 4.9 uses no memory operands, the Execution Unit should never have to wait for data from memory, and DRAM refresh should have no impact on performance. (Remember that the Execution Unit can operate normally during DRAM refreshes so long as it doesn’t need to request a memory access from the Bus Interface Unit.)

-

LISTING 4.9 LST4-9.ASM

+

LISTING 4.9 LST4-9.ASM

 ; Measures the performance of repeated MUL instructions,
 ; which allow the prefetch queue to be full at all times,
@@ -54,13 +47,13 @@
      mul  ax
      endm
      call ZTimerOff
-
+

Running Listing 4.9, we find that each MUL executes in 24.72 µs, or exactly 118 cycles. Since that’s the shortest time in which MUL can execute, we can see that no performance is lost to DRAM refresh. Listing 4.9 clearly illustrates that DRAM refresh only affects code performance when a DRAM refresh forces the Execution Unit of the 8088 to wait for a memory access.

Now let’s look at the series of SHR instructions shown in Listing 4.10. Since SHR executes in 2 cycles but is 2 bytes long, the prefetch queue should be empty while Listing 4.10 executes, with the 8088 prefetching instruction bytes non-stop. As a result, the time per instruction of Listing 4.10 should precisely reflect the time required to fetch the instruction bytes.

-

LISTING 4.10 LST4-10.ASM

+

LISTING 4.10 LST4-10.ASM

 ; Measures the performance of repeated SHR instructions,
 ; which empty the prefetch queue, to demonstrate the
@@ -71,7 +64,7 @@
      shr  ax,1
      endm
      call ZTimerOff
-
+

Since 4 cycles are required to read each instruction byte, we’d expect each SHR to execute in 8 cycles, or 1.676 µs, if there were no DRAM refresh. In fact, each SHR in Listing 4.10 executes in 1.81 µs, indicating that DRAM refresh is taking 7.4 percent of the program’s execution time. That’s nearly 2 percent more than our worst-case estimate of the loss to DRAM refresh overhead! In fact, the result indicates that DRAM refresh is stealing not 4, but 5.33 cycles out of every 72 cycles. How can this be?

@@ -79,7 +72,7 @@

Which of the two cases we’ve examined reflects reality? While either case can happen, the latter case—significant performance reduction, ranging as high as 8.33 percent—is far more likely to occur. This is especially true for high-performance assembly code, which uses fast instructions that tend to cause non-stop instruction fetching.

-

What to Do About the DRAM Refresh Cycle-Eater?

+

What to Do About the DRAM Refresh Cycle-Eater?

Hmmm. When we discovered the prefetch queue cycle-eater, we learned to use short instructions. When we discovered the 8-bit bus cycle-eater, we learned to use byte-sized memory operands whenever possible, and to keep word-sized variables in registers. What can we do to work around the DRAM refresh cycle-eater?

@@ -91,7 +84,7 @@

The important thing to understand about DRAM refresh is that it generally slows your code down, and that the extent of that performance reduction can vary considerably and unpredictably, depending on how the DRAM refreshes interact with your code’s pattern of memory accesses. When you use the Zen timer and get a fractional cycle count for the execution time of an instruction, that’s often the DRAM refresh cycle-eater at work. (The display adapter cycleis another possible culprit, and, on 386s and later processors, cache misses and pipeline execution hazards produce this sort of effect as well.) Whenever you get two timing results that differ less or more than they seemingly should, that’s usually DRAM refresh too. Thanks to DRAM refresh, variations of up to 8.33 percent in PC code performance are par for the course.

-

Wait States

+

Wait States

Wait states are cycles during which a bus access by the CPU to a device on the PC’s bus is temporarily halted by that device while the device gets ready to complete the read or write. Wait states are well and truly the lowest level of code performance. Everything we have discussed (and will discuss)—even DMA accesses—can be affected by wait states.

@@ -112,10 +105,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-08.html b/04-08.html index 364daba..0a1d345 100644 --- a/04-08.html +++ b/04-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -45,15 +38,14 @@

First, let’s learn a bit more about wait states by contrast with DRAM refresh. Unlike DRAM refresh, wait states do not occur on any regularly scheduled basis, and are of no particular duration. Wait states can only occur when an instruction performs a memory or I/O read or write. Both the presence of wait states and the number of wait states inserted on any given bus access are entirely controlled by the device being accessed. When it comes to wait states, the CPU is passive, merely accepting whatever wait states the accessed device chooses to insert during the course of the access. All of this makes perfect sense given that the whole point of the wait state mechanism is to allow a device to stretch out any access to itself for however much time it needs to perform the access.

-


- Figure 4.6
  Video wait states inserted by the display adapter.

+


+ Figure 4.6
  Video wait states inserted by the display adapter.

As with DRAM refresh, wait states don’t stop the 8088 completely. The Execution Unit can continue processing while wait states are inserted, so long as the EU doesn’t need to perform a bus access. However, in the PC, wait states most often occur when an instruction accesses a memory operand, so in fact the Execution Unit usually is stopped by wait states. (Instruction fetches rarely wait in an 8088-based PC because system memory is zero-wait-state. AT-class memory systems routinely insert 1 or more wait states, however.)

As it turns out, wait states pose a serious problem in just one area in the PC. While any adapter can insert wait states, in the PC only display adapters do so to the extent that performance is seriously affected.

-

The Display Adapter Cycle-Eater

+

The Display Adapter Cycle-Eater

Display adapters must serve two masters, and that creates a fundamental performance problem. Master #1 is the circuitry that drives the display screen. This circuitry must constantly read display memory in order to obtain the information used to draw the characters or dots displayed on the screen. Since the screen must be redrawn between 50 and 70 times per second, and since each redraw of the screen can require as many as 36,000 reads of display memory (more in Super VGA modes), master #1 is a demanding master indeed. No matter how demanding master #1 gets, however, its needs must always be met—otherwise the quality of the picture on the screen would suffer.

@@ -63,17 +55,15 @@

It turns out that the 8088 CPU has to do a lot of waiting, for three reasons. First, the video circuitry can take as much as about 90 percent of the available display memory access time, as shown in Figure 4.7, leaving as little as about 10 percent of all display memory accesses for the 8088. (These percentages vary considerably among the many EGA and VGA clones.)

-


- Figure 4.7
  Allocation of display memory access.

+


+ Figure 4.7
  Allocation of display memory access.

Second, because the displayed dots (or pixels, short for “picture elements”) must be drawn on the screen at a constant speed, many display adapters provide memory accesses only at fixed intervals. As a result, time can be lost while the 8088 synchronizes with the start of the next display adapter memory access, even if the video circuitry isn’t accessing display memory at that time, as shown in Figure 4.8.

Finally, the time it takes a display adapter to complete a memory access is related to the speed of the clock which generates pixels on the screen rather than to the memory access speed of the 8088. Consequently, the time taken for display memory to complete an 8088 read or write access is often longer than the time taken for system memory to complete an access, even if the 8088 lucks into hitting a free display memory access just as it becomes available, again as shown in Figure 4.8. Any or all of the three factors I’ve described can result in wait states, slowing the 8088 and creating the display adapter cycle.

-


- Figure 4.8
  Display memory access slots.

+


+ Figure 4.8
  Display memory access slots.

If some of this is Greek to you, don’t worry. The important point is that display memory is not very fast compared to normal system memory. How slow is it? Incredibly slow. Remember how slow IBM’s ill-fated PCjrwas? In case you’ve forgotten, I’ll refresh your memory: The PCjrwas at best only half as fast as the PC. The PCjr had an 8088 running at 4.77 MHz, just like the PC—why do you suppose it was so much slower? I’ll tell you why: All the memory in the PCjr was display memory.

@@ -96,10 +86,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-09.html b/04-09.html index b27ff58..2ea20bc 100644 --- a/04-09.html +++ b/04-09.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -39,7 +32,7 @@

The answer varies considerably depending on what display adapter and what display mode we’re talking about. The display adapter cycle-eater is worst with the Enhanced Graphics Adapter (EGA) and the original Video Graphics Array (VGA). (Many VGAs, especially newer ones, insert many fewer wait states than IBM’s original VGA. On the other hand, Super VGAs have more bytes of display memory to be accessed in high-resolution mode.) While the Color/Graphics Adapter (CGA), Monochrome Display Adapter (MDA), and Hercules Graphics Card (HGC) all suffer from the display adapter cycle-eater as well, they suffer to a lesser degree. Since the VGA represents the base standard for PC graphics now and for the foreseeable future, and since it is the hardest graphics adapter to wring performance from, we’ll restrict our discussion to the VGA (and its close relative, the EGA) for the remainder of this chapter.

-

The Impact of the Display Adapter Cycle-Eater

+

The Impact of the Display Adapter Cycle-Eater

Even on the EGA and VGA, the effect of the display adapter cycle-eater depends on the display mode selected. In text mode, the display adapter cycle-eater is rarely a major factor. It’s not that the cycle-eater isn’t present; however, a mere 4,000 bytes control the entire text mode display, and even with the display adapter cycle-eater it just doesn’t take that long to manipulate 4,000 bytes. Even if the display adapter cycle-eater were to cause the 8088 to take as much as 5µs per display memory access—more than five times normal—it would still take only 4,000x 2x 5µs, or 40 µs, to read and write every byte of display memory. That’s a lot of time as measured in 8088 cycles, but it’s less than the blink of an eye in human time, and video performance only matters in human time. After all, the whole point of drawing graphics is to convey visual information, and if that information can be presented faster than the eye can see, that is by definition fast enough.

@@ -51,7 +44,7 @@

That sounds pretty serious, but we did make an unfounded assumption about memory access speed. Let’s get some hard numbers. Listing 4.11 accesses display memory at the 8088’s maximum speed, by way of a REP MOVSW with display memory as both source and destination. The code in Listing 4.11 executes in 3.18 µs per access to display memory—not as long as we had assumed, but a long time nonetheless.

-

LISTING 4.11 LST4-11.ASM

+

LISTING 4.11 LST4-11.ASM

 ; Times speed of memory access to Enhanced Graphics
 ; Adapter graphics mode display memory at A000:0000.
@@ -81,11 +74,11 @@
 ;
      mov  ax,0003h
      int  10h         ;return to text mode
-
+

For comparison, let’s see how long the same code takes when accessing normal system RAM instead of display memory. The code in Listing 4.12, which performs a REP MOVSW from the code segment to the code segment, executes in 1.39 µs per display memory access. That means that on average, 1.79 µs (more than 8 cycles!) are lost to the display adapter cycle-eater on each access. In other words, the display adapter cycle-eater can more than double the execution time of 8088 code!

-

LISTING 4.12 LST4-12.ASM

+

LISTING 4.12 LST4-12.ASM

 ; Times speed of memory access to normal system
 ; memory.
@@ -105,7 +98,7 @@
                       ; is just to measure memory access
                       ; times
      call ZTimerOff
-
+

Bear in mind that we’re talking about a worst case here; the impact of the display adapter cycle-eater is proportional to the percent of time a given code sequence spends accessing display memory.

@@ -126,10 +119,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/04-10.html b/04-10.html index 161a1db..9106ff7 100644 --- a/04-10.html +++ b/04-10.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: In the Lair of the Cycle-Eaters - - @@ -49,7 +42,7 @@

Nonetheless, the display adapter cycle-eater always takes its toll on graphics code. Interestingly, that toll becomes much higher on ATs and 80386 machines because while those computers can execute many more instructions per microsecond than can the 8088-based PC, it takes just as long to access display memory on those computers as on the 8088-based PC. Remember, the limited speed of access to a graphics adapter is an inherent characteristic of the adapter, so the fastest computer around can’t access display memory one iota faster than the adapter will allow.

-

What to Do about the Display Adapter Cycle-Eater?

+

What to Do about the Display Adapter Cycle-Eater?

What can we do about the display adapter cycle-eater? Well, we can minimize display memory accesses whenever possible. In particular, we can try to avoid read/modify/write display memory operations of the sort used to mask individual pixels and clip images. Why? Because read/modify/write operations require two display memory accesses (one read and one write) each time display memory is manipulated. Instead, we should try to use writes of the sort that set all the pixels in a given byte of display memory at once, since such writes don’t require accompanying read accesses. The key here is that only half as many display memory accesses are required to write a byte to display memory as are required to read a byte from display memory, mask part of it off and alter the rest, and write the byte back to display memory. Half as many display memory accesses means half as many display memory wait states.

@@ -67,7 +60,7 @@

It would be handy to explore the display adapter cycle-eater issue in depth, with lots of example code and execution timings, but alas, I don’t have the space for that right now. For the time being, all you really need to know about the display adapter cycle-eater is that on the 8088 you can lose more than 8 cycles of execution time on each access to display memory. For intensive access to display memory, the loss really can be as high as 8cycles (and up to 50, 100, or even more on 486s and Pentiums paired with slow VGAs), while for average graphics code the loss is closer to 4 cycles; in either case, the impact on performance is significant. There is only one way to discover just how significant the impact of the display adapter cycle-eater is for any particular graphics code, and that is of course to measure the performance of that code.

-

Cycle-Eaters: A Summary

+

Cycle-Eaters: A Summary

We’ve covered a great deal of sophisticated material in this chapter, so don’t feel bad if you haven’t understood everything you’ve read; it will all become clear from further reading, especially once you study, time, and tune code that you have written yourself. What’s really important is that you come away from this chapter understanding that on the 8088:

@@ -83,7 +76,7 @@

This basic knowledge about cycle-eaters puts you in a good position to understand the results reported by the Zen timer, and that means that you’re well on your way to writing high-performance assembler code.

-

What Does It All Mean?

+

What Does It All Mean?

There you have it: life under the programming interface. It’s not a particularly pretty picture for the inhabitants of that strange realm where hardware and software meet are little-known cycle-eaters that sap the speed from your unsuspecting code. Still, some of those cycle-eaters can be minimized by keeping instructions short, using the registers, using byte-sized memory operands, and accessing display memory as little as possible. None of the cycle-eaters can be eliminated, and dynamic RAM refresh can scarcely be addressed at all; still, aren’t you better off knowing how fast your code really runs—and why—than you were reading the official execution times and guessing? And while specific cycle-eaters vary in importance on later x86-family processors, with some cycle-eaters vanishing altogether and new ones appearing, the concept that understanding these obscure gremlins is a key to performance remains unchanged, as we’ll see again and again in later chapters.

@@ -104,10 +97,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/05-01.html b/05-01.html index cd609b2..343592e 100644 --- a/05-01.html +++ b/05-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - @@ -37,10 +30,10 @@


-

Chapter 5
+

Chapter 5
Crossing the Border

-

Searching Files with Restartable Blocks

+

Searching Files with Restartable Blocks

We just moved. Those three little words should strike terror into the heart of anyone who owns more than a sleeping bag and a toothbrush. Our last move was the usual zoo—and then some. Because the distance from the old house to the new was only five miles, we used cars to move everything smaller than a washing machine. We have a sizable household—cats, dogs, kids, com, you name it—so the moving process took a number of car trips. A large number—33, to be exact. I personally spent about 15 hours just driving back and forth between the two houses. The move took days to complete.

@@ -66,7 +59,7 @@

And with that, let’s look at a fairly complex application of restartable blocks.

-

Searching for Text

+

Searching for Text

The application we’re going to examine searches a file for a specified string. We’ll develop a program that will search the file specified on the command line for a string (also specified on the comline), then report whether the string was found or not. (Because the searched-for string is obtained via argv, it can’t contain any whitespace characters.)

@@ -78,7 +71,7 @@

Well, it might be instructive to consider how we would search if our search involved only one buffer, already resident in memory. In other words, suppose we don’t have to bother with file handling at all, and further suppose that we don’t have to deal with searching through multiple blocks. After all, that’s a good description of the all-important inner loop of our searching program, where the program will spend virtually all of its time (aside from the unavoidable disk access overhead).

-

Avoiding the String Trap

+

Avoiding the String Trap

The easiest approach would be to use a C/C++ library function. The closest match to what we need is strstr(), which searches one string for the first occurrence of a second string. However, while strstr() would work, it isn’t ideal for our purposes. The problem is this: Where we want to search a fixed-length buffer for the first occurrence of a string, strstr() searches a string for the first occurrence of another string.

@@ -99,10 +92,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/05-02.html b/05-02.html index 434f4e8..a8ca012 100644 --- a/05-02.html +++ b/05-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - @@ -47,33 +40,31 @@ -

Brute-Force Techniques

+

Brute-Force Techniques

Given that no C/C++ library function meets our needs precisely, an obvious alternative approach is the brute-force technique that uses memcmp() to compare every potential matching location in the buffer to the string we’re searching for, as illustrated in Figure 5.1.

By the way, we could, of course, use our own code, working with pointers in a loop, to perform the comparison in place of memcmp(). But memcmp() will almost certainly use the very fast REPZ CMPS instruction. However, never assume! It wouldn’t hurt to use a debugger to check out the actual machine-code implementation of memcmp() from your compiler. If necessary, you could always write your own assembly language implementation of memcmp().

-


- Figure 5.1
  The brute-force searching technique.

+


+ Figure 5.1
  The brute-force searching technique.

Invoking memcmp() for each potential match location works, but entails considerable overhead. Each comparison requires that parameters be pushed and that a call to and return from memcmp() be performed, along with a pass through the comparison loop. Surely there’s a better way!

Indeed there is. We can eliminate most calls to memcmp() by performing a simple test on each potential match location that will reject most such locations right off the bat. We’ll just check whether the first character of the potentially matching buffer location matches the first character of the string we’re searching for. We could make this check by using a pointer in a loop to scan the buffer for the next match for the first character, stopping to check for a match with the rest of the string only when the first character matches, as shown in Figure 5.2.

-

Using memchr()

+

Using memchr()

There’s yet a better way to implement this approach, however. Use the memchr() function, which does nothing more or less than find the next occurrence of a specified character in a fixed-length buffer (presumably by using the extremely efficient REPNZ SCASB instruction, although again it wouldn’t hurt to check). By using memchr() to scan for potential matches that can then be fully tested with memcmp(), we can build a highly efficient search engine that takes good advantage of the information we have about the buffer being searched and the string we’re searching for. Our engine also relies heavily on repeated string instructions, assuming that the memchr() and memcmp() library functions are properly coded.

-


- Figure 5.2
  The faster string-searching technique.

+


+ Figure 5.2
  The faster string-searching technique.

We’re going to go with the this approach in our file-searching program; the only trick lies in deciding how to integrate this approach with restartable blocks in order to search through files larger than our buffer. This certainly isn’t the fastest-possible searching algorithm; as one example, the Boyer-Moore algorithm, which cleverly eliminates many buffer locations as potential matches in the process of checking preceding locations, can be considerably faster. However, the Boyer-Moore algorithm is quite complex to understand and implement, and would distract us from our main focus, restartable blocks, so we’ll save it for a later chapter (Chapter 14, to be precise). Besides, I suspect you’ll find the approach we’ll use to be fast enough for most purposes.

Now that we’ve selected a searching approach, let’s integrate it with file handling and searching through multiple blocks. In other words, let’s make it restartable.

-

Making a Search Restartable

+

Making a Search Restartable

As it happens, there’s no great trick to putting the pieces of this search program together. Basically, we’ll read in a buffer of data (we’ll work with 16K at a time to avoid signed overflow problems with integers), search it for a match with the memchr()/memcmp() engine described, and exit with a “string found” response if the desired string is found.

@@ -102,10 +93,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/05-03.html b/05-03.html index beefdaa..6e76f01 100644 --- a/05-03.html +++ b/05-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - @@ -37,7 +30,7 @@


-

LISTING 5.1 SEARCH.C

+

LISTING 5.1 SEARCH.C

 /* Program to search the file specified by the first command-line
  * argument for the string specified by the second command-line
@@ -204,7 +197,7 @@ main(int argc, char *argv[]) {
    exit(Found);   /* Return the found/not found status as the
                      DOS errorlevel */
 }
-
+


@@ -223,10 +216,6 @@ main(int argc, char *argv[]) {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/05-04.html b/05-04.html index 99b2da6..ec1e233 100644 --- a/05-04.html +++ b/05-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - @@ -37,7 +30,7 @@


-

Interpreting Where the Cycles Go

+

Interpreting Where the Cycles Go

To boost the overall performance of Listing 5.1, I would normally convert SearchForString() to assembly language at this point. However, I’m not going to do that, and the reason is as important a lesson as any discussion of optimized assembly code is likely to be. Take a moment to examine some interesting performance aspects of the C implementation, and all should become much clearer.

@@ -53,7 +46,7 @@

Not likely.

-

Knowing When Assembly Is Pointless

+

Knowing When Assembly Is Pointless

So that’s why we’re not going to go to assembly language in this example—which is not to say it would never be worth converting the search engine in Listing 5.1 to assembly.

@@ -78,10 +71,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/05-05.html b/05-05.html index 5c1ad72..c00fcea 100644 --- a/05-05.html +++ b/05-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Crossing the Border - - @@ -53,7 +46,7 @@

Restartable blocks do minimize the overhead of DOS file-access calls in Listing 5.1; it’s just that there’s no way to reduce that overhead to the point where it becomes worth attempting to further improve the performance of our relatively efficient search engine. Although the search engine is by no means fully optimized, it’s nonetheless as fast as there’s any reason for it to be, given the balance of performance among the components of this program.

-

Always Look Where Execution Is Going

+

Always Look Where Execution Is Going

I’ve explained two important lessons: Know when it’s worth optimizing further, and use restartable blocks to process large data sets as a series of blocks, with each block handled at high speed. The first lesson is less obvious than it seems.

@@ -90,10 +83,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/06-01.html b/06-01.html index bfdd52f..7e41cbb 100644 --- a/06-01.html +++ b/06-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Looking Past Face Value - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Looking Past Face Value - - @@ -37,10 +30,10 @@


-

Chapter 6
+

Chapter 6
Looking Past Face Value

-

How Machine Instructions May Do More Than You Think

+

How Machine Instructions May Do More Than You Think

I first met Jeff Duntemann at an authors’ dinner hosted by PC Tech Journal at Fall Comdex, back in 1985. Jeff was already reasonably well-known as a computer editor and writer, although not as famous as Complete Turbo Pascal, editions 1 through 672 (or thereabouts), TURBO TECHNIX, and PC TECHNIQUES would soon make him. I was fortunate enough to be seated next to Jeff at the dinner table, and, not surprisingly, our often animated conversation revolved around computers, computer writing, and more computers (not necessarily in that order).

@@ -66,7 +59,7 @@

In short, the x86 family can do much more than you think—if you’ll use everything it has to offer. Give it a shot!

-

Memory Addressing and Arithmetic

+

Memory Addressing and Arithmetic

Years ago, I saw a clip on the David Letterman show in which Letterman walked into a store by the name of “Just Lamps” and asked, “So what do you sell here?”

@@ -82,16 +75,16 @@

They perform arithmetic, that’s what they do, and that’s a distinctly different and often useful perspective on memory address calculations.

-

For example, suppose you have an array base address in BX and an index into the array in SI. You could add the two registers together to address memory, like this:

+

For example, suppose you have an array base address in BX and an index into the array in SI. You could add the two registers together to address memory, like this:

 add  bx,si
 mov  al,[bx]
-
+ -

Or you could let the processor do the arithmetic for you in a single instruction:

+

Or you could let the processor do the arithmetic for you in a single instruction:

 mov  al,[bx+si]
-
+


@@ -110,10 +103,6 @@ mov al,[bx+si]
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/06-02.html b/06-02.html index 3375638..ae946b7 100644 --- a/06-02.html +++ b/06-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Looking Past Face Value - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Looking Past Face Value - - @@ -37,14 +30,14 @@


-

The two approaches are functionally interchangeable but not equivalent from a performance standpoint, and which is better depends on the particular context. If it’s a one-shot memory access, it’s best to let the processor perform the addition; it’s generally faster at doing this than a separate ADD instruction would be. If it’s a memory access within a loop, however, it’s advantageous on the 8088 CPU to perform the addition outside the loop, if possible, reducing effective address calculation time inside the loop, as in the following:

+

The two approaches are functionally interchangeable but not equivalent from a performance standpoint, and which is better depends on the particular context. If it’s a one-shot memory access, it’s best to let the processor perform the addition; it’s generally faster at doing this than a separate ADD instruction would be. If it’s a memory access within a loop, however, it’s advantageous on the 8088 CPU to perform the addition outside the loop, if possible, reducing effective address calculation time inside the loop, as in the following:

       add   bx,si
 LoopTop:
       mov   al,[bx]
       inc   bx
       loop  LoopTop
-
+

Here, MOV AL,[BX] is two cycles faster than MOV AL,[BX+SI].

@@ -52,7 +45,7 @@ LoopTop:

The 486 is an odd case, in which the use of an index register or the use of a base register that’s the destination of the previous instruction may slow things down, so it is generally but not always better to perform the addition outside the loop on the 486. All memory addressing calculations are free on the Pentium, however. I’ll discuss 486 performance issues in Chapters 12 and 13, and the Pentium in Chapters 19 through 21.

-

Math via Memory Addressing

+

Math via Memory Addressing

You’re probably not particularly wowed to hear that you can use addressing modes to perform memory addressing arithmetic that would otherwise have to be performed with separate arithmetic instructions. You may, however, be a tad more interested to hear that you can also use addressing modes to perform arithmetic that has nothing to do with memory addressing, and with a couple of advantages over arithmetic instructions, at that.

@@ -62,75 +55,73 @@ LoopTop:

What does that give us? Two things that ADD doesn’t provide: the ability to perform addition with either two or three operands, and the ability to store the result in any register, not just in one of the source operands.

-

Imagine that we want to add BX to DI, add two to the result, and store the result in AX. The obvious solution is this:

+

Imagine that we want to add BX to DI, add two to the result, and store the result in AX. The obvious solution is this:

 mov  ax,bx
 add  ax,di
 add  ax,2
-
+ -

(It would be more compact to increment AX twice than to add two to it, and would probably be faster on an 8088, but that’s not what we’re after at the moment.) An elegant alternative solution is simply:

+

(It would be more compact to increment AX twice than to add two to it, and would probably be faster on an 8088, but that’s not what we’re after at the moment.) An elegant alternative solution is simply:

 lea  ax,[bx+di+2]
-
+ -

Likewise, either of the following would copy SI plus two to DI

+

Likewise, either of the following would copy SI plus two to DI

 mov  di,si
 add  di,2
-
+ -

or:

+

or:

 lea  di,[si+2]
-
+

Mind you, the only components LEA can add are BX or BP, SI or DI, and a constant displacement, so it’s not going to replace ADD most of the time. Also, LEA is considerably slower than ADD on an 8088, although it is just as fast as ADD on a 286 or 386 when fewer than three memory addressing components are used. LEA is 1 cycle slower than ADD on a 486 if the sum of two registers is used to point to memory, but no slower than ADD on a Pentium. On both a 486 and Pentium, LEA can also be slowed down by addressing interlocks.

-


- Figure 6.1
  Operation of ADD Reg,Reg vs. LEA Reg,{Addr}.

+


+ Figure 6.1
  Operation of ADD Reg,Reg vs. LEA Reg,{Addr}.

-

The Wonders of LEA on the 386

+

The Wonders of LEA on the 386

LEA really comes into its own as a “super-ADD” instruction on the 386, 486, and Pentium, where it can take advantage of the enhanced memory addressing modes of those processors. (The 486 and Pentium offer the same modes as the 386, so I’ll refer only to the 386 from now on.) The 386 can do two very interesting things: It can use any 32-bit register (EAX, EBX, and so on) as the memory addressing base register and/or the memory addressing index register, and it can multiply any 32-bit register used as an index by two, four, or eight in the process of calculating a memory address, as shown in Figure 6.2. Let’s see what that’s good for.

Well, the obvious advantage is that any two 32-bit registers, or any 32-bit register and any constant, or any two 32-bit registers and any constant, can be added together, with the result stored in any register. This makes the 32-bit LEA much more generally useful than the standard 16-bit LEA in the role of an ADD with an independent destination.

-


- Figure 6.2
  Operation of the 32-bit LEA reg,[Addr].

+


+ Figure 6.2
  Operation of the 32-bit LEA reg,[Addr].

But what else can LEA do on a 386, besides add?

-

It can multiply any register used as an index. LEA can multiply only by the power-of-two values 2, 4, or 8, but that’s useful more often than you might imagine, especially when dealing with pointers into tables. Besides, multiplying by 2, 4, or 8 amounts to a left shift of 1, 2, or 3 bits, so we can now add up to two 32-bit registers and a constant, and shift (or multiply) one of the registers to some extent—all with a single instruction. For example,

+

It can multiply any register used as an index. LEA can multiply only by the power-of-two values 2, 4, or 8, but that’s useful more often than you might imagine, especially when dealing with pointers into tables. Besides, multiplying by 2, 4, or 8 amounts to a left shift of 1, 2, or 3 bits, so we can now add up to two 32-bit registers and a constant, and shift (or multiply) one of the registers to some extent—all with a single instruction. For example,

 lea  edi,TableBase[ecx+edx*4]
-
+ -

replaces all this

+

replaces all this

 mov  edi,edx
 shl  edi,2
 add  edi,ecx
 add  edi,offset TableBase
-
+

when pointing to an entry in a doubly indexed table.

-

Multiplication with LEA Using Non-Powers of Two

+

Multiplication with LEA Using Non-Powers of Two

-

Are you impressed yet with all that LEA can do on the 386? Believe it or not, one more feature still awaits us. LEA can actually perform a fast multiply of a 32-bit register by some values other than powers of two. You see, the same 32-bit register can be both base and index on the 386, and can be scaled as the index while being used unchanged as the base. That means that you can, for example, multiply EBX by 5 with:

+

Are you impressed yet with all that LEA can do on the 386? Believe it or not, one more feature still awaits us. LEA can actually perform a fast multiply of a 32-bit register by some values other than powers of two. You see, the same 32-bit register can be both base and index on the 386, and can be scaled as the index while being used unchanged as the base. That means that you can, for example, multiply EBX by 5 with:

 lea ebx,[ebx+ebx*4]
-
+ -

Without LEA and scaling, multiplication of EBX by 5 would require either a relatively slow MUL, along with a set-up instruction or two, or three separate instructions along the lines of the following

+

Without LEA and scaling, multiplication of EBX by 5 would require either a relatively slow MUL, along with a set-up instruction or two, or three separate instructions along the lines of the following

 mov  edx,ebx
 shl  ebx,2
 add  ebx,edx
-
+

and would in either case require the destruction of the contents of another register.

@@ -163,10 +154,6 @@ add ebx,edx
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/07-01.html b/07-01.html index 9afe4e0..6a41bd7 100644 --- a/07-01.html +++ b/07-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - @@ -37,10 +30,10 @@


-

Chapter 7
+

Chapter 7
Local Optimization

-

Optimizing Halfway between Algorithms and Cycle Counting

+

Optimizing Halfway between Algorithms and Cycle Counting

You might not think it, but there’s much to learn about performance programming from the Great Buffalo Sauna Fiasco. To wit:

@@ -74,7 +67,7 @@

And yes, in case you’re wondering, the above story is indeed true. Was I there? Let me put it this way: If I were, I’d never admit it!

-

When LOOP Is a Bad Idea

+

When LOOP Is a Bad Idea

Let’s examine first an instruction that is less than it appears to be: LOOP. There’s no mystery about what LOOP does; it decrements CX and branches if CX doesn’t decrement to zero. It’s so beautifully suited to the task of counting down loops that any experienced x86 programmer instinctively stuffs the loop count in CX and reaches for LOOP when setting up a loop. That’s fine—LOOP does, of course, work as advertised—but there is one problem:

@@ -113,10 +106,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/07-02.html b/07-02.html index fb490d8..c0a9229 100644 --- a/07-02.html +++ b/07-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - @@ -37,31 +30,31 @@


-

By the way, don’t fall victim to the lures of JCXZ and do something like this:

+

By the way, don’t fall victim to the lures of JCXZ and do something like this:

 and     cx,ofh          ;Isolate the desired field
 jcxz    SkipLoop        ;If field is 0, don’t bother
-
+ -

The AND instruction has already set the Zero flag, so this

+

The AND instruction has already set the Zero flag, so this

 and     cx,0fh           ;Isolate the desired field
 jz      SkipLoop         ;If field is 0, don’t bother
-
+

will do just fine and is faster on all processors. Use JCXZ only when the Zero flag isn’t already set to reflect the status of CX.

-

The Lessons of LOOP and JCXZ

+

The Lessons of LOOP and JCXZ

What can we learn from LOOP and JCXZ? First, that a single instruction that is intended to do a complex task is not necessarily faster than several instructions that together do the same thing. Second, that the relative merits of instructions and optimization rules vary to a surprisingly large degree across the x86 family.

In particular, if you’re going to write 386 protected mode code, which will run only on the 386, 486, and Pentium, you’d be well advised to rethink your use of the more esoteric members of the x86 instruction set. LOOP, JCXZ, the various accumulator-specific instructions, and even the string instructions in many circumstances no longer offer the advantages they did on the 8088. Sometimes they’re just not any faster than more general instructions, so they’re not worth going out of your way to use; sometimes, as with LOOP, they’re actually slower, and you’d do well to avoid them altogether in the 386/486 world. Reviewing the instruction cycle times in the MASM or TASM manuals, or looking over the cycle times in Intel’s literature, is a good place to start; published cycle times are closer to actual execution times on the 386 and 486 than on the 8088, and are reasonably reliable indicators of the relative performance levels of x86 instructions.

-

Avoiding LOOPS of Any Stripe

+

Avoiding LOOPS of Any Stripe

Cycle counting and directly substituting instructions (DEC CX/JNZ for LOOP, for example) are techniques that belong at the lowest level of optimization. It’s an important level, but it’s fairly mechanical; once you’ve learned the capabilities and relative performance levels of the various instructions, you should be able to select the best instructions fairly easily. What’s more, this is a task at which compilers excel. What I’m saying is that you shouldn’t get too caught up in counting cycles because that’s a small (albeit important) part of the optimization picture, and not the area in which your greatest advantage lies.

-

Local Optimization

+

Local Optimization

One level at which assembly language programming pays off handsomely is that of local optimization; that is, selecting the best sequence of instructions for a task. The key to local optimization is viewing the 80x86 instruction set as a set of building blocks, each with unique characteristics. Your job is to sequence those blocks so that they perform well. It doesn’t matter what the instructions are intended to do or what their names are; all that matters is what they do.

@@ -98,10 +91,6 @@ jz SkipLoop ;If field is 0, don’t bother
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/07-03.html b/07-03.html index 5fa347c..455846e 100644 --- a/07-03.html +++ b/07-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - @@ -37,7 +30,7 @@


-

LISTING 7.1 L7-1.ASM

+

LISTING 7.1 L7-1.ASM

 ; Program to illustrate searching through a buffer of a specified
 ; length until either a specified byte or a zero byte is
@@ -129,9 +122,9 @@ ByteFound:
       ret
 SearchMaxLengthendp
       end   Start
-
+ -

Unrolling Loops

+

Unrolling Loops

Listing 7.2 takes a different tack, unrolling the loop so that four bytes are checked for each LOOP performed. The same instructions are used inside the loop in each listing, but Listing 7.2 is arranged so that three-quarters of the LOOPs are eliminated. Listings 7.1 and 7.2 perform exactly the same task, and they use the same instructions in the loop—the searching algorithm hasn’t changed in any way—but we have sequenced the instructions differently in Listing 7.2, and that makes all the difference.

@@ -152,10 +145,6 @@ SearchMaxLengthendp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/07-04.html b/07-04.html index d402118..0730f35 100644 --- a/07-04.html +++ b/07-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - @@ -37,7 +30,7 @@


-

LISTING 7.2 L7-2.ASM

+

LISTING 7.2 L7-2.ASM

 ; Program to illustrate searching through a buffer of a specified
 ; length until a specified zero byte is encountered.
@@ -167,7 +160,7 @@ ByteFound:
      ret
 SearchMaxLengthendp
      end    Start
-
+

How much difference? Listing 7.2 runs in 121 µs—40 percent faster than Listing 7.1, even though Listing 7.2 still uses LOOP rather than DEC CX/JNZ. (The loop in Listing 7.2 could be unrolled further, too; it’s just a question of how much more memory you want to trade for ever-decreasing performance benefits.) That’s typical of local optimization; it won’t often yield the order-of-magnitude improvements that algorithmic improvements can produce, but it can get you a critical 50 percent or 100 percent improvement when you’ve exhausted all other avenues.

@@ -196,10 +189,6 @@ SearchMaxLengthendp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/07-05.html b/07-05.html index ab61b7e..8cf1670 100644 --- a/07-05.html +++ b/07-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Local Optimization - - @@ -37,20 +30,20 @@


-

Rotating and Shifting with Tables

+

Rotating and Shifting with Tables

As another example of local optimization, consider the matter of rotating or shifting a mask into position. First, let’s look at the simple task of setting bit N of AX to 1.

-

The obvious way to do this is to place N in CL, rotate the bit into position, and OR it with AX, as follows:

+

The obvious way to do this is to place N in CL, rotate the bit into position, and OR it with AX, as follows:

 MOV  BX,1
 SHL  BX,CL
 OR   AX,BX
-
+

This solution is obvious because it takes good advantage of the special ability of the x86 family to shift or rotate by the variable number of bits specified by CL. However, it takes an average of about 45 cycles on an 8088. It’s actually far faster to precalculate the results, pass the bit number in BX, and look the shifted bit up, as shown in Listing 7.3.

-

LISTING 7.3 L7-3.ASM

+

LISTING 7.3 L7-3.ASM

      SHL  BX,1                ;prepare for word sized look up
      OR   AX,ShiftTable[BX]   ;look up the bit and OR it in
@@ -61,13 +54,13 @@ BIT_PATTERN=0001H
      DW   BIT_PATTERN
 BIT_PATTERN=BIT_PATTERN SHL 1
      ENDM
-
+

Even though it accesses memory, this approach takes only 20 cycles—more than twice as fast as the variable shift. Once again, we were able to improve performance considerably—not by knowing the fastest instructions, but by selecting the fastest sequence of instructions.

In the particular example above, we once again run into the difficulty of optimizing across the x86 family. The table lookup is faster on the 8088 and 286, but it’s slightly slower on the 386 and no faster on the 486. However, 386/486-specific code could use enhanced addressing to accomplish the whole job in just one instruction, along the lines of the code snippet in Listing 7.4.

-

LISTING 7.4 L7-4.ASM

+

LISTING 7.4 L7-4.ASM

      OR   EAX,ShiftTable[EBX*4]    ;look up the bit and OR it in
           :
@@ -77,7 +70,7 @@ BIT_PATTERN=0001H
      DD   BIT_PATTERN
 BIT_PATTERN=BIT_PATTERN SHL 1
      ENDM
-
+ @@ -87,7 +80,7 @@ BIT_PATTERN=BIT_PATTERN SHL 1
-

NOT Flips Bits—Not Flags

+

NOT Flips Bits—Not Flags

The NOT instruction flips all the bits in the operand, from 0 to 1 or from 1 to 0. That’s as simple as could be, but NOT nonetheless has a minor but interesting talent: It doesn’t affect the flags. That can be irritating; I once spent a good hour tracking down a bug caused by my unconscious assumption that NOT does set the flags. After all, every other arithmetic and logical instruction sets the flags; why not NOT? Probably because NOT isn’t considered to be an arithmetic or logical instruction at all; rather, it’s a data manipulation instruction, like MOV and the various rotates. (These are RCR, RCL, ROR, and ROL, which affect only the Carry and Overflow flags.) NOT is often used for tasks, such as flipping masks, where there’s no reason to test the state of the result, and in that context it can be handy to keep the flags unmodified for later testing.

@@ -101,13 +94,13 @@ BIT_PATTERN=BIT_PATTERN SHL 1

The x86 instruction set offers many ways to accomplish almost any task. Understanding the subtle distinctions between the instructions—whether and which flags are set, for example—can be critical when you’re trying to optimize a code sequence and you’re running out of registers, or when you’re trying to minimize branching.

-

Incrementing with and without Carry

+

Incrementing with and without Carry

Another case in which there are two slightly different ways to perform a task involves adding 1 to an operand. You can do this with INC, as in INC AX, or you can do it with ADD, as in ADD AX,1. What’s the difference? The obvious difference is that INC is usually a byte or two shorter (the exception being ADD AL,1, which at two bytes is the same length as INC AL), and is faster on some processors. Less obvious, but no less important, is that ADD sets the Carry flag while INC leaves the Carry flag untouched.

Why is that important? Because it allows INC to function as a data pointer manipulation instruction for multi-word arithmetic. You can use INC to advance the pointers in code like that shown in Listing 7.5 without having to do any work to preserve the Carry status from one addition to the next.

-

LISTING 7.5 L7-5.ASM

+

LISTING 7.5 L7-5.ASM

         CLC                  ;clear the Carry for the initial addition
 LOOP_TOP:
@@ -118,11 +111,11 @@ LOOP_TOP:
         INC    DI            ;point to next dest operand word
         INC    DI
         LOOP   LOOP_TOP
-
+

If ADD were used, the Carry flag would have to be saved between additions, with code along the lines shown in Listing 7.6.

-

LISTING 7.6 L7-6.ASM

+

LISTING 7.6 L7-6.ASM

      CLC            ;clear the carry for the initial addition
 LOOP_TOP:
@@ -133,7 +126,7 @@ LOOP_TOP:
      ADD  DI,2      ;point to next dest operand word
      SAHF           ;restore the carry flag
      LOOP LOOP_TOP
-
+

It’s not that the Listing 7.6 approach is necessarily better or worse; that depends on the processor and the situation. The Listing 7.6 approach is different, and if you understand the differences, you’ll be able to choose the best approach for whatever code you happen to write. (DEC has the same property of preserving the Carry flag, by the way.)

@@ -151,17 +144,17 @@ LOOP_TOP:

The other interesting point about the last example is the use of LAHF and SAHF, which transfer the low byte of the FLAGS register to and from AH, respectively. These instructions were created to help provide compatibility with the 8080’s (that’s 8080, not 8088) PUSH PSW and POP PSW instructions, but turn out to be compact (one byte) instructions for saving and restoring the arithmetic flags. A word of caution, however: SAHF restores the Carry, Zero, Sign, Auxiliary Carry, and Parity flags—but not the Overflow flag, which resides in the high byte of the FLAGS register. Also, be aware that LAHF and SAHF provide a fast way to preserve the flags on an 8088 but are relatively slow instructions on the 486 and Pentium.

-

There are times when it’s a clear liability that INC doesn’t set the Carry flag. For instance

+

There are times when it’s a clear liability that INC doesn’t set the Carry flag. For instance

 INC   AX
 ADC   DX,0
-
+ -

does not increment the 32-bit value in DX:AX. To do that, you’d need the following:

+

does not increment the 32-bit value in DX:AX. To do that, you’d need the following:

 ADD   AX,1
 ADC   DX,0
-
+

As always, pay attention!

@@ -182,10 +175,6 @@ ADC DX,0
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/08-01.html b/08-01.html index c5fb971..1e21b85 100644 --- a/08-01.html +++ b/08-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - @@ -37,10 +30,10 @@


-

Chapter 8
+

Chapter 8
Speeding Up C with Assembly Language

-

Jumping Languages When You Know It’ll Help

+

Jumping Languages When You Know It’ll Help

When I was a senior in high school, a pop song called “Seasons in the Sun,” sung by one Terry Jacks, soared up the pop charts and spent, as best I can recall, two straight weeks atop Kasey Kasem’s American Top 40. “Seasons in the Sun” wasn’t a particularly good song, primarily because the lyrics were silly. I’ve never understood why the song was a hit, but, as so often happens with undistinguished but popular music by forgotten one- or two-shot groups (“Don’t Pull Your Love Out on Me Baby,” “Billy Don’t Be a Hero,” et al.), I heard it everywhere for a month or so, then gave it not another thought for 15 years.

@@ -62,7 +55,7 @@

Apropos of which, when was the last time you heard of Terry Jacks?

-

Billy, Don’t Be a Compiler

+

Billy, Don’t Be a Compiler

The key to optimizing C programs with assembly language is, as always, writing good assembly language code, but with an added twist. Rule 1 when converting C code to assembly is this: Don’t think like a compiler. That’s more easily said than done, especially when the C code you’re converting is readily available as a model and the assembly code that the compiler generates is available as well. Nevertheless, the principle of not thinking like a compiler is essential, and is, in one form or another, the basis for all that I’ll discuss below.

@@ -80,7 +73,7 @@ -

Don’t Call Your Functions on Me, Baby

+

Don’t Call Your Functions on Me, Baby

In order to think differently from a compiler, you must understand both what compilers and C programmers tend to do and how that differs from what assembly language does well. In this pursuit, it can be useful to examine the code your compiler generates, either by viewing the code in a debugger or by having the compiler generate an assembly language output file. (The latter is done with /Fa or /Fc in Microsoft C/C++ and -S in Borland C++.)

@@ -105,10 +98,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/08-02.html b/08-02.html index 6812edc..b8d55f3 100644 --- a/08-02.html +++ b/08-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - @@ -37,13 +30,13 @@


-

Stack Frames Slow So Much

+

Stack Frames Slow So Much

C compilers work within the stack frame model, whereby variables reside in a block of stack memory and are accessed via offsets from BP. Compilers may store a couple of variables in registers and may briefly keep other variables in registers when they’re used repeatedly, but the stack frame is the underlying architecture. It’s a nice architecture; it’s flexible, convenient, easy to program, and makes for fairly compact code. However, stack frames have a few drawbacks. They must be constructed and destroyed, which takes both time and code. They are so easy to use that they tend to bias the assembly language programmer in favor of accessing memory variables more often than might be necessary. Finally, you cannot use BP as a general-purpose register if you intend to access a stack frame, and having that seventh register available is sometimes useful indeed.

That doesn’t mean you shouldn’t use stack frames, which are useful and often necessary. Just don’t fall victim to their undeniable charms.

-

Torn Between Two Segments

+

Torn Between Two Segments

C compilers are not terrific at handling segments. Some compilers can efficiently handle a single far pointer used in a loop by leaving ES set for the duration of the loop. But two far pointers used in the same loop confuse every compiler I’ve seen, causing the full segment:offset address to be reloaded each time either pointer is used.

@@ -57,7 +50,7 @@

In assembly language you have full control over segments. Use it, and, if necessary, reorganize your code to minimize segment loading.

-

Why Speeding Up Is Hard to Do

+

Why Speeding Up Is Hard to Do

You might think that the most obvious advantage assembly language has over C is that it allows the use of all forms of instructions and all registers in all ways, whereas C compilers tend to use a subset of registers and instructions in a limited number of ways. Yes and no. It’s true that C compilers typically don’t generate instructions such as XLAT, rotates, or the string instructions. On the other hand, XLAT and rotates are useful in a limited set of circumstances, and string instructions are used in the C library functions. In fact, C library code is likely to be carefully optimized by experts, and may be much better than equivalent code you’d produce yourself.

@@ -75,21 +68,20 @@

True optimization requires rethinking your code to take advantage of assembly language. A C loop that searches through an integer array for matches might compile

-


- Figure 8.1
  Tweaked compiler output for a loop.

+


+ Figure 8.1
  Tweaked compiler output for a loop.

to something like Figure 8.1A. You might look at that and tweak it to the code shown in Figure 8.1B.

-

Congratulations! You’ve successfully eliminated all stack frame access, you’ve used LOOP (although DEC SI/JNZ is actually faster on 386 and later machines, as I explained in the last chapter), and you’ve used a string instruction. Unfortunately, the new code isn’t going to run very much faster. Maybe 25 percent faster, maybe a little more. Big deal. You’ve eliminated the trappings of the compiler—the stack frame and the restricted register usage—but you’re still thinking like the compiler. Try this:

+

Congratulations! You’ve successfully eliminated all stack frame access, you’ve used LOOP (although DEC SI/JNZ is actually faster on 386 and later machines, as I explained in the last chapter), and you’ve used a string instruction. Unfortunately, the new code isn’t going to run very much faster. Maybe 25 percent faster, maybe a little more. Big deal. You’ve eliminated the trappings of the compiler—the stack frame and the restricted register usage—but you’re still thinking like the compiler. Try this:

 repnz scasw
 jz    Match
-
+

It’s a simple example—but, I hope, a convincing one. Stretch your brain when you optimize.

-

Taking It to the Limit

+

Taking It to the Limit

The ultimate in assembly language optimization comes when you change the rules; that is, when you reorganize the entire program to allow the use of better assembly language code in the small section of code that most affects overall performance. For example, consider that the data searched in the last example is stored in an array of structures, with each structure in the array containing other information as well. In this situation, REP SCASW couldn’t be used because the data searched through wouldn’t be contiguous.

@@ -121,7 +113,7 @@ jz Match

That said, let me show some of these precepts in action.

-

A C-to-Assembly Case Study

+

A C-to-Assembly Case Study

Listing 8.1 is the sample C application I’m going to use to examine optimization in action. Listing 8.1 isn’t really complete—it doesn’t handle the “no-matches” case well, and it assumes that the sum of all matches will fit into an int—but it will do just fine as an optimization example.

@@ -142,10 +134,6 @@ jz Match
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/08-03.html b/08-03.html index 979b426..4e12a5e 100644 --- a/08-03.html +++ b/08-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - @@ -37,7 +30,7 @@


-

LISTING 8.1 L8-1.C

+

LISTING 8.1 L8-1.C

 /* Program to search an array spanning a linked list of variable-
    sized blocks, for all entries with a specified ID number,
@@ -158,17 +151,16 @@ unsigned int FindIDAverage(unsigned int SearchedForID,
    else
       return(IDMatchSum / IDMatchCount);
 }
-
+

The main body of Listing 8.1 constructs a linked list of memory blocks of various sizes and stores an array of structures across those blocks, as shown in Figure 8.2. The function FindIDAverage in Listing 8.1 searches through that array for all matches to a specified ID number and returns the average value of all such matches. FindIDAverage contains two nested loops, the outer one repeating once for each linked block and the inner one repeating once for each array element in each block. The inner loop—the critical one—is compact, containing only four statements, and should lend itself rather well to compiler optimization.

-


- Figure 8.2
  Linked array storage format (version 1).

+


+ Figure 8.2
  Linked array storage format (version 1).

As it happens, Microsoft C/C++ does optimize the inner loop of FindIDAverage nicely. Listing 8.2 shows the code Microsoft C/C++ generates for the inner loop, consisting of a mere seven assembly language instructions inside the loop. The compiler is smart enough to convert the loop index variable, which counts up but is used for nothing but counting loops, into a count-down variable so that the LOOP instruction can be used.

-

LISTING 8.2 L8-2.COD

+

LISTING 8.2 L8-2.COD

 ; Code generated by Microsoft C for inner loop of FindIDAverage.
 ;|*** for (WorkingBlockCount=0;
@@ -199,7 +191,7 @@ $I265:
           mov     WORD PTR [bp-2],di        ;IDMatchSum
           mov     WORD PTR [bp-4],dx        ;IDMatchCount
 $FB264:
-
+


@@ -218,10 +210,6 @@ $FB264:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/08-04.html b/08-04.html index ecc8a42..0fd1acc 100644 --- a/08-04.html +++ b/08-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - @@ -39,7 +32,7 @@

It’s hard to squeeze much more performance from this code by tweaking it, as exemplified by Listing 8.3, a fine-tuned assembly version of FindIDAverage that was produced by looking at the assembly output of MS C/C++ and tightening it. Listing 8.3 eliminates all stack frame access in the inner loop, but that’s about all the tightening there is to do. The result, as shown in Table 8.1, is that Listing 8.3 runs a modest 11 percent faster than Listing 8.1 on a 386. The results could vary considerably, depending on the nature of the data set searched through (average block size and frequency of matches). But, then, understanding the typical and worst case conditions is part of optimization, isn’t it?

-

LISTING 8.3 L8-3.ASM

+

LISTING 8.3 L8-3.ASM

 ; Typically optimized assembly language version of FindIDAverage.
 SearchedForID   equ     4      ;Passed parameter offsets in the
@@ -54,7 +47,7 @@ DATA_ELEMENT_SIZE equ   4      ;Number of bytes in struct DataElement
         .model  small
         .code
         public  _FindIDAverage
-
+ @@ -156,7 +149,7 @@ DATA_ELEMENT_SIZE equ 4 ;Number of bytes in struct DataElement
-
+
 _FindIDAverage  proc    near
         push    bp              ;Save caller’s stack frame
@@ -201,11 +194,11 @@ Done:   pop     si              ;Restore C register variables
         ret
 _FindIDAverage  ENDP
         end
-
+

Listing 8.4 tosses some sophisticated optimization techniques into the mix. The loop is unrolled eight times, eliminating a good deal of branching, and SCASW is used instead of CMP [DI],AX. (Note, however, that SCASW is in fact slower than CMP [DI],AX on the 386 and 486, and is sometimes faster on the 286 and 8088 only because it’s shorter and therefore may prefetch faster.) This advanced tweaking produces a 39 percent improvement over the original C code—substantial, but not a tremendous return for the optimization effort invested.

-

LISTING 8.4 L8-4.ASM

+

LISTING 8.4 L8-4.ASM

 ; Heavily optimized assembly language version of FindIDAverage.
 ; Features an unrolled loop and more efficient pointer use.
@@ -295,7 +288,7 @@ Done:   pop     si              ;Restore C register variables
         ret
 _FindIDAverage  ENDP
         end
-
+


@@ -314,10 +307,6 @@ _FindIDAverage ENDP
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/08-05.html b/08-05.html index 7089f25..1f332c0 100644 --- a/08-05.html +++ b/08-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Speeding Up C with Assembly Language - - @@ -39,7 +32,7 @@

Listings 8.5 and 8.6 together go the final step and change the rules in favor of assembly language. Listing 8.5 creates the same list of linked blocks as Listing 8.1. However, instead of storing an array of structures within each block, it stores two arrays in each block, one consisting of ID numbers and the other consisting of the corresponding values, as shown in Figure 8.3. No information is lost; the data is merely rearranged.

-

LISTING 8.5 L8-5.C

+

LISTING 8.5 L8-5.C

 /* Program to search an array spanning a linked list of variable-
    sized blocks, for all entries with a specified ID number,
@@ -59,11 +52,10 @@ void main(void);
 void exit(int);
 extern unsigned int FindIDAverage2(unsigned int,
                                    struct BlockHeader *);
-
+ -


- Figure 8.3
  Linked array storage format (version 2).

+


+ Figure 8.3
  Linked array storage format (version 2).

 /* Structure that starts each variable-sized block */
 struct BlockHeader {
@@ -117,9 +109,9 @@ void main(void) {
          IDToFind, FindIDAverage2(IDToFind, BaseArrayBlockPointer));
    exit(0);
 }
-
+ -

LISTING 8.6 L8-6.ASM

+

LISTING 8.6 L8-6.ASM

 ; Alternative optimized assembly language version of FindIDAverage
 ; requires data organized as two arrays within each block rather
@@ -188,7 +180,7 @@ Done:   pop     si              ;Restore C register variables
         ret
 _FindIDAverage2 ENDP
         end
-
+

The whole point of this rearrangement is to allow us to use REP SCASW to search through each block, and that’s exactly what FindIDAverage2 in Listing 8.6 does. The result: Listing 8.6 calculates the average about three times as fast as the original C implementation and more than twice as fast as Listing 8.4, heavily optimized as the latter code is.

@@ -211,10 +203,6 @@ _FindIDAverage2 ENDP
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-01.html b/09-01.html index 5cf685e..5cfafb6 100644 --- a/09-01.html +++ b/09-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -37,19 +30,19 @@


-

Chapter 9
+

Chapter 9
Hints My Readers Gave Me

-

Optimization Odds and Ends from the Field

+

Optimization Odds and Ends from the Field

Back in high school, I took a pre-calculus class from Mr. Bourgeis, whose most notable characteristics were incessant pacing and truly enormous feet. My friend Barry, who sat in the back row, right behind me, claimed that it was because of his large feet that Mr. Bourgeis was so restless. Those feet were so heavy, Barry hypothesized, that if Mr. Bourgeis remained in any one place for too long, the floor would give way under the strain, plunging the unfortunate teacher deep into the mantle of the Earth and possibly all the way through to China. Many amusing cartoons were drawn to this effect.

Unfortunately, Barry was too busy drawing cartoons, or, alternatively, sleeping, to actually learn any math. In the long run, that didn’t turn out to be a handicap for Barry, who went on to become vice-president of sales for a ham-packing company, where presumably he was rarely called upon to derive the quadratic equation. Barry’s lack of scholarship caused some problems back then, though. On one memorable occasion, Barry was half-asleep, with his eyes open but unfocused and his chin balanced on his hand in the classic “if I fall asleep my head will fall off my hand and I’ll wake up” posture, when Mr. Bourgeis popped a killer problem:

-

“Barry, solve this for X, please.” On the blackboard lay the equation:

+

“Barry, solve this for X, please.” On the blackboard lay the equation:

 X - 1 = 0
-
+

“Minus 1,” Barry said promptly.

@@ -71,23 +64,23 @@ X - 1 = 0

I like to think I know more about performance programming than Barry knew about math. Nonetheless, I always welcome good ideas and comments, and many readers have sent me a slew of those over the years. So in this chapter, I think I’ll return the favor by devoting a chapter to reader feedback.

-

Another Look at LEA

+

Another Look at LEA

-

Several people have pointed out that while LEA is great for performing certain additions (see Chapter 6), it isn’t a perfect replacement for ADD. What’s the difference? LEA, an addressing instruction by trade, doesn’t affect the flags, while the arithmetic ADD instruction most certainly does. This is no problem when performing additions that involve only quantities that fit in one machine word (32 bits in 386 protected mode, 16 bits otherwise), but it renders LEA useless for multiword operations, which use the Carry flag to tie together partial results. For example, these instructions

+

Several people have pointed out that while LEA is great for performing certain additions (see Chapter 6), it isn’t a perfect replacement for ADD. What’s the difference? LEA, an addressing instruction by trade, doesn’t affect the flags, while the arithmetic ADD instruction most certainly does. This is no problem when performing additions that involve only quantities that fit in one machine word (32 bits in 386 protected mode, 16 bits otherwise), but it renders LEA useless for multiword operations, which use the Carry flag to tie together partial results. For example, these instructions

 ADD  EAX,EBX
 ADC  EDX,ECX
-
+ -

could not be replaced

+

could not be replaced

 LEA  EAX,[EAX+EBX]
 ADC  EDX,ECX
-
+

because LEA doesn’t affect the Carry flag.

-

The no-carry characteristic of LEA becomes a distinct advantage when performing pointer arithmetic, however. For instance, the following code uses LEA to advance the pointers while adding one 128-bit memory variable to another such variable:

+

The no-carry characteristic of LEA becomes a distinct advantage when performing pointer arithmetic, however. For instance, the following code uses LEA to advance the pointers while adding one 128-bit memory variable to another such variable:

    MOV   ECX,4   ;# of 32-bit words to add
    CLC
@@ -99,7 +92,7 @@ ADDLOOP:
    LEA   ESI,[ESI+4]  ;advance one array’s pointer
    LEA   EDI,[EDI+4]  ;advance the other array’s pointer
          LOOP ADDLOOP
-
+

(Yes, I could use LODSD instead of MOV/LEA; I’m just illustrating a point here. Besides, LODS is only 1 cycle faster than MOV/LEA on the 386, and is actually more than twice as slow on the 486.) If we used ADD rather than LEA to advance the pointers, the carry from one ADC to the next would have to be preserved with either PUSHF/POPF or LAHF/SAHF. (Alternatively, we could use multiple INCs, since INC doesn’t affect the Carry flag.)

@@ -107,37 +100,37 @@ ADDLOOP:

But there sure are a lot of interesting options, aren’t there?

-

The Kennedy Portfolio

+

The Kennedy Portfolio

Reader John Kennedy regularly passes along intriguing assembly programming tricks, many of which I’ve never seen mentioned anywhere else. John likes to optimize for size, whereas I lean more toward speed, but many of his optimizations are good for both purposes. Here are a few of my favorites:

-

John’s code for setting AX to its absolute value is:

+

John’s code for setting AX to its absolute value is:

 CWD
 XOR   AX,DX
 SUB   AX,DX
-
+ -

This does nothing when bit 15 of AX is 0 (that is, if AX is positive). When AX is negative, the code “nots” it and adds 1, which is exactly how you perform a two’s complement negate. For the case where AX is not negative, this trick usually beats the stuffing out of the standard absolute value code:

+

This does nothing when bit 15 of AX is 0 (that is, if AX is positive). When AX is negative, the code “nots” it and adds 1, which is exactly how you perform a two’s complement negate. For the case where AX is not negative, this trick usually beats the stuffing out of the standard absolute value code:

    AND   AX,AX        ;negative?
    JNS   IsPositive   ;no
    NEG   AX           ;yes,negate it
 IsPositive:
-
+

However, John’s code is slower on a 486; as you’re no doubt coming to realize (and as I’ll explain in Chapters 12 and 13), the 486 is an optimization world unto itself.

-

Here’s how John copies a block of bytes from DS:SI to ES:DI, moving as much data as possible a word at a time:

+

Here’s how John copies a block of bytes from DS:SI to ES:DI, moving as much data as possible a word at a time:

 SHR   CX,1      ;word count
 REP   MOVSW     ;copy as many words as possible
 ADC   CX,CX     ;CX=1 if copy length was odd,
                 ;0 else
 REP   MOVSB     ;copy any odd byte
-
+ -

(ADC CX,CX can be replaced with RCL CX,1; which is faster depends on the processor type.) It might be hard to believe that the above is faster than this:

+

(ADC CX,CX can be replaced with RCL CX,1; which is faster depends on the processor type.) It might be hard to believe that the above is faster than this:

    SHR   CX,1      ;word count
    REP   MOVSW     ;copy as many words as
@@ -145,7 +138,7 @@ REP   MOVSB     ;copy any odd byte
    JNC   CopyDone  ;done if even copy length
    MOVSB           ;copy the odd byte
 CopyDone:
-
+


@@ -164,10 +157,6 @@ CopyDone:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-02.html b/09-02.html index 4b722d3..7c7f550 100644 --- a/09-02.html +++ b/09-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -39,7 +32,7 @@

However, it generally is. Sure, if the length is odd, John’s approach incurs a penalty approximately equal to the REP startup time for MOVSB. However, if the length is even, John’s approach doesn’t branch, saving cycles and not emptying the prefetch queue. If copy lengths are evenly distributed between even and odd, John’s approach is faster in most x86 systems. (Not on the 486, though.)

-

John also points out that on the 386, multiple LEAs can be combined to perform multiplications that can’t be handled by a single LEA, much as multiple shifts and adds can be used for multiplication, only faster. LEA can be used to multiply in a single instruction on the 386, but only by the values 2, 3, 4, 5, 8, and 9; several LEAs strung together can handle a much wider range of values. For example, video programmers are undoubtedly familiar with the following code to multiply AX times 80 (the width in bytes of the bitmap in most PC display modes):

+

John also points out that on the 386, multiple LEAs can be combined to perform multiplications that can’t be handled by a single LEA, much as multiple shifts and adds can be used for multiplication, only faster. LEA can be used to multiply in a single instruction on the 386, but only by the values 2, 3, 4, 5, 8, and 9; several LEAs strung together can handle a much wider range of values. For example, video programmers are undoubtedly familiar with the following code to multiply AX times 80 (the width in bytes of the bitmap in most PC display modes):

 SHL   AX,1        ;*2
 SH   LAX,1        ;*4
@@ -49,31 +42,31 @@ MO   VBX,AX
 SH   LAX,1        ;*32
 SH   LAX,1        ;*64
 ADD  AX,BX        ;*80
-
+ -

Using LEA on the 386, the above could be reduced to

+

Using LEA on the 386, the above could be reduced to

 LEA   EAX,[EAX*2]     ;*2
 LEA   EAX,[EAX*8]     ;*16
 LEA   EAX,[EAX+EAX*4] ;*80
-
+ -

which still isn’t as fast as using a lookup table like

+

which still isn’t as fast as using a lookup table like

 MOV   EAX,MultiplesOf80Table[EAX*4]
-
+

but is close and takes a great deal less space.

-

Of course, on the 386, the shift and add version could also be reduced to this considerably more efficient code:

+

Of course, on the 386, the shift and add version could also be reduced to this considerably more efficient code:

 SH    LAX,4      ;*16
 MOV   BX,AX
 SHL   AX,2       ;*64
 ADD   AX,BX      ;*80
-
+ -

Speeding Up Multiplication

+

Speeding Up Multiplication

That brings us to multiplication, one of the slowest of x86 operations and one that allows for considerable optimization. One way to speed up multiplication is to use shift and add, LEA, or a lookup table to hard-code a multiplication operation for a fixed multiplier, as shown above. Another is to take advantage of the early-out feature of the 386 (and the 486, but in the interests of brevity I’ll just say “386” from now on) by arranging your operands so that the multiplier (always the rightmost operand following MUL or IMUL) is no larger than the other operand.

@@ -103,13 +96,12 @@ ADD AX,BX ;*80

That doesn’t mean that your code should test and swap operands to make sure the smaller one is the multiplier; that rarely pays off. I’m speaking more of the case where you’re scaling an array up by a value that’s always in the range of, say, 2 to 10; because the scale value will always be small and the array elements may have any value, the scale value is the logical choice for the multiplier.

-

Optimizing Optimized Searching

+

Optimizing Optimized Searching

Rob Williams writes with a wonderful optimization to the REPNZ SCASB-based optimized searching routine I discussed in Chapter 5. As a quick refresher, I described searching a buffer for a text string as follows: Scan for the first byte of the text string with REPNZ SCASB, then use REPZ CMPS to check for a full match whenever REPNZ SCASB finds a match for the first character, as shown in Figure 9.1. The principle is that most buffer characters won’t match the first character of any given string, so REPNZ SCASB, by far the fastest way to search on the PC, can be used to eliminate most potential matches; each remaining potential match can then be checked in its entirety with REPZ CMPS.

-


- Figure 9.1
  Simple searching method for locating a text string.

+


+ Figure 9.1
  Simple searching method for locating a text string.

Rob’s revelation, which he credits without explanation to Edgar Allen Poe (search nevermore?), was that by far the slowest part of the whole deal is handling REPNZ SCASB matches, which require checking the remainder of the string with REPZ CMPS and restarting REPNZ SCASB if no match is found.

@@ -140,10 +132,6 @@ ADD AX,BX ;*80
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-03.html b/09-03.html index b891d4b..6c7fb6d 100644 --- a/09-03.html +++ b/09-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -49,11 +42,10 @@ -


- Figure 9.2
  Faster searching method for locating a text string.

+


+ Figure 9.2
  Faster searching method for locating a text string.

-

LISTING 9.1 L9-1.ASM

+

LISTING 9.1 L9-1.ASM

 ; Searches a text buffer for a text string. Uses REPNZ SCASB to sca"n
 ; the buffer for locations that match the first character of the
@@ -146,7 +138,7 @@ FindStringDone:
 ret
 _FindStringendp
       end
-
+


@@ -165,10 +157,6 @@ _FindStringendp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-04.html b/09-04.html index 41a7855..80ba8d9 100644 --- a/09-04.html +++ b/09-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -37,7 +30,7 @@


-

LISTING 9.2 L9-2.ASM

+

LISTING 9.2 L9-2.ASM

 ; Searches a text buffer for a text string. Uses REPNZ SCASB to scan
 ; the buffer for locations that match a specified character of the
@@ -135,9 +128,9 @@ FindStringDone:
       ret
 _FindStringendp
       end
-
+ -

LISTING 9.3 L9-3.C

+

LISTING 9.3 L9-3.C

 /* Program to exercise buffer-search routines in Listings 9.1 & 9.2 */
 #include <stdio.h>
@@ -171,7 +164,7 @@ void main() {
             strncpy(TempBuffer, MatchPtr, DISPLAY_LENGTH));
    }
 }
-
+


@@ -190,10 +183,6 @@ void main() {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-05.html b/09-05.html index b8a0025..7497267 100644 --- a/09-05.html +++ b/09-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -41,11 +34,11 @@

The point is that you can improve performance dramatically by understanding the nature of the data with which you work. (This is equally true for high-level language programming, by the way.) Listing 9.2 is very similar to and only slightly more complex than Listing 9.1; the difference lies not in elbow grease or cycle counting but in the organic integrating optimizer technology we all carry around in our heads.

-

Short Sorts

+

Short Sorts

David Stafford (recently of Borland and Borland Japan) who happens to be one of the best assembly language programmers I’ve ever met, has written a C-callable routine that sorts an array of integers in ascending order. That wouldn’t be particularly noteworthy, except that David’s routine, shown in Listing 9.4, is exactly 25 bytes long. Look at the code; you’ll keep saying to yourself, “But this doesn’t work...oh, yes, I guess it does.” As they say in the Prego spaghetti sauce ads, it’s in there—and what a job of packing. Anyway, David says that a 24-byte sort routine eludes him, and he’d like to know if anyone can come up with one.

-

LISTING 9.4 L9-4.ASM

+

LISTING 9.4 L9-4.ASM

 .
 ;--------------------------------------------------------------------------
@@ -79,9 +72,9 @@ _sort:  pop     dx              ;get return address (entry point)
         ret
 
       end
-
+ -

Full 32-Bit Division

+

Full 32-Bit Division

One of the most annoying limitations of the x86 is that while the dividend operand to the DIV instruction can be 32 bits in size, both the divisor and the result must be 16 bits. That’s particularly annoying in regards to the result because sometimes you just don’t know whether the ratio of the dividend to the divisor is greater than 64K-1 or not—and if you guess wrong, you get that godawful Divide By Zero interrupt. So, what is one to do when the result might not fit in 16 bits, or when the dividend is larger than 32 bits? Fall back to a software division approach? That will work—but oh so slowly.

@@ -89,9 +82,8 @@ _sort: pop dx ;get return address (entry point)

This technique involves nothing more complicated than breaking up the division into word-sized chunks, starting with the most significant word of the dividend. The most significant word is divided by the divisor (with no chance of overflow because there are only 16 bits in each); then the remainder is prepended to the next 16 bits of dividend, and the process is repeated, as shown in Figure 9.3. This process is equivalent to dividing by hand, except that here we stop to carry the remainder manually only after each word of the dividend; the hardware divide takes care of the rest. Listing 9.5 shows a function to divide an arbitrarily large dividend by a 16-bit divisor, and Listing 9.6 shows a sample division of a large dividend. Note that the same principle can be applied to handling arbitrarily large dividends in 386 native mode code, but in that case the operation can proceed a dword, rather than a word, at a time.

-


- Figure 9.3
  Fast multiword division on the 386.

+


+ Figure 9.3
  Fast multiword division on the 386.

As for handling signed division with arbitrarily large dividends, that can be done easily enough by remembering the signs of the dividend and divisor, dividing the absolute value of the dividend by the absolute value of the divisor, and applying the stored signs to set the proper signs for the quotient and remainder. There may be more clever ways to produce the same result, by using IDIV, for example; if you know of one, drop me a line c/o Coriolis Group Books.

@@ -112,10 +104,6 @@ _sort: pop dx ;get return address (entry point)
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-06.html b/09-06.html index 2fbb2bd..4eb5c48 100644 --- a/09-06.html +++ b/09-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -37,7 +30,7 @@


-

LISTING 9.5 L9-5.ASM

+

LISTING 9.5 L9-5.ASM

 ; Divides an arbitrarily long unsigned dividend by a 16-bit unsigned
 ; divisor. C near-callable as:
@@ -105,9 +98,9 @@ DivLoop:
                ret
 _Divendp
                end
-
+ -

LISTING 9.6 L9-6.C

+

LISTING 9.6 L9-6.C

 /* Sample use of Div function to perform division when the result
    doesn’t fit in 16 bits */
@@ -125,9 +118,9 @@ main() {
    k = Div((unsigned int *)&i, sizeof(i), j, (unsigned int *)&m);
    printf(“%lu / %u = %lu r %u\n”, i, j, m, k);
 }
-
+ -

Sweet Spot Revisited

+

Sweet Spot Revisited

Way back in Volume 1, Number 1 of PC TECHNIQUES, (April/May 1990) I wrote the very first of that magazine’s HAX (#1), which extolled the virtues of placing your most commonly-used automatic (stack-based) variables within the stack’s “sweet spot,” the area between +127 to -128 bytes away from BP, the stack frame pointer. The reason was that the 8088 can store addressing displacements that fall within that range in a single byte; larger displacements require a full word of storage, increasing code size by a byte per instruction, and thereby slowing down performance due to increased instruction fetching time.

@@ -162,10 +155,6 @@ main() {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/09-07.html b/09-07.html index 0d01066..b0a7171 100644 --- a/09-07.html +++ b/09-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Hints My Readers Gave Me - - @@ -37,30 +30,29 @@


-

Hard-Core Cycle Counting

+

Hard-Core Cycle Counting

Next, we come to an item that cycle counters will love, especially since it involves apparently incorrect documentation on Intel’s part. According to Intel’s documents, all RCR and RCL instructions, which perform rotations through the Carry flag, as shown in Figure 9.4, take 9 cycles on the 386 when working with a register operand. My measurements indicate that the 9-cycle execution time almost holds true for multibit rotate-through-carries, which I’ve timed at 8 cycles apiece; for example, RCR AX,CL takes 8 cycles on my 386, as does RCL DX,2. Contrast that with ROR and ROL, which can rotate the contents of a register any number of bits in just 3 cycles.

However, rotating by one bit through the Carry flag does not take 9 cycles, contrary to Intel’s 80386 Programmer’s Reference Manual, or even 8 cycles. In fact, RCR reg,1 and RCL reg,1 take 3 cycles, just like ROR, ROL, SHR, and SHL. At least, that’s how fast they run on my 386, and I very much doubt that you’ll find different execution times on other 386s. (Please let me know if you do, though!)

-


- Figure 9.4
  Performing rotate instructions using the Carry flag.

+


+ Figure 9.4
  Performing rotate instructions using the Carry flag.

Interestingly, according to Intel’s i486 Microprocessor Programmer’s Reference Manual, the 486 can RCR or RCL a register by one bit in 3 cycles, but takes between 8 and 30 cycles to perform a multibit register RCR or RCL!

No great lesson here, just a caution to be leery of multibit RCR and RCL when performance matters—and to take cycle-time documentation with a grain of salt.

-

Hardwired Far Jumps

+

Hardwired Far Jumps

-

Did you ever wonder how to code a far jump to an absolute address in assembly language? Probably not, but if you ever do, you’re going to be glad for this next item, because the obvious solution doesn’t work. You might think all it would take to jump to, say, 1000:5 would be JMP FAR PTR 1000:5, but you’d be wrong. That won’t even assemble. You might then think to construct in memory a far pointer containing 1000:5, as in the following:

+

Did you ever wonder how to code a far jump to an absolute address in assembly language? Probably not, but if you ever do, you’re going to be glad for this next item, because the obvious solution doesn’t work. You might think all it would take to jump to, say, 1000:5 would be JMP FAR PTR 1000:5, but you’d be wrong. That won’t even assemble. You might then think to construct in memory a far pointer containing 1000:5, as in the following:

 Ptr  dd   ?
      :
      mov  word ptr [Ptr],5
      mov  word ptr [Ptr+2],1000h
      jmp  [Ptr]
-
+

That will work, but at a price in performance. On an 8088, JMP DWORD PTR [mem] (an indirect far jump) takes at least 37 cycles; JMP DWORD PTR label (a direct far jump) takes only 15 cycles (plus, almost certainly, some cycles for instruction fetching). On a 386, an indirect far jump is documented to take at least 43 cycles in real mode (31 in protected mode); a direct far jump is documented to take at least 12 cycles, about three times faster. In truth, the difference between those two is nowhere near that big; the fastest I’ve measured for a direct far jump is 21 cycles, and I’ve measured indirect far jumps as fast as 30 cycles, so direct is still faster, but not by so much. (Oh, those cycle-time documentation blues!) Also, a direct far jump is documented to take at least 27 cycles in protected mode; why the big difference in protected mode, I have no idea.

@@ -68,7 +60,7 @@ Ptr dd ?

Listing 9.7 shows a short program that performs a direct far call to 1000:5. (Don’t run it, unless you want to crash your system!) It does this by creating a dummy segment at 1000H, so that the label FarLabel can be created with the desired far attribute at the proper location. (Segments created with “AT” don’t cause the generation of any actual bytes or the allocation of any memory; they’re just templates.) It’s a little kludgey, but at least it does work. There may be a better solution; if you have one, pass it along.

-

LISTING 9.7 L9-7.ASM

+

LISTING 9.7 L9-7.ASM

 ; Program to perform a direct far jump to address 1000:5.
 ; *** Do not run this program! It’s just an example of how ***
@@ -86,34 +78,34 @@ FarSeg      ends
 start:
       jmp     FarLabel
       end     start
-
+

By the way, if you’re wondering how I figured this out, I merely applied my good friend Dan Illowsky’s long-standing rule for dealing with MASM:

If the obvious doesn’t work (and it usually doesn’t), just try everything you can think of, no matter how ridiculous, until you find something that does—a rule with plenty of history on its side.

-

Setting 32-Bit Registers: Time versus Space

+

Setting 32-Bit Registers: Time versus Space

-

To finish up this chapter, consider these two items. First, in 32-bit protected mode,

+

To finish up this chapter, consider these two items. First, in 32-bit protected mode,

 sub  eax,eax
 inc  eax
-
+ -

takes 4 cycles to execute, but is only 3 bytes long, while

+

takes 4 cycles to execute, but is only 3 bytes long, while

 mov  eax,1
-
+ -

takes only 2 cycles to execute, but is 5 bytes long (because native mode constants are dwords and the MOV instruction doesn’t sign-extend). Both code fragments are ways to set EAX to 1 (although the first affects the flags and the second doesn’t); this is a classic trade-off of speed for space. Second,

+

takes only 2 cycles to execute, but is 5 bytes long (because native mode constants are dwords and the MOV instruction doesn’t sign-extend). Both code fragments are ways to set EAX to 1 (although the first affects the flags and the second doesn’t); this is a classic trade-off of speed for space. Second,

 or    ebx,-1
-
+ -

takes 2 cycles to execute and is 3 bytes long, while

+

takes 2 cycles to execute and is 3 bytes long, while

 move  bx,-1
-
+

takes 2 cycles to execute and is 5 bytes long. Both instructions set EBX to -1; this is a classic trade-off of—gee, it’s not a trade-off at all, is it? OR is a better way to set a 32-bit register to all 1-bits, just as SUB or XOR is a better way to set a register to all 0-bits. Who woulda thunk it? Just goes to show how the 32-bit displacements and constants of 386 native mode change the familiar landscape of 80x86 optimization.

@@ -136,10 +128,6 @@ move bx,-1
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/10-01.html b/10-01.html index 23cc904..91b8f25 100644 --- a/10-01.html +++ b/10-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - @@ -37,10 +30,10 @@


-

Chapter 10
+

Chapter 10
Patient Coding, Faster Code

-

How Working Quickly Can Bring Execution to a Crawl

+

How Working Quickly Can Bring Execution to a Crawl

My grandfather does The New York Times crossword puzzle every Sunday. In ink. With nary a blemish.

@@ -64,7 +57,7 @@

In this chapter, I’m going to walk you through a simple but illustrative case history that nicely points up the wisdom of delaying gratification when faced with programming problems, so that your mind has time to chew on the problems from other angles. The alternative solutions you find by doing this may seem obvious, once you’ve come up with them. They may not even differ greatly from your initial solutions. Often, however, they will be much better—and you’ll never even have the chance to decide whether they’re better or not if you take the first thing that comes into your head and run with it.

-

The Case for Delayed Gratification

+

The Case for Delayed Gratification

Once upon a time, I set out to read Algorithms, by Robert Sedgewick (Addison-Wesley), which turned out to be a wonderful, stimulating, and most useful book, one that I recommend highly. My story, however, involves only what happened in the first 12 pages, for it was in those pages that Sedgewick discussed Euclid’s algorithm.

@@ -72,7 +65,7 @@

The problem at hand, then, is simply this: Find the largest integer value that evenly divides two arbitrary positive integers. That’s all there is to it. So warm up your pattern matchers...and go!

-

The Brute-Force Syndrome

+

The Brute-Force Syndrome

I have a funny feeling that you’d already figured out how to find the GCD before I even said “go.” That’s what I did when reading Algorithms; before I read another word, I had to figure it out for myself. Programmers are like that; give them a problem and their eyes immediately glaze over as they try to solve it before you’ve even shut your mouth. That sort of instant response can certainly be impressive, but it can backfire, too, as it did in my case.

@@ -95,10 +88,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/10-02.html b/10-02.html index 69dd59d..7044396 100644 --- a/10-02.html +++ b/10-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - @@ -43,9 +36,8 @@

Listing 10.1 is an implementation of the brute-force approach to GCD calculation. Table 10.1 shows how long it takes this approach to find the GCD for several integer pairs. As expected, performance is extremely poor when iS is large.

-


- Figure 10.1
  Using a brute-force algorithm to find a GCD.

+


+ Figure 10.1
  Using a brute-force algorithm to find a GCD.

@@ -222,7 +214,7 @@
-

LISTING 10.1 L10-1.C

+

LISTING 10.1 L10-1.C

 /* Finds and returns the greatest common divisor of two positive
    integers. Works by trying every integral divisor between the
@@ -246,13 +238,13 @@ unsigned int gcd(unsigned int int1, unsigned int int2) {
       ;
    return(trial_divisor);
 }
-
+ -

Wasted Breakthroughs

+

Wasted Breakthroughs

Sedgewick’s first solution to the GCD problem was pretty much the one I came up with. He then pointed out that the GCD of iL and iS is the same as the GCD of iL-iS and iS. This was obvious (once Sedgewick pointed it out); by the very nature of division, any number that divides iL evenly nL times and iS evenly nS times must divide iL-iS evenly nL-nS times. Given that insight, I immediately designed a new, faster approach, shown in Listing 10.2.

-

LISTING 10.2 L10-2.C

+

LISTING 10.2 L10-2.C

 /* Finds and returns the greatest common divisor of two positive
    integers. Works by subtracting the smaller integer from the
@@ -283,7 +275,7 @@ unsigned int gcd(unsigned int int1, unsigned int int2) {
    /* Now recursively call this function to continue the process */
    return(gcd(int1, int2));
 }
-
+


@@ -302,10 +294,6 @@ unsigned int gcd(unsigned int int1, unsigned int int2) {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/10-03.html b/10-03.html index 9fda6d0..6bf750b 100644 --- a/10-03.html +++ b/10-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - @@ -39,15 +32,14 @@

Listing 10.2 repeatedly subtracts iS from iL until iL becomes less than or equal to iS. If iL becomes equal to iS, then that’s the GCD; alternatively, if iL becomes less than iS, iL and iS switch values, and the process is repeated, as shown in Figure 10.2. The number of iterations this approach requires relative to Listing 10.1 depends heavily on the values of iL and iS, so it’s not always faster, but, as Table 10.1 indicates, Listing 10.2 is generally much better code.

-


- Figure 10.2
  Using repeated subtraction algorithm to find a GCD.

+


+ Figure 10.2
  Using repeated subtraction algorithm to find a GCD.

Listing 10.2 is a far graver misstep than Listing 10.1, for all that it’s faster. Listing 10.1 is obviously a hacked-up, brute-force approach; no one could mistake it for anything else. It could be speeded up in any of a number of ways with a little thought. (Simply skipping testing all the divisors between iS and iS/2, not inclusive, would cut the worst-case time in half, for example; that’s not a particularly good optimization, but it illustrates how easily Listing 10.1 can be improved.) Listing 10.1 is a hack job, crying out for inspiration.

Listing 10.2, on the other hand, has gotten the inspiration—and largely wasted it through haste. Had Sedgewick not told me otherwise, I might well have assumed that Listing 10.2 was optimized, a mistake I would never have made with Listing 10.1. I experienced a conceptual breakthrough when I understood Sedgewick’s point: A smaller number can be subtracted from a larger number without affecting their GCD, thereby inexpensively reducing the scale of the problem. And, in my hurry to make this breakthrough reality, I missed its full scope. As Sedgewick says on the very next page, the number that one gets by subtracting iS from iL until iL is less than iS is precisely the same as the remainder that one gets by dividing iL by iS—again, this is inherent in the nature of division—and that is the basis for Euclid’s algorithm, shown in Figure 10.3. Listing 10.3 is an implementation of Euclid’s algorithm.

-

LISTING 10.3 L10-3.C

+

LISTING 10.3 L10-3.C

 /* Finds and returns the greatest common divisor of two integers.
    Uses Euclid’s algorithm: divides the larger integer by the
@@ -92,7 +84,7 @@ static unsigned int gcd_recurs(unsigned int larger_int,
       continue the process */
    return(gcd_recurs(smaller_int, temp));
 }
-
+

As you can see from Table 10.1, Euclid’s algorithm is superior, especially for large numbers (and imagine if we were working with large longs!).

@@ -104,17 +96,16 @@ static unsigned int gcd_recurs(unsigned int larger_int, -


- Figure 10.3
  Using Euclid’s algorithm to find a GCD.

+


+ Figure 10.3
  Using Euclid’s algorithm to find a GCD.

Give your mind time and space to wander around the edges of important programming problems before you settle on any one approach. I titled this book’s first chapter “The Best Optimizer Is between Your Ears,” and that’s still true; what’s even more true is that the optimizer between your ears does its best work not at the implementation stage, but at the very beginning, when you try to imagine how what you want to do and what a computer is capable of doing can best be brought together.

-

Recursion

+

Recursion

Euclid’s algorithm lends itself to recursion beautifully, so much so that an implementation like Listing 10.3 comes almost without thought. Again, though, take a moment to stop and consider what’s really going on, at the assembly language level, in Listing 10.3. There’s recursion and then there’s recursion; code recursion and data recursion, to be exact. Listing 10.3 is code recursion—recursion through calls—the sort most often used because it is conceptually simplest. However, code recursion tends to be slow because it pushes parameters and calls a subroutine for every iteration. Listing 10.4, which uses data recursion, is much faster and no more complicated than Listing 10.3. Actually, you could just say that Listing 10.4 uses a loop and ignore any mention of recursion; conceptually, though, Listing 10.4 performs the same recursive operations that Listing 10.3 does.

-

LISTING 10.4 L10-4.C

+

LISTING 10.4 L10-4.C

 /* Finds and returns the greatest common divisor of two integers.
    Uses Euclid’s algorithm: divides the larger integer by the
@@ -148,9 +139,9 @@ unsigned int gcd(unsigned int int1, unsigned int int2) {
       int2 = temp;
    }
 }
-
+ -

Patient Optimization

+

Patient Optimization

At long last, we’re ready to optimize GCD determination in the classic sense. Table 10.1 shows the performance of Listing 10.4 with and without Microsoft C/C++’s maximum optimization, and also shows the performance of Listing 10.5, an assembly language version of Listing 10.4. Sure, the optimized versions are faster than the unoptimized version of Listing 10.4—but the gains are small compared to those realized from the higher-level optimizations in Listings 10.2 through 10.4.

@@ -171,10 +162,6 @@ unsigned int gcd(unsigned int int1, unsigned int int2) {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/10-04.html b/10-04.html index 9083327..1a349d5 100644 --- a/10-04.html +++ b/10-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Patient Coding, Faster Code - - @@ -37,7 +30,7 @@


-

LISTING 10.5 L10-5.ASM

+

LISTING 10.5 L10-5.ASM

 ; Finds and returns the greatest common divisor of two integers.
 ; Uses Euclid’s algorithm: divides the larger integer by the
@@ -126,7 +119,7 @@ Done:
       ret
 _gcd  endp
       end
-
+

Assembly language optimization is pattern matching on a local scale. Frankly, it’s also the sort of boring, brute-force work that people are lousy at; compilers could out-optimize you at this level with one pass tied behind their back if they knew as much about the code you’re writing as you do, which they don’t.

@@ -163,10 +156,6 @@ _gcd endp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-01.html b/11-01.html index a27de4d..c41abef 100644 --- a/11-01.html +++ b/11-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -37,16 +30,16 @@


-

Chapter 11
+

Chapter 11
Pushing the 286 and 386

-

New Registers, New Instructions, New Timings, New Complications

+

New Registers, New Instructions, New Timings, New Complications

This chapter, adapted from my earlier book Zen of Assembly Language (1989; now out of print), provides an overview of the 286 and 386, often contrasting those processors with the 8088. At the time I originally wrote this, the 8088 was the king of processors, and the 286 and 386 were the new kids on the block. Today, of course, all three processors are past their primes, but many millions of each are still in use, and the 386 in particular is still well worth considering when optimizing software.

This chapter provides an interesting look at the evolution of the x86 architecture, to a greater degree than you might expect, for the x86 family came into full maturity with the 386; the 486 and the Pentium are really nothing more than faster 386s, with very little in the way of new functionality. In contrast, the 286 added a number of instructions, respectable performance, and protected mode to the 8088’s capabilities, and the 386 added more instructions and a whole new set of addressing modes, and brought the x86 family into the 32-bit world that represents the future (and, increasingly, the present) of personal computing. This chapter also provides insight into the effects on optimization of the variations in processors and memory architectures that are common in the PC world. So, although the 286 and 386 no longer represent the mainstream of computing, this chapter is a useful mix of history lesson, x86 overview, and details on two workhorse processors that are still in wide use.

-

Family Matters

+

Family Matters

While the x86 family is a large one, only a few members of the family—which includes the 8088, 8086, 80188, 80186, 286, 386SX, 386DX, numerous permutations of the 486, and now the Pentium—really matter.

@@ -60,21 +53,21 @@

This leaves us with just two processors: the 286 and the 386. Each was the PC standard in its day. The 286 is no longer used in new systems, but there are millions of 286-based systems still in daily use. The 386 is still being used in new systems, although it’s on the downhill leg of its lifespan, and it is in even wider use than the 286. The future clearly belongs to the 486 and Pentium, but the 286 and 386 are still very much a part of the present-day landscape.

-

Crossing the Gulf to the 286 and the 386

+

Crossing the Gulf to the 286 and the 386

Apart from vastly improved performance, the biggest difference between the 8088 and the 286 and 386 (as well as the later Intel CPUs) is that the 286 introduced protected mode, and the 386 greatly expanded the capabilities of protected mode. We’re only going to talk about real-mode operation of the 286 and 386 in this book, however. Protected mode offers a whole new memory management scheme, one that isn’t supported by the 8088. Only code specifically written for protected mode can run in that mode; it’s an alien and hostile environment for MS-DOS programs.

-

In particular, segments are different creatures in protected mode. They’re selectors—indexes into a table of segment descriptors—rather than plain old registers, and can’t be set to arbitrary values. That means that segments can’t be used for temporary storage or as part of a fast indivisible 32-bit load from memory, as in

+

In particular, segments are different creatures in protected mode. They’re selectors—indexes into a table of segment descriptors—rather than plain old registers, and can’t be set to arbitrary values. That means that segments can’t be used for temporary storage or as part of a fast indivisible 32-bit load from memory, as in

 les  ax,dword ptr [LongVar]
 mov  dx,es
-
+ -

which loads LongVar into DX:AX faster than this:

+

which loads LongVar into DX:AX faster than this:

 mov  ax,word ptr [LongVar]
 mov  dx,word ptr [LongVar+2]
-
+

Protected mode uses those altered segment registers to offer access to a great deal more memory than real mode: The 286 supports 16 megabytes of memory, while the 386 supports 4 gigabytes (4K megabytes) of physical memory and 64 terabytes (64K gigabytes!) of virtual memory.

@@ -82,7 +75,7 @@ mov dx,word ptr [LongVar+2]

In short, taken as a whole, protected mode programming is a different kettle of fish altogether from what I’ve been describing in this book. There’s certainly a knack to optimizing specifically for protected mode under a given operating system...but it’s not what we’ve been learning, and now is not the time to pursue it further. In general, though, the optimization strategies discussed in this book still hold true in protected mode; it’s just issues specific to protected mode or a particular operating system that we won’t discuss.

-

In the Lair of the Cycle-Eaters, Part II

+

In the Lair of the Cycle-Eaters, Part II

Under the programming interface, the 286 and 386 differ considerably from the 8088. Nonetheless, with one exception and one addition, the cycle-eaters remain much the same on computers built around the 286 and 386. Next, we’ll review each of the familiar cycle-eaters I covered in Chapter 4 as they apply to the 286 and 386, and we’ll look at the new member of the gang, the data alignment cycle-eater.

@@ -105,10 +98,6 @@ mov dx,word ptr [LongVar+2]
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-02.html b/11-02.html index f047239..d7eb15e 100644 --- a/11-02.html +++ b/11-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -47,7 +40,7 @@

The most significant reason that the prefetch queue cycle-eater not only survives but prospers on the 286 and 386, however, lies in the various memory architectures used in computers built around the 286 and 386. Due to the memory architectures, the 8-bit bus cycle-eater is replaced by a new form of the wait state cycle-eater: wait states on accesses to normal system memory.

-

System Wait States

+

System Wait States

The 286 and 386 were designed to lose relatively little performance to the prefetch queue cycle-eater...when used with zero-wait-state memory: memory that can complete memory accesses so rapidly that no wait states are needed. However, true zero-wait-state memory is almost never used with those processors. Why? Because memory that can keep up with a 286 is fairly expensive, and memory that can keep up with a 386 is very expensive. Instead, computer designers use alternative memory architectures that offer more performance for the dollar—but less performance overall—than zero-wait-state memory. (It is possible to build zero-wait-state systems for the 286 and 386; it’s just so expensive that it’s rarely done.)

@@ -71,7 +64,7 @@

Let’s check out the prefetch queue cycle-eater in action. Listing 11.1 times MOV [WordVar],0. The Zen timer reports that on a one-wait-state 10 MHz 286-based AT clone (the computer used for all tests in this chapter), Listing 11.1 runs in 1.27 µs per instruction. That’s 12.7 cycles per instruction, just as we calculated. (That extra seven-tenths of a cycle comes from DRAM refresh, which we’ll get to shortly.)

-

LISTING 11.1 L11-1.ASM

+

LISTING 11.1 L11-1.ASM

 ;
 ; *** Listing 11.1 ***
@@ -92,7 +85,7 @@ Skip:
         mov     [WordVar],0
         endm
         call    ZTimerOff
-
+

What does this mean? It means that, practically speaking, the 286 as used in the AT doesn’t have a 16-bit bus. From a performance perspective, the 286 in an AT has two-thirds of a 16-bit bus (a 10.7-bit bus?), since every bus access on an AT takes 50 percent longer than it should. A 286 running at 10 MHz should be able to access memory at a maximum rate of 1 word every 200 ns; in a 10 MHz AT, however, that rate is reduced to 1 word every 300 ns by the one-wait-state memory.

@@ -113,10 +106,6 @@ Skip:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-03.html b/11-03.html index a79bace..c6eefaf 100644 --- a/11-03.html +++ b/11-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -65,7 +58,7 @@

Of course, those are exactly the rules that apply to 8088 optimization as well. Isn’t it convenient that the same general rules apply across the board?

-

Data Alignment

+

Data Alignment

Thanks to its 16-bit bus, the 286 can access word-sized memory variables just as fast as byte-sized variables. There’s a catch, however: That’s only true for word-sized variables that start at even addresses. When the 286 is asked to perform a word-sized access starting at an odd address, it actually performs two separate accesses, each of which fetches 1 byte, just as the 8088 does for all word-sized accesses.

@@ -83,15 +76,14 @@

That, in a nutshell, is the data alignment cycle-eater, the one new cycle-eater of the 286 and 386. (The data alignment cycle-eater is a close relative of the 8088’s 8-bit bus cycle-eater, but since it behaves differently—occurring only at odd addresses—and is avoided with a different workaround, we’ll consider it to be a new cycle-eater.)

-


- Figure 11.1
  The data alignment cycle-eater.

+


+ Figure 11.1
  The data alignment cycle-eater.

The way to deal with the data alignment cycle-eater is straightforward: Don’t perform word-sized accesses to odd addresses on the 286 if you can help it. The easiest way to avoid the data alignment cycle-eater is to place the directive EVEN before each of your word-sized variables. EVEN forces the offset of the next byte assembled to be even by inserting a NOP if the current offset is odd; consequently, you can ensure that any word-sized variable can be accessed efficiently by the 286 simply by preceding it with EVEN.

Listing 11.2, which accesses memory a word at a time with each word starting at an odd address, runs on a 10 MHz AT clone in 1.27 ms per repetition of MOVSW, or 0.64 ms per word-sized memory access. That’s 6-plus cycles per word-sized access, which breaks down to two separate memory accesses—3 cycles to access the high byte of each word and 3 cycles to access the low byte of each word, the inevitable result of non-word-aligned word-sized memory accesses—plus a bit extra for DRAM refresh.

-

LISTING 11.2 L11-2.ASM

+

LISTING 11.2 L11-2.ASM

 ;
 ; *** Listing 11.2 ***
@@ -110,11 +102,11 @@ Skip:
         call    ZTimerOn
         rep     movsw
         call    ZTimerOff
-
+

On the other hand, Listing 11.3, which is exactly the same as Listing 11.2 save that the memory accesses are word-aligned (start at even addresses), runs in 0.64 ms per repetition of MOVSW, or 0.32 µs per word-sized memory access. That’s 3 cycles per word-sized access—exactly twice as fast as the non-word-aligned accesses of Listing 11.2, just as we predicted.

-

LISTING 11.3 L11-3.ASM

+

LISTING 11.3 L11-3.ASM

 ;
 ; *** Listing 11.3 ***
@@ -132,11 +124,11 @@ Skip:
         call    ZTimerOn
         rep     movsw
         call    ZTimerOff
-
+

The data alignment cycle-eater has intriguing implications for speeding up 286/386 code. The expenditure of a little care and a few bytes to make sure that word-sized variables and memory blocks are word-aligned can literally double the performance of certain code running on the 286. Even if it doesn’t double performance, word alignment usually helps and never hurts.

-

Code Alignment

+

Code Alignment

Lack of word alignment can also interfere with instruction fetching on the 286, although not to the extent that it interferes with access to word-sized memory variables. The 286 prefetches instructions a word at a time; even if a given instruction doesn’t begin at an even address, the 286 simply fetches the first byte of that instruction at the same time that it fetches the last byte of the previous instruction, as shown in Figure 11.2, then separates the bytes internally. That means that in most cases, instructions run just as fast whether they’re word-aligned or not.

@@ -159,10 +151,6 @@ Skip:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-04.html b/11-04.html index 2921472..2df2e9f 100644 --- a/11-04.html +++ b/11-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -37,7 +30,7 @@


-

When I was developing the Zen timer, I used my trusty 10 MHz 286-based AT clone to verify the basic functionality of the timer by measuring the performance of simple instruction sequences. I was cruising along with no problems until I timed the following code:

+

When I was developing the Zen timer, I used my trusty 10 MHz 286-based AT clone to verify the basic functionality of the timer by measuring the performance of simple instruction sequences. I was cruising along with no problems until I timed the following code:

 
     mov    cx,1000
@@ -45,19 +38,17 @@
 LoopTop:
     loop   LoopTop
     call   ZTimerOff
-
+ -


- Figure 11.2
  Word-aligned prefetching on the 286.

+


+ Figure 11.2
  Word-aligned prefetching on the 286.

-


- Figure 11.3
  How instruction bytes are fetched after a branch.

+


+ Figure 11.3
  How instruction bytes are fetched after a branch.

Now, this code should run in, say, about 12 cycles per loop at most. Instead, it took over 14 cycles per loop, an execution time that I could not explain in any way. After rolling it around in my head for a while, I took a look at the code under a debugger...and the answer leaped out at me. The loop began at an odd address! That meant that two instruction fetches were required each time through the loop; one to get the opcode byte of the LOOP instruction, which resided at the end of one word-aligned word, and another to get the displacement byte, which resided at the start of the next word-aligned word.

-

One simple change brought the execution time down to a reasonable 12.5 cycles per loop:

+

One simple change brought the execution time down to a reasonable 12.5 cycles per loop:

   mov   cx,1000
   call  ZTimerOn
@@ -65,7 +56,7 @@ LoopTop:
 LoopTop:
   loop  LoopTop
   call  ZTimerOff
-
+

While word-aligning branch destinations can improve branching performance, it’s a nuisance and can increase code size a good deal, so it’s not worth doing in most code. Besides, EVEN inserts a NOP instruction if necessary, and the time required to execute a NOP can sometimes cancel the performance advantage of having a word-aligned branch destination.

@@ -77,16 +68,16 @@ LoopTop: -

I recommend that you only go out of your way to word-align the start offsets of your subroutines, as in:

+

I recommend that you only go out of your way to word-align the start offsets of your subroutines, as in:

           even
 FindChar  proc near
           :
-
+

In my experience, this simple practice is the one form of code alignment that consistently provides a reasonable return for bytes and effort expended, although sometimes it also pays to word-align tight time-critical loops.

-

Alignment and the 386

+

Alignment and the 386

So far we’ve only discussed alignment as it pertains to the 286. What, you may well ask, of the 386?

@@ -94,7 +85,7 @@ FindChar proc near

As for code alignment...the subroutine-start word-alignment rule of the 286 serves reasonably well there too since it avoids the worst case, where just 1 byte is fetched on entry to a subroutine. While optimum performance would dictate doubleword alignment of subroutines, that takes 3 bytes, a high price to pay for an optimization that improves performance only on the post 286 processors.

-

Alignment and the Stack

+

Alignment and the Stack

One side-effect of the data alignment cycle-eater of the 286 and 386 is that you should never allow the stack pointer to become odd. (You can make the stack pointer odd by adding an odd value to it or subtracting an odd value from it, or by loading it with an odd value.) An odd stack pointer on the 286 or 386 (or a non-doubleword-aligned stack in 32-bit protected mode on the 386, 486, or Pentium) will significantly reduce the performance of PUSH, POP, CALL, and RET, as well as INT and IRET, which are executed to invoke DOS and BIOS functions, handle keystrokes and incoming serial characters, and manage the mouse. I know of a Forth programmer who vastly improved the performance of a complex application on the AT simply by forcing the Forth interpreter to maintain an even stack pointer at all times.

@@ -108,7 +99,7 @@ FindChar proc near -

The DRAM Refresh Cycle-Eater: Still an Act of God

+

The DRAM Refresh Cycle-Eater: Still an Act of God

The DRAM refresh cycle-eater is the cycle-eater that’s least changed from its 8088 form on the 286 and 386. In the AT, DRAM refresh uses a little over five percent of all available memory accesses, slightly less than it uses in the PC, but in the same ballpark. While the DRAM refresh penalty varies somewhat on various AT clones and 386 computers (in fact, a few computers are built around static RAM, which requires no refresh at all; likewise, caches are made of static RAM so cached systems generally suffer less from DRAM refresh), the 5 percent figure is a good rule of thumb.

@@ -116,7 +107,7 @@ FindChar proc near

There’s nothing much new with DRAM refresh on 286/386 computers, then. Be aware of it, but don’t overly concern yourself—DRAM refresh is still an act of God, and there’s not a blessed thing you can do about it. Happily, the internal caches of the 486 and Pentium make DRAM refresh largely a performance non-issue on those processors.

-

The Display Adapter Cycle-Eater

+

The Display Adapter Cycle-Eater

Finally we come to the last of the cycle-eaters, the display adapter cycle-eater. There are two ways of looking at this cycle-eater on 286/386 computers: (1) It’s much worse than it was on the PC, or (2) it’s just about the same as it was on the PC.

@@ -141,10 +132,6 @@ FindChar proc near
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-05.html b/11-05.html index f1b605f..1b66ca3 100644 --- a/11-05.html +++ b/11-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -61,7 +54,7 @@

What can we do about this new, more virulent form of the display adapter cycle-eater? The workaround is the same as it was on the PC: Access display memory as little as you possibly can.

-

New Instructions and Features: The 286

+

New Instructions and Features: The 286

The 286 and 386 offer a number of new instructions. The 286 has a relatively small number of instructions that the 8088 lacks, while the 386 has those instructions and quite a few more, along with new addressing modes and data sizes. We’ll discuss the 286 and the 386 separately in this regard.

@@ -71,7 +64,7 @@

A couple of old instructions gain new features on the 286. For one, the 286 version of PUSH is capable of pushing a constant on the stack. For another, the 286 allows all shifts and rotates to be performed for not just 1 bit or the number of bits specified by CL, but for any constant number of bits.

-

New Instructions and Features: The 386

+

New Instructions and Features: The 386

The 386 is somewhat more complex than the 286 regarding new features. Once again, we won’t discuss protected mode, which on the 386 comes with the ability to address up to 4 gigabytes per segment and 64 terabytes in all. In real mode (and in virtual-86 mode, which allows the 386 to multitask MS-DOS applications, and which is identical to real mode so far as MS-DOS programs are concerned), programs running on the 386 are still limited to 1 MB of addressable memory and 64K per segment.

@@ -98,10 +91,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-06.html b/11-06.html index 2d77dd7..8640755 100644 --- a/11-06.html +++ b/11-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -47,7 +40,7 @@ -

Optimization Rules: The More Things Change...

+

Optimization Rules: The More Things Change...

Let’s see what we’ve learned about 286/386 optimization. Mostly what we’ve learned is that our familiar PC cycle-eaters still apply, although in somewhat different forms, and that the major optimization rules for the PC hold true on ATs and 386-based computers. You won’t go wrong on any of these computers if you keep your instructions short, use the registers heavily and avoid memory, don’t branch, and avoid accessing display memory like the plague.

@@ -55,7 +48,7 @@

There’s one cycle-eater with new implications on the 286 and 386, and that’s the data alignment cycle-eater. From the data alignment cycle-eater we get a new rule: Word-align your word-sized variables, and start your subroutines at even addresses.

-

Detailed Optimization

+

Detailed Optimization

While the major 8088 optimization rules hold true on computers built around the 286 and 386, many of the instruction-specific optimizations no longer hold, for the execution times of most instructions are quite different on the 286 and 386 than on the 8088. We have already seen one such example of the sometimes vast difference between 8088 and 286/386 instruction execution times: MOV [WordVar],0, which has an Execution Unit execution time of 20 cycles on the 8088, has an EU execution time of just 3 cycles on the 286 and 2 cycles on the 386.

@@ -71,7 +64,7 @@

Theory confirmed.

-

LISTING 11.4 L11-4.ASM

+

LISTING 11.4 L11-4.ASM

 ;
 ; *** Listing 11.4 ***
@@ -85,7 +78,7 @@
         add     dx,100h
         endm
         call    ZTimerOff
-
+


@@ -104,10 +97,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-07.html b/11-07.html index dc565ae..885c180 100644 --- a/11-07.html +++ b/11-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -37,7 +30,7 @@


-

LISTING 11.5 L11-5.ASM

+

LISTING 11.5 L11-5.ASM

 ;
 ; *** Listing 11.5 ***
@@ -58,7 +51,7 @@ Skip:
         add     [WordVar]100h
         endm
         call    ZTimerOff
-
+

What’s going on? Simply this: Instruction fetching is controlling overall execution time on both processors. Both the 8088 in a PC and the 286 in an AT can execute the bytes of the instructions in Listings 11.4 and 11.5 faster than they can be fetched. Since the instructions are exactly the same lengths on both processors, it stands to reason that the ratio of the overall execution times of the instructions should be the same on both processors as well. Instruction length controls execution time, and the instruction lengths are the same—therefore the ratios of the execution times are the same. The 286 can both fetch and execute instruction bytes faster than the 8088 can, so code executes much faster on the 286; nonetheless, because the 286 can also execute those instruction bytes much faster than it can fetch them, overall performance is still largely determined by the size of the instructions.

@@ -68,7 +61,7 @@ Skip:

The more things change, the more they remain the same....

-

POPF and the 286

+

POPF and the 286

We’ve one final 286-related item to discuss: the hardware malfunction of POPF under certain circumstances on the 286.

@@ -97,10 +90,6 @@ Skip:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/11-08.html b/11-08.html index cce6015..f22716a 100644 --- a/11-08.html +++ b/11-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 286 and 386 - - @@ -41,11 +34,10 @@

Obviously, the segment:offset that IRET expects to find on the stack above the pushed flags isn’t present when the stack is set up for POPF, so we’ll have to adjust the stack a bit before we can substitute IRET for POPF. What we’ll have to do is push the segment:offset of the instruction after our workaround code onto the stack right above the pushed flags. IRET will then branch to that address and pop the flags, ending up at the instruction after the workaround code with the flags popped. That’s just the result that would have occurred had we executed POPF—WITH the bonus that no interrupts can accidentally occur when the Interrupt flag is 0 both before and after the pop.

-


- Figure 11.4
  The operation of POPF.

+


+ Figure 11.4
  The operation of POPF.

-

How can we push the segment:offset of the next instruction? Well, finding the offset of the next instruction by performing a near call to that instruction is a tried-and-true trick. We can do something similar here, but in this case we need a far call, since IRET requires both a segment and an offset. We’ll also branch backward so that the address pushed on the stack will point to the instruction we want to continue with. The code works out like this:

+

How can we push the segment:offset of the next instruction? Well, finding the offset of the next instruction by performing a near call to that instruction is a tried-and-true trick. We can do something similar here, but in this case we need a far call, since IRET requires both a segment and an offset. We’ll also branch backward so that the address pushed on the stack will point to the instruction we want to continue with. The code works out like this:

       jmpshort popfskip
 popfiret:
@@ -63,15 +55,14 @@ popfskip:
 ; the word that was on top of the stack when JMP SHORT POPFSKIP
 ; was reached has been popped into the FLAGS register, just as
 ; if a POPF instruction had been executed.
-
+ -


- Figure 11.5
  The operation of IRET.

+


+ Figure 11.5
  The operation of IRET.

The operation of this code is illustrated in Figure 11.6.

-

The POPF workaround can best be implemented as a macro; we can also emulate a far call by pushing CS and performing a near call, thereby shrinking the workaround code by 1 byte:

+

The POPF workaround can best be implemented as a macro; we can also emulate a far call by pushing CS and performing a near call, thereby shrinking the workaround code by 1 byte:

 EMULATE_POPF             macro
      local popfskip, popfiret
@@ -82,9 +73,9 @@ popfskip:
      push  cs
      call  popfiret
      endm
-
+ -

By the way, the flags can be popped much more quickly if you’re willing to alter a register in the process. For example, the following macro emulates POPF with just one branch, but wipes out AX:

+

By the way, the flags can be popped much more quickly if you’re willing to alter a register in the process. For example, the following macro emulates POPF with just one branch, but wipes out AX:

 EMULATE_POPF_TRASH_AX   macro
    push  cs
@@ -92,9 +83,9 @@ EMULATE_POPF_TRASH_AX   macro
    push  ax
    iret
    endm
-
+ -

It’s not a perfect substitute for POPF, since POPF doesn’t alter any registers, but it’s faster and shorter than EMULATE_POPF when you can spare the register. If you’re using 286-specific instructions, you can use which is shorter still, alters no registers, and branches just once. (Of course, this version of EMULATE_POPF won’t work on an 8088.)

+

It’s not a perfect substitute for POPF, since POPF doesn’t alter any registers, but it’s faster and shorter than EMULATE_POPF when you can spare the register. If you’re using 286-specific instructions, you can use which is shorter still, alters no registers, and branches just once. (Of course, this version of EMULATE_POPF won’t work on an 8088.)

       .286
                  :
@@ -103,11 +94,10 @@ EMULATE_POPFmacro
       pushoffset $+4
       iret
       endm
-
+ -


- Figure 11.6
  Workaround code for the POPF bug.

+


+ Figure 11.6
  Workaround code for the POPF bug.

The standard version of EMULATE_POPF is 6 bytes longer than POPF and much slower, as you’d expect given that it involves three branches. Anyone in his/her right mind would prefer POPF to a larger, slower, three-branch macro—given a choice. In noncode, however, there’s no choice here; the safer—if slower—approach is the best. (Having people associate your programs with crashed computers is not a desirable situation, no matter how unfair the circumstances under which it occurs.)

@@ -130,10 +120,6 @@ EMULATE_POPFmacro
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/12-01.html b/12-01.html index 34fb3e1..506541a 100644 --- a/12-01.html +++ b/12-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - @@ -37,10 +30,10 @@


-

Chapter 12
+

Chapter 12
Pushing the 486

-

It’s Not Just a Bigger 386

+

It’s Not Just a Bigger 386

So this traveling salesman is walking down a road, and he sees a group of men digging a ditch with their bare hands. “Whoa, there!” he says. “What you guys need is a Model 8088 ditch digger!” And he whips out a trowel and sells it to them.

@@ -52,13 +45,13 @@

Substitute “processor” for the various digging implements, and you get an idea of just how different the optimization rules for the 486 are from what you’re used to. Okay, it’s not quite that bad—but upon encountering a processor where string instructions are often to be avoided and memory-to-register MOVs are frequently as fast as register-to-register MOVs, Dorothy was heard to exclaim (before she sank out of sight in a swirl of hopelessly mixed metaphors), “I don’t think we’re in Kansas anymore, Toto.”

-

Enter the 486

+

Enter the 486

No chip that is a direct, fully compatible descendant of the 8088, 286, and 386 could ever be called a RISC chip, but the 486 certainly contains RISC elements, and it’s those elements that are most responsible for making 486 optimization unique. Simple, common instructions are executed in a single cycle by a RISC-like core processor, but other instructions are executed pretty much as they were on the 386, where every instruction takes at least 2 cycles. For example, MOV AL, [TestChar] takes only 1 cycle on the 486, assuming both instruction and data are in the cache—3 cycles faster than the 386—but STOSB takes 5 cycles, 1 cycle slower than on the 386. The floating-point execution unit inside the 486 is also much faster than the 387 math coprocessor, largely because, being in the same silicon as the CPU (the 486 has a math coprocessor built in), it is more tightly coupled. The results are sometimes startling: FMUL (floating point multiply) is usually faster on the 486 than IMUL (integer multiply)!

An encyclopedic approach to 486 optimization would take a book all by itself, so in this chapter I’m only going to hit the highlights of 486 optimization, touching on several optimization rules, some documented, some not. You might also want to check out the following sources of 486 information: i486 Microprocessor Programmer’s Reference Manual, from Intel; “8086 Optimization: Aim Down the Middle and Pray,” in the March, 1991 Dr. Dobb’s Journal; and “Peak Performance: On to the 486,” in the November, 1990 Programmer’s Journal.

-

Rules to Optimize By

+

Rules to Optimize By

In Appendix G of the i486 Microprocessor Programmers Reference Manual, Intel lists a number of optimization techniques for the 486. While neither exhaustive (we’ll look at two undocumented optimizations shortly) nor entirely accurate (we’ll correct two of the rules here), Intel’s list is certainly a good starting point. In particular, the list conveys the extent to which 486 optimization differs from optimization for earlier x86 processors. Generally, I’ll be discussing optimization for real mode (it being the most widely used mode at the moment), although many of the rules should apply to protected mode as well.

@@ -72,7 +65,7 @@

In other words, for cached code (which time-critical code almost always is), performance is predictable and can be calculated with good precision, and those calculations will apply on any 486. However, “predictable” doesn’t mean “trivial”; the cycle times printed for the various instructions are not the whole story. You must be aware of all the rules, documented and undocumented, that go into calculating actual execution times—and uncovering some of those rules is exactly what this chapter is about.

-

The Hazards of Indexed Addressing

+

The Hazards of Indexed Addressing

Rule #1: Avoid indexed addressing (that is, try not to use either two registers or scaled addressing to point to memory).

@@ -88,16 +81,16 @@ -

As an example, you might adhere to this rule by replacing the code

+

As an example, you might adhere to this rule by replacing the code

 LoopTop:
     add  ax,[bx+si]
     add  si,2
     dec  cx
     jnz  LoopTop
-
+ -

with this

+

with this

     add  si,bx
 LoopTop:
@@ -106,7 +99,7 @@ LoopTop:
     dec  cx
     jnz  LoopTop
     sub  si,bx
-
+


@@ -125,10 +118,6 @@ LoopTop:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/12-02.html b/12-02.html index 147321d..96449a6 100644 --- a/12-02.html +++ b/12-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - @@ -47,35 +40,35 @@

In a key loop on the 486, 1 cycle can indeed matter.

-

Calculate Memory Pointers Ahead of Time

+

Calculate Memory Pointers Ahead of Time

Rule #2: Don’t use a register as a memory pointer during the next two cycles after loading it.

Intel states that if the destination of one instruction is used as the base addressing component of the next instruction, then a one-cycle penalty is imposed. This rule, unlike anything ever before seen in the x86 family, reflects the heavily pipelined nature of the 486. Apparently, the 486 starts each effective address calculation before the start of the instruction that will need it, as shown in Figure 12.1; this effectively makes the address calculation time vanish, because it happens while the preceding instruction executes.

-

Of course, the 486 can’t perform an effective address calculation for a target instruction ahead of time if one of the address components isn’t known until the instruction starts, and that’s exactly the case when the preceding instruction modifies one of the target instruction’s addressing registers. For example, in the code

+

Of course, the 486 can’t perform an effective address calculation for a target instruction ahead of time if one of the address components isn’t known until the instruction starts, and that’s exactly the case when the preceding instruction modifies one of the target instruction’s addressing registers. For example, in the code

 MOV  BX,OFFSET MemVar
 MOV  AX,[BX]
-
+ -

there’s no way that the 486 can calculate the address referenced by MOV AX,[BX] until MOV BX,OFFSET MemVar finishes, so pipelining that calculation ahead of time is not possible. A good workaround is rearranging your code so that at least one instruction lies between the loading of the memory pointer and its use. For example, postdecrementing, as in the following

+

there’s no way that the 486 can calculate the address referenced by MOV AX,[BX] until MOV BX,OFFSET MemVar finishes, so pipelining that calculation ahead of time is not possible. A good workaround is rearranging your code so that at least one instruction lies between the loading of the memory pointer and its use. For example, postdecrementing, as in the following

 LoopTop:
     add    ax,[si]
     add    si,2
     dec    cx
     jnz    LoopTop
-
+ -

is faster than preincrementing, as in:

+

is faster than preincrementing, as in:

 LoopTop:
     add    si,2
     add    ax,[SI]
     dec    cx
     jnz    LoopTop
-
+

Now that we understand what Intel means by this rule, let me make a very important comment: My observations indicate that for real-mode code, the documentation understates the extent of the penalty for interrupting the address calculation pipeline by loading a memory pointer just before it’s used.

@@ -91,46 +84,44 @@ LoopTop:

Considering that MOV normally takes only one cycle total, that’s quite a loss. For example, the postdecrement loop shown above is 2 full cycles faster than the preincrement loop, resulting in a 29 percent improvement in the performance of the entire loop. But wait, there’s more. If a register is loaded 2 cycles (which generally means 2 instructions, but, because some 486 instructions take more than 1 cycle,

-


- Figure 12.1
  One-cycle-ahead address pipelining.

+


+ Figure 12.1
  One-cycle-ahead address pipelining.

-

the 2 are not always equivalent) before it’s used to point to memory, 1 cycle is lost. Therefore, whereas this code

+

the 2 are not always equivalent) before it’s used to point to memory, 1 cycle is lost. Therefore, whereas this code

 mov    bx,offset MemVar
 mov    ax,[bx]
 inc    dx
 dec    cx
 jnz    LoopTop
-
+ -

loses two cycles from interrupting the address calculation pipeline, this code

+

loses two cycles from interrupting the address calculation pipeline, this code

 mov    bx,offset MemVar
 inc    dx
 mov    ax,[bx]
 dec    cx
 jnz    LoopTop
-
+ -

loses only one cycle, and this code

+

loses only one cycle, and this code

 mov    bx,offset MemVar
 inc    dx
 dec    cx
 mov    ax,[bx]
 jnz    LoopTop
-
+

loses no cycles at all. Apparently, the 486’s addressing calculation pipeline actually starts 2 cycles ahead, as shown in Figure 12.2. (In truth, my best guess at the moment is that the addressing pipeline really does start only 1 cycle ahead; the additional cycle crops up when the addressing pipeline has to wait for a register to be written into the register file before it can read it out for use in addressing calculations. However, I’m guessing here, and the 2-cycle-ahead model in Figure 12.2 will do just fine for optimization purposes.)

Clearly, there’s considerable optimization potential in careful rearrangement of 486 code.

-


- Figure 12.2
  Two-cycle-ahead address pipelining.

+


+ Figure 12.2
  Two-cycle-ahead address pipelining.

-

Caveat Programmor

+

Caveat Programmor

A caution: I’m quite certain that the 2-cycle-ahead addressing pipeline interruption penalty I’ve described exists in the two 486s I’ve tested. However, there’s no guarantee that Intel won’t change this aspect of the 486 in the future, especially given that the documentation indicates otherwise. Perhaps the 2-cycle penalty is the result of a bug in the initial steps of the 486, and will revert to the documented 1-cycle penalty someday; likewise for the undocumented optimizations I’ll describe below. Nonetheless, none of the optimizations I suggest would hurt performance even if the undocumented performance characteristics of the 486 were to vanish, and they certainly will help performance on at least some 486s right now, so I feel they’re well worth using.

@@ -151,10 +142,6 @@ jnz LoopTop
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/12-03.html b/12-03.html index 7faf1d6..53b5db0 100644 --- a/12-03.html +++ b/12-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - @@ -39,13 +32,13 @@

There is, of course, no guarantee that I’m entirely correct about the optimizations discussed in this chapter. Without knowing the internals of the 486, all I can do is time code and make inferences from the results; I invite you to deduce your own rules and cross-check them against mine. Also, most likely there are other optimizations that I’m unaware of. If you have further information on these or any other undocumented optimizations, please write and let me know. And, of course, if anyone from Intel is reading this and wants to give us the gospel truth, please do!

-

Stack Addressing and Address Pipelining

+

Stack Addressing and Address Pipelining

Rule #2A: Rule #2 sometimes, but not always, applies to the stack pointer when it is implicitly used to point to memory.

Intel states that the stack pointer is an implied destination register for CALL, ENTER, LEAVE, RET, PUSH, and POP (which alter (E)SP), and that it is the implied base addressing register for PUSH, POP, and RET (which use (E)SP to address memory). Intel then implies that the aforementioned addressing pipeline penalty is incurred whenever the stack pointer is used as a destination by one of the first set of instructions and is then immediately used to address memory by one of the second set. This raises the specter of unpleasant programming contortions such as intermixing PUSHes and POPs with other instructions to avoid interrupting the addressing pipeline. Fortunately, matters are actually not so grim as Intel’s documentation would indicate; my tests indicate that the addressing pipeline penalty pops up only spottily when the stack pointer is involved.

-

For example, you’d certainly expect a sequence such as

+

For example, you’d certainly expect a sequence such as

 :
 pop    ax
@@ -53,67 +46,67 @@ ret
 pop    ax
 et
 :
-
+ -

to exhibit the addressing pipeline interruption phenomenon (SP is both destination and addressing register for both instructions, according to Intel), but this code runs in six cycles per POP/RET pair, matching the official execution times exactly. Likewise, a sequence like

+

to exhibit the addressing pipeline interruption phenomenon (SP is both destination and addressing register for both instructions, according to Intel), but this code runs in six cycles per POP/RET pair, matching the official execution times exactly. Likewise, a sequence like

 pop    dx
 pop    cx
 pop    bx
 pop    ax
-
+

runs in one cycle per instruction, just as it should.

-

On the other hand, performing arithmetic directly on SP as an explicit destination—for example, to deallocate local variables—and then using PUSH, POP, or RET, definitely can interrupt the addressing pipeline. For example

+

On the other hand, performing arithmetic directly on SP as an explicit destination—for example, to deallocate local variables—and then using PUSH, POP, or RET, definitely can interrupt the addressing pipeline. For example

 add    sp,10h
 ret
-
+ -

loses two cycles because SP is the explicit destination of one instruction and then the implied addressing register for the next, and the sequence

+

loses two cycles because SP is the explicit destination of one instruction and then the implied addressing register for the next, and the sequence

 add    sp,10h
 pop    ax
-
+

loses two cycles for the same reason.

I certainly haven’t tried all possible combinations, but the results so far indicate that the stack pointer incurs the addressing pipeline penalty only if (E)SP is the explicit destination of one instruction and is then used by one of the two following instructions to address memory. So, for instance, SP isn’t the explicit operand of POP AX—AX is—and no cycles are lost if POP AX is followed by POP or RET. Happily, then, we need not worry about the sequence in which we use PUSH and POP. However, adding to, moving to, or subtracting from the stack pointer should ideally be done at least two cycles before PUSH, POP, RET, or any other instruction that uses the stack pointer to address memory.

-

Problems with Byte Registers

+

Problems with Byte Registers

There are two ways to lose cycles by using byte registers, and neither of them is documented by Intel, so far as I know. Let’s start with the lesser and simpler of the two.

Rule #3: Do not load a byte portion of a register during one instruction, then use that register in its entirety as a source register during the next instruction.

-

So, for example, it would be a bad idea to do this

+

So, for example, it would be a bad idea to do this

 mov    ah,o
             :
 mov    cx,[MemVar1]
 mov    al,[MemVar2]
 add    cx,ax
-
+ -

because AL is loaded by one instruction, then AX is used as the source register for the next instruction. A cycle can be saved simply by rearranging the instructions so that the byte register load isn’t immediately followed by the word register usage, like so:

+

because AL is loaded by one instruction, then AX is used as the source register for the next instruction. A cycle can be saved simply by rearranging the instructions so that the byte register load isn’t immediately followed by the word register usage, like so:

 mov    ah,o
             :
 mov    al,[MemVar2]
 mov    cx,[MemVar1]
 add    cx,ax
-
+

Strange as it may seem, this rule is neither arbitrary nor nonsensical. Basically, when a byte destination register is part of a word source register for the next instruction, the 486 is unable to directly use the result from the first instruction as the source for the second instruction, because only part of the register required by the second instruction is contained in the first instruction’s result. The full, updated register value must be read from the register file, and that value can’t be read out until the result from the first instruction has been written into the register file, a process that takes an extra cycle. I’m not going to explain this in great detail because it’s not important that you understand why this rule exists (only that it does in fact exist), but it is an interesting window on the way the 486 works.

-

In case you’re curious, there’s no such penalty for the typical XLAT sequence like

+

In case you’re curious, there’s no such penalty for the typical XLAT sequence like

 mov    bx,offset MemTable
        :
 mov    al,[si]
 xlat
-
+

even though AL must be converted to a word by XLAT before it can be added to BX and used to address memory. In fact, none of the penalties mentioned in this chapter apply to XLAT, apparently because XLAT is so slow—4 cycles—that it gives the 486 time to perform addressing calculations during the course of the instruction.

@@ -144,10 +137,6 @@ xlat
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/12-04.html b/12-04.html index cc56b4b..a3d7d8f 100644 --- a/12-04.html +++ b/12-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pushing the 486 - - @@ -39,41 +32,41 @@

You don’t need to understand every corner of the 486 universe unless you’re a diehard ASMhead who does this stuff for fun. Just learn enough to be able to speed up the key portions of your programs, and spend the rest of your time on a fast design and overall implementation.

-

More Fun with Byte Registers

+

More Fun with Byte Registers

Rule #4: Don’t load any byte register exactly 2 cycles before using any register to address memory.

-

This, the last of this chapter’s rules, is the strangest of the lot. If any byte register is loaded, and then two cycles later any register is used to point to memory, one cycle is lost. So, for example, this code

+

This, the last of this chapter’s rules, is the strangest of the lot. If any byte register is loaded, and then two cycles later any register is used to point to memory, one cycle is lost. So, for example, this code

 mov    al,bl
 mov    cx,dx
 mov    si,[di]
-
+

takes four rather than the expected three cycles to execute. Note that it is not required that the byte register be part of the register used to address memory; any byte register will do the trick.

-

Worse still, loading byte registers both one and two cycles before a register is used to address memory costs two cycles, as in

+

Worse still, loading byte registers both one and two cycles before a register is used to address memory costs two cycles, as in

 mov    bl,al
 mov    cl,3
 mov    bx,[si]
-
+ -

which takes five rather than three cycles to run. However, there is no penalty if a byte register is loaded one cycle but not two cycles before a register is used to address memory. Therefore,

+

which takes five rather than three cycles to run. However, there is no penalty if a byte register is loaded one cycle but not two cycles before a register is used to address memory. Therefore,

 mov    cx,3
 mov    dl,al
 mov    si,[bx]
-
+

runs in the expected three cycles.

-

In truth, I do not know why this happens. Clearly, it has something to do with interrupting the start of the addressing pipeline, and I have my theories about how this works, but at this point they’re pure speculation. Whatever the reason for this rule, ignorance of it—and of its interaction with the other rules—could lead to considerable performance loss in seemingly air-tight code. For instance, a casual observer would expect the following code to run in 3 cycles:

+

In truth, I do not know why this happens. Clearly, it has something to do with interrupting the start of the addressing pipeline, and I have my theories about how this works, but at this point they’re pure speculation. Whatever the reason for this rule, ignorance of it—and of its interaction with the other rules—could lead to considerable performance loss in seemingly air-tight code. For instance, a casual observer would expect the following code to run in 3 cycles:

 mov    bx,offset MemVar
 mov    cl,al
 mov    ax,[bx]
-
+

A more sophisticated programmer would expect to lose one cycle, because BX is loaded two cycles before being used to address memory. In fact, though, this code takes 5 cycles—2 cycles, or 67 percent, longer than normal. Why? Well, under normal conditions, loading a byte register—CL in this case—one cycle before using a register to address memory produces no penalty; loading 2 cycles ahead is the only case that normally incurs a penalty. However, think of Rule #4 as meaning that loading a byte register disrupts the memory addressing pipeline as it starts up. Viewed that way, we can see that MOV BX,OFFSET MemVar interrupts the addressing pipeline, forcing it to start again, and then, presumably, MOV CL,AL interrupts the pipeline again because the pipeline is now on its first cycle: the one that loading a byte register can affect.

@@ -85,11 +78,11 @@ mov ax,[bx] -

Timing Your Own 486 Code

+

Timing Your Own 486 Code

In case you want to do some 486 performance analysis of your own, let me show you how I arrived at one of the above conclusions; at the same time, I can warn you of the timing hazards of the cache. Listings 12.1 and 12.2 show the code I ran through the Zen timer in order to establish the effects of loading a byte register before using a register to address memory. Listing 12.1 ran in 120 µs on a 33 MHz 486, or 4 cycles per repetition (120 µs/1000 repetitions = 120 ns per repetition; 120 ns per repetition/30 ns per cycle = 4 cycles per repetition); Listing 12.2 ran in 90 µs, or 3 cycles, establishing that loading a byte register costs a cycle only when it’s performed exactly 2 cycles before addressing memory.

-

LISTING 12.1 LST12-1.ASM

+

LISTING 12.1 LST12-1.ASM

 ; Measures the effect of loading a byte register 2 cycles before
 ; using a register to address memory.
@@ -108,9 +101,9 @@ CacheFillLoop:
     jz      Done
     jmp     CacheFillLoop
 Done:
-
+ -

LISTING 12.2 LST12-2.ASM

+

LISTING 12.2 LST12-2.ASM

 ; Measures the effect of loading a byte register 1 cycle before
 ; using a register to address memory.
@@ -129,7 +122,7 @@ CacheFillLoop:
     jz     Done
     jmp    CacheFillLoop
 Done:
-
+

Note that Listings 12.1 and 12.2 each repeat the timing of the code under test a second time, to make sure that the instructions are in the cache on the second pass, the one for which results are displayed. Also note that the code is less than 8K in size, so that it can all fit in the 486’s 8K internal cache. If I double the REPT value in Listing 12.2 to 2,000, making the test code larger than 8K, the execution time more than doubles to 224 µs, or 3.7 cycles per repetition; the extra seven-tenths of a cycle comes from fetching non-cached instruction bytes.

@@ -141,7 +134,7 @@ Done: -

The Story Continues

+

The Story Continues

There’s certainly plenty more 486 lore to explore, including the 486’s unique prefetch queue, more optimization rules, branching optimizations, performance implications of the cache, the cost of cache misses for reads, and the implications of cache write-through for writes. Nonetheless, we’ve covered quite a bit of ground in this chapter, and I trust you’ve gotten a feel for the considerable extent to which 486 optimization differs from what you’re used to. Odd as 486 optimization is, though, it’s well worth mastering, for the 486 is, at its best, so staggeringly fast that carefully crafted 486 code can do more than twice as much per cycle as the best 386 code—which makes it perhaps 50 times as fast as optimized code for the original PC.

@@ -164,10 +157,6 @@ Done:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/13-01.html b/13-01.html index 141abd6..9fab9fc 100644 --- a/13-01.html +++ b/13-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - @@ -37,10 +30,10 @@


-

Chapter 13
+

Chapter 13
Aiming the 486

-

Pipelines and Other Hazards of the High End

+

Pipelines and Other Hazards of the High End

It’s a sad but true fact that 84 percent of American schoolchildren are ignorant of 92 percent of American history. Not my daughter, though. We recently visited historical Revolutionary-War-vintage Fort Ticonderoga, and she’s now 97 percent aware of a key element of our national heritage: that the basic uniform for soldiers in those days was what appears to be underwear, plus a hat so that no one could complain that they were undermining family values. Ha! Just kidding! Actually, what she learned was that in those days, it was pure coincidence if a cannonball actually hit anything it was aimed at, which isn’t surprising considering the lack of rifling, precision parts, and ballistics. The guides at the fort shot off three cannons; the closest they came to the target was about 50 feet, and that was only because the wind helped. I think the idea in early wars was just to put so much lead in the air that some of it was bound to hit something; preferably, but not necessarily, the enemy.

@@ -50,27 +43,26 @@

For example, consider how Terje Mathisen doubled the speed of his word-counting program on a 486 simply by shuffling a couple of instructions.

-

486 Pipeline Optimization

+

486 Pipeline Optimization

I’ve mentioned Terje Mathisen in my writings before. Terje is an assembly language programmer extraordinaire, and author of the incredibly fast public-domain word-counting program WC (which comes complete with source code; well worth a look, if you want to see what really fast code looks like). Terje’s a regular participant in the ibm.pc/fast.code topic on Bix. In a thread titled “486 Pipeline Optimization, or TANSTATFC (There Ain’t No Such Thing As The Fastest Code),” he detailed the following optimization to WC, perhaps the best example of 486 pipeline optimization I’ve yet seen.

Terje’s inner loop originally looked something like the code in Listing 13.1. (I’ve taken a few liberties for illustrative purposes.) Of course, Terje unrolls this loop a few times (128 times, to be exact). By the way, in Listing 13.1 you’ll notice that Terje counts not only words but also lines, at a rate of three instructions for every two characters!

-

LISTING 13.1 L13-1.ASM

+

LISTING 13.1 L13-1.ASM

 mov di,[bp+OFFS]    ;get the next pair of characters
 mov bl,[di]         ;get the state value for the pair
 add dx,[bx+8000h]   ;increment word and line count
                     ; appropriately for the pair
-
+

Listing 13.1 looks as tight as it could be, with just two one-cycle instructions, one two-cycle instruction, and no branches. It is tight, but those three instructions actually take a minimum of 8 cycles to execute, as shown in Figure 13.1. The problem is that DI is loaded just before being used to address memory, and that costs 2 cycles because it interrupts the 486’s internal instruction pipeline. Likewise, BX is loaded just before being used to address memory, costing another two cycles. Thus, this loop takes twice as long as cycle counts would seem to indicate, simply because two registers are loaded immediately before being used, disrupting the 486’s pipeline.

Listing 13.2 shows Terje’s immediate response to these pipelining problems; he simply swapped the instructions that load DI and BL. This one change cut execution time per character pair from eight cycles to five cycles! The load of BL is now separated by one instruction from the use of BX to address memory, so the pipeline penalty is reduced from two cycles to one cycle. The load of DI is also separated by one instruction from the use of DI to address memory (remember, the loop is unrolled, so the last instruction is followed by the first instruction), but because the intervening instruction takes two cycles, there’s no penalty at all.

-


- Figure 13.1
  Cycle-eaters in the original WC.

+


+ Figure 13.1
  Cycle-eaters in the original WC.

@@ -80,13 +72,13 @@ add dx,[bx+8000h] ;increment word and line count
-

LISTING 13.2 L13-2.ASM

+

LISTING 13.2 L13-2.ASM

 mov bl,[di]         ;get the state value for the pair
 mov di,[bp+OFFS]    ;get the next pair of characters
 add dx,[bx+8000h]   ;increment word and line count
                     ; appropriately for the pair
-
+

At this point, Terje had nearly doubled the performance of this code simply by moving one instruction. (Note that swapping the instructions also made it necessary to preload DI at the start of the loop; Listing 13.2 is not exactly equivalent to Listing 13.1.) I’ll let Terje describe his next optimization in his own words:

@@ -107,10 +99,6 @@ add dx,[bx+8000h] ;increment word and line count
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/13-02.html b/13-02.html index 07f4aa9..33aa119 100644 --- a/13-02.html +++ b/13-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - @@ -39,47 +32,46 @@

“When I looked closely as this, I realized that the two cycles for the final ADD is just the sum of 1 cycle to load the data from memory, and 1 cycle to add it to DX, so the code could just as well have been written as shown in Listing 13.3. The final breakthrough came when I realized that by initializing AX to zero outside the loop, I could rearrange it as shown in Listing 13.4 and do the final ADD DX,AX after the loop. This way there are two single-cycle instructions between the first and the fourth line, avoiding all pipeline stalls, for a total throughput of two cycles/char.”

-

LISTING 13.3 L13-3.ASM

+

LISTING 13.3 L13-3.ASM

 mov bl,[di]         ;get the state value for the pair
 mov di,[bp+OFFS]    ;get the next pair of characters
 mov ax,[bx+8000h]   ;increment word and line count
 add dx,ax           ; appropriately for the pair
-
+ -

LISTING 13.4 L13-4.ASM

+

LISTING 13.4 L13-4.ASM

 mov bl,[di]         ;get the state value for the pair
 mov di,[bp+OFFS]    ;get the next pair of characters
 add dx,ax           ;increment word and line count
                     ; appropriately for the pair
 mov ax,[bx+8000h]   ;get increments for next time
-
+

I’d like to point out two fairly remarkable things. First, the single cycle that Terje saved in Listing 13.4 sped up his entire word-counting engine by 25 percent or more; Listing 13.4 is fully twice as fast as Listing 13.1—all the result of nothing more than shifting an instruction and splitting another into two operations. Second, Terje’s word-counting engine can process more than 16 million characters per second on a 486/33.

Clever 486 optimization can pay off big. QED.

-

BSWAP: More Useful Than You Might Think

+

BSWAP: More Useful Than You Might Think

-

There are only 3 non-system instructions unique to the 486. None is earthshaking, but they have their uses. Consider BSWAP. BSWAP does just what its name implies, swapping the bytes (not bits) of a 32-bit register from one end of the register to the other, as shown in Figure 13.2. (BSWAP can only work with 32-bit registers; memory locations and 16-bit registers are not valid operands.) The obvious use of BSWAP is to convert data from Intel format (least significant byte first in memory, also called little endian) to Motorola format (most significant byte first in memory, or big endian), like so:

+

There are only 3 non-system instructions unique to the 486. None is earthshaking, but they have their uses. Consider BSWAP. BSWAP does just what its name implies, swapping the bytes (not bits) of a 32-bit register from one end of the register to the other, as shown in Figure 13.2. (BSWAP can only work with 32-bit registers; memory locations and 16-bit registers are not valid operands.) The obvious use of BSWAP is to convert data from Intel format (least significant byte first in memory, also called little endian) to Motorola format (most significant byte first in memory, or big endian), like so:

 lodsd
 bswap
 stosd
-
+

BSWAP can also be useful for reversing the order of pixel bits from a bitmap so that they can be rotated 32 bits at a time with an instruction such as ROR EAX,1. Intel’s byte ordering for multiword values (least-significant byte first) loads pixels in the wrong order, so far as word rotation is concerned, but BSWAP can take care of that.

-


- Figure 13.2
  BSWAP in operation.

+


+ Figure 13.2
  BSWAP in operation.

As it turns out, though, BSWAP is also useful in an unexpected way, having to do with making efficient use of the upper half of 32-bit registers. As any assembly language programmer knows, the x86 register set is too small; or, to phrase that another way, it sure would be nice if the register set were bigger. As any 386/486 assembly language programmer knows, there are many cases in which 16 bits is plenty. For example, a 16-bit scan-line counter generally does the trick nicely in a video driver, because there are very few video devices with more than 65,535 addressable scan lines. Combining these two observations yields the obvious conclusion that it would be great if there were some way to use the upper and lower 16 bits of selected 386 registers as separate 16-bit registers, effectively increasing the available register space.

Unfortunately, the x86 instruction set doesn’t provide any way to work directly with only the upper half of a 32-bit register. The next best solution is to rotate the register to give you access in the lower 16 bits to the half you need at any particular time, with code along the lines of that in Listing 13.5. Having to rotate the 16-bit fields into position certainly isn’t as good as having direct access to the upper half, but surely it’s better than having to get the values out of memory, isn’t it?

-

LISTING 13.5 L13-5.ASM

+

LISTING 13.5 L13-5.ASM

 mov   cx,[initialskip]
 shl   ecx,16       ;put skip value in upper half of ECX
@@ -92,7 +84,7 @@ looptop:
       ror   ecx,16      ;put loop count in CX
       dec   cx          ;count down loop
       jnz   looptop
-
+

Not necessarily. Shifts and rotates are among the worst performing instructions of the 486, taking 2 to 3 cycles to execute. Thus, it takes 2 cycles to rotate the skip value into CX in Listing 13.5, and 2 more cycles to rotate it back to the upper half of ECX. I’d say four cycles is a pretty steep price to pay, especially considering that a MOV to or from memory takes only one cycle. Basically, using ROR to access a 16-bit value in the upper half of a 16-bit register is a pretty marginal technique, unless for some reason you can’t access memory at all (for example, if you’re using BP as a working register, temporarily making the stack frame inaccessible).

@@ -113,10 +105,6 @@ looptop:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/13-03.html b/13-03.html index fde0f12..a383085 100644 --- a/13-03.html +++ b/13-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - @@ -39,7 +32,7 @@

On the 386, ROR was the only way to split a 32-bit register into two 16-bit registers. On the 486, however, BSWAP can not only do the job, but can do it better, because BSWAP executes in just one cycle. BSWAP has the added benefit of not affecting any flags, unlike ROR. With BSWAP-based code like that in Listing 13.6, the upper 16 bits of a register can be accessed with only 2 cycles of overhead and without altering any flags, making the technique of packing two 16-bit registers into one 32-bit register much more useful.

-

LISTING 13.6 L13-6.ASM

+

LISTING 13.6 L13-6.ASM

       mov    cx,[initialskip]
       bswap  ecx        ;put skip value in upper half of ECX
@@ -52,20 +45,20 @@ looptop:
       bswap  ecx        ;put loop count in CX
       dec    cx         ;count down loop
       jnz    looptop
-
+ -

Pushing and Popping Memory

+

Pushing and Popping Memory

-

Pushing or popping a memory location, as in PUSH WORD PTR [BX] or POP [MemVar], is a compact, easy way to get a value onto or off of the stack, especially when pushing parameters for calling a C-compatible function. However, on a 486, these are unattractive instructions from a performance perspective. Pushing a memory location takes four cycles; by contrast, loading a memory location into a register takes only one cycle, and pushing a register takes just 1 more cycle, for a total of two cycles. Therefore,

+

Pushing or popping a memory location, as in PUSH WORD PTR [BX] or POP [MemVar], is a compact, easy way to get a value onto or off of the stack, especially when pushing parameters for calling a C-compatible function. However, on a 486, these are unattractive instructions from a performance perspective. Pushing a memory location takes four cycles; by contrast, loading a memory location into a register takes only one cycle, and pushing a register takes just 1 more cycle, for a total of two cycles. Therefore,

 mov   ax,[bx]
 push  ax
-
+ -

is twice as fast as

+

is twice as fast as

 push   word ptr [bx]
-
+

and the only cost is that the previous contents of AX are destroyed.

@@ -81,16 +74,16 @@ push word ptr [bx] -

Optimal 1-Bit Shifts and Rotates

+

Optimal 1-Bit Shifts and Rotates

On a 486, the n-bit forms of the shift and rotate instructions—as in ROR AX,2 and SHL BX,9—are 2-cycle instructions, but the 1-bit forms—as in ROR AX,1 and SHL BX,1—are 3-cycle instructions. Go figure.

-

Assemblers default to the 1-bit instruction for 1-bit shifts and rotates. That’s not unreasonable since the 1-bit form is a byte shorter and is just as fast as the n-bit forms on a 386 and faster on a 286, and the n-bit form doesn’t even exist on an 8088. In a really critical loop, however, it might be worth hand-assembling the n-bit form of a single-bit shift or rotate in order to save that cycle. The easiest way to do this is to assemble a 2-bit form of the desired instruction, as in SHL AX,2, then look at the hex codes that the assembler generates and use DB to insert them in your program code, with the value two replaced with the value one. For example, you could determine that SHL AX,2 assembles to the bytes 0C1H 0E0H 002H, either by looking at the disassembly in a debugger or by having the assembler generate a listing file. You could then insert the n-bit version of SHL AX,1 in your code as follows:

+

Assemblers default to the 1-bit instruction for 1-bit shifts and rotates. That’s not unreasonable since the 1-bit form is a byte shorter and is just as fast as the n-bit forms on a 386 and faster on a 286, and the n-bit form doesn’t even exist on an 8088. In a really critical loop, however, it might be worth hand-assembling the n-bit form of a single-bit shift or rotate in order to save that cycle. The easiest way to do this is to assemble a 2-bit form of the desired instruction, as in SHL AX,2, then look at the hex codes that the assembler generates and use DB to insert them in your program code, with the value two replaced with the value one. For example, you could determine that SHL AX,2 assembles to the bytes 0C1H 0E0H 002H, either by looking at the disassembly in a debugger or by having the assembler generate a listing file. You could then insert the n-bit version of SHL AX,1 in your code as follows:

 mov   ax,1
 db    0c1h, 0e0h, 001h
 mov   dx,ax
-
+

At the end of this sequence, DX will contain 2, and the fast n-bit version of SHL AX,1 will have executed. If you use this approach, I’d recommend using a macro, rather than sticking DBs in the middle of your code.

@@ -113,10 +106,6 @@ mov dx,ax
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/13-04.html b/13-04.html index a2146e4..093a597 100644 --- a/13-04.html +++ b/13-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Aiming the 486 - - @@ -37,12 +30,12 @@


-

32-Bit Addressing Modes

+

32-Bit Addressing Modes

-

The 386 and 486 both support 32-bit addressing modes, in which any register may serve as the base memory addressing register, and almost any register may serve as the potentially scaled index register. For example,

+

The 386 and 486 both support 32-bit addressing modes, in which any register may serve as the base memory addressing register, and almost any register may serve as the potentially scaled index register. For example,

 mov al,BaseTable[ecx+edx*4]
-
+

uses a perfectly valid 32-bit address, with the byte accessed being the one at the offset in DS pointed to by the sum of EDX times 4 plus the offset of BaseTable plus ECX. This is a very powerful memory addressing scheme, far superior to 8088-style 16-bit addressing, but it’s not without its quirks and costs, so let’s take a quick look at 32-bit addressing. (By the way, 32-bit addressing is not limited to protected mode; 32-bit instructions may be used in real mode, although each instruction that uses 32-bit addressing must have an address-size prefix byte, and the presence of a prefix byte costs a cycle on a 486.)

@@ -60,16 +53,16 @@ mov al,BaseTable[ecx+edx*4] -

However, because 32-bit addressing supports many more addressing combinations than 16-bit addressing, the Mod-R/M byte can’t describe all the combinations. Therefore, whenever an index register (as described above) is involved, a second byte, the SIB byte, follows the Mod-R/M byte to provide additional address information. Consequently, whenever you use a scaled memory addressing register or use the sum of two registers to point to memory, you automatically add 1 cycle and 1 byte to that instruction. This is not to say that you shouldn’t use index registers when they’re needed, but if you find yourself using them inside key loops, you should see if it’s possible to move the index calculation outside the loop as, for example, in a loop like this:

+

However, because 32-bit addressing supports many more addressing combinations than 16-bit addressing, the Mod-R/M byte can’t describe all the combinations. Therefore, whenever an index register (as described above) is involved, a second byte, the SIB byte, follows the Mod-R/M byte to provide additional address information. Consequently, whenever you use a scaled memory addressing register or use the sum of two registers to point to memory, you automatically add 1 cycle and 1 byte to that instruction. This is not to say that you shouldn’t use index registers when they’re needed, but if you find yourself using them inside key loops, you should see if it’s possible to move the index calculation outside the loop as, for example, in a loop like this:

 LoopTop:
       add   ax,DataTable[ebx*2]
       inc   ebx
       dec   cx
       jnz   LoopTop
-
+ -

You could change this to the following for greater performance:

+

You could change this to the following for greater performance:

       add   ebx,ebx      ;ebx*2
 LoopTop:
@@ -78,7 +71,7 @@ LoopTop:
       dec   cx
       jnz   LoopTop
       shr   ebx,1 ;ebx*2/2
-
+

I’ll end this chapter with two more quirks of 32-bit addressing. First, as with 16-bit addressing, addressing that uses EBP as a base register both accesses the SS segment by default and always has a displacement of at least 1 byte. This reflects the common use of EBP to address a stack frame, but is worth keeping in mind if you should happen to use EBP to address non-stack memory.

@@ -101,10 +94,6 @@ LoopTop:
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/14-01.html b/14-01.html index 905a800..15ac494 100644 --- a/14-01.html +++ b/14-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - @@ -37,10 +30,10 @@


-

Chapter 14
+

Chapter 14
Boyer-Moore String Searching

-

Optimizing a Pretty Optimum Search Algorithm

+

Optimizing a Pretty Optimum Search Algorithm

When you seem to be stumped, stop for a minute and think. All the information you need may be right in front of your nose if you just look at things a little differently. Here’s a case in point:

@@ -58,7 +51,7 @@

As I said, sometimes everything you need to know is right in front of your nose. Which brings us to Boyer-Moore string searching.

-

String Searching Refresher

+

String Searching Refresher

I’ve discussed string searching earlier in this book, in Chapters 5 and 9. You may want to refer back to these chapters for some background on string searching in general. I’m also going to use some of the code from that chapter as part of this chapter’s test suite. For further information, you may want to refer to the discussion of string searching in the excellent Algorithms in C, by Robert Sedgewick (Addison-Wesley), which served as the primary reference for this chapter. (If you look at Sedgewick, be aware that in the Boyer-Moore listing on page 288, there is a mistake: “j > 0” in the for loop should be “j >= 0,” unless I’m missing something.)

@@ -101,10 +94,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/14-02.html b/14-02.html index b55d312..9ad7519 100644 --- a/14-02.html +++ b/14-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - @@ -37,7 +30,7 @@


-

The Boyer-Moore Algorithm

+

The Boyer-Moore Algorithm

All our a priori knowledge of string searching is stated above, but there’s another sort of knowledge—knowledge that’s generated dynamically. As we search through the buffer, we acquire information each time we check for a match. One sort of information that we acquire is based on partial matches; we can often skip ahead after partial matches because (take a deep breath!) by partially matching, we have already implicitly done a comparison of the partially matched buffer characters with all possible pattern start locations that overlap those partially-matched bytes.

@@ -59,25 +52,22 @@

Figure 14.1 illustrates the operation of a Boyer-Moore search when the rightcharacter of the search pattern (which is the first character that’s compared at each location because we’re comparing backwards) mismatches with a buffer character that appears nowhere in the pattern. Figure 14.2 illustrates the operation of a partial match when the mismatch occurs with a character that’s not a pattern member. In this case, we can only skip ahead past the mismatch location, resulting in an advance of fewer bytes than the pattern length, and potentially as little as the same single byte distance by which the standard search approach advances.

-


- Figure 14.1
  Mismatch on first character checked.

+


+ Figure 14.1
  Mismatch on first character checked.

What if the mismatch occurs with a buffer character that does occur in the pattern? Then we can’t skip past the mismatch location, but we can skip to whatever location aligns the rightmost occurrence of that character in the pattern with the mismatch location, as shown in Figure 14.3.

Basically, we exercise our right as members of a free society to compare strings in whichever direction we choose, and we choose to do so right to left, rather than the more intuitive left to right. Whenever we find a mismatch, we see what we can learn from the buffer character that failed to match the pattern. Imagine that we move the pattern to the right across the mismatch location until we find a start location that the mismatch does not eliminate as a possible match for the pattern. If the mismatch character doesn’t appear in the pattern, the pattern can move clear past the mismatch location. Otherwise, the pattern moves until a matching pattern byte lies atop the mismatch. That’s all there is to it!

-


- Figure 14.2
  Mismatch on third character checked.

+


+ Figure 14.2
  Mismatch on third character checked.

-

Boyer-Moore: The Good and the Bad

+

Boyer-Moore: The Good and the Bad

The worst case for this version of Boyer-Moore is that the pattern mismatches on the leftmost character—the last character compared—every time. Again, not very likely, but it is true that this version of Boyer-Moore performs better as there are fewer and shorter partial matches; ideally, the rightmost character would never match until the full match location was reached. Longer patterns, which make for longer skips, help Boyer-Moore, as does a long distance to the match location, which helps diffuse the overhead of building the table of distances to skip ahead on all the possible mismatch values.

-


- Figure 14.3
  Mismatch on character that appears in pattern.

+


+ Figure 14.3
  Mismatch on character that appears in pattern.

The best case for Boyer-Moore is good indeed: About N/M comparisons are required, where N is the buffer length and M is the pattern length. This reflects the ability of Boyer-Moore to skip ahead by a full pattern length on a complete mismatch.

@@ -98,10 +88,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/14-03.html b/14-03.html index 5671bc4..0ec8d4e 100644 --- a/14-03.html +++ b/14-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - @@ -221,10 +214,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/14-04.html b/14-04.html index d2d1cd7..f83e2af 100644 --- a/14-04.html +++ b/14-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - @@ -37,7 +30,7 @@


-

LISTING 14.1 L14-1.C

+

LISTING 14.1 L14-1.C

 /* Searches a buffer for a specified pattern. In case of a mismatch,
    uses the value of the mismatched byte to skip across as many
@@ -124,9 +117,9 @@ unsigned char * FindString(unsigned char * BufferPtr,
       BufferPtr += Skip;
    }
 }
-
+ -

LISTING 14.2 L14-2.C

+

LISTING 14.2 L14-2.C

 /* Program to exercise buffer-search routines in Listings 14.1 & 14.3.
    (Must be modified to put copy of pattern as sentinel at end of the
@@ -182,7 +175,7 @@ void main() {
    }
    exit(0);
 }
-
+

Well, architecture carries a lot of weight, but it sure as heck isn’t destiny. I had simply fallen into the trap of figuring that the algorithm was so clever that I didn’t have to do any thinking myself. The path leading to REPNZ SCASB from the original brute-force approach of REPZ CMPSB at every location had been based on my observation that the first character comparison at each buffer location usually fails. Why not apply the same concept to Boyer-Moore? Listing 14.3 is just like the standard implementation—except that it’s optimized to handle a first-comparison mismatch as quickly as possible in the loop at QuickSearchLoop, much as REPNZ SCASB optimizes first-comparison mismatches for the brute-force approach. The results in Table 14.1 speak for themselves; Listing 14.3 is more than twice as fast as what I assure you was already a nice, tight assembly implementation (and unrolling QuickSearchLoop could boost performance by up to 10 percent more). Listing 14.3 is also four times faster than REPNZ SCASB in one case.

@@ -203,10 +196,6 @@ void main() {
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/14-05.html b/14-05.html index 8494ec6..faf4b1a 100644 --- a/14-05.html +++ b/14-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - @@ -37,7 +30,7 @@


-

LISTING 14.3 L14-3.ASM

+

LISTING 14.3 L14-3.ASM

 ; Searches a buffer for a specified pattern. In case of a mismatch,
 ; uses the value of the mismatched byte to skip across as many
@@ -199,7 +192,7 @@ Done:
         ret
 _FindString     endp
         end
-
+


@@ -218,10 +211,6 @@ _FindString endp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/14-06.html b/14-06.html index 78e5ac1..2e6afda 100644 --- a/14-06.html +++ b/14-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Boyer-Moore String Searching - - @@ -41,11 +34,11 @@

Know your data and use your smarts. Don’t stop thinking just because you’re implementing a big-name algorithm; you know more than it does.

-

Further Optimization of Boyer-Moore

+

Further Optimization of Boyer-Moore

We can do substantially better yet than Listing 14.3 if we’re willing to accept tighter limits on the data. Limiting the length of the searched-for pattern to a maximum of 255 bytes allows us to use the XLAT instruction and generally tighten the critical loop. (Be aware, however, that XLAT is a relatively expensive instruction on the 486 and Pentium.) Putting a copy of the searched-for string at the end of the search buffer as a sentinel, so that the search never fails, frees us from counting down the buffer length, and makes it easy to unroll the critical loop. Listing 14.4, which implements these optimizations, is about 60 percent faster than Listing 14.3.

-

LISTING 14.4 L14-4.ASM

+

LISTING 14.4 L14-4.ASM

 ; Searches a buffer for a specified pattern. In case of a mismatch,
 ; uses the value of the mismatched byte to skip across as many
@@ -193,11 +186,11 @@ Done:
         ret
 _FindString     endp
         end
-
+

Note that Table 14.1 includes the time required to build the skip table each time FindString is called. This time could be eliminated for all but the first search when repeatedly searching for a particular pattern, by building the skip table externally and passing a pointer to it as a parameter.

-

Know What You Know

+

Know What You Know

Here we’ve turned up our nose at a repeated string instruction, we’ve gone against the grain by comparing backward, and yet we’ve speeded up our code quite a bit. All this without any restrictions or special requirements (excluding Listing 14.4)—and without any new information. Everything we needed was sitting there all along; we just needed to think to look at it.

@@ -220,10 +213,6 @@ _FindString endp
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/15-01.html b/15-01.html index 451c902..d1e966c 100644 --- a/15-01.html +++ b/15-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - @@ -37,10 +30,10 @@


-

Chapter 15
+

Chapter 15
Linked Lists and plain Unintended Challenges

-

Unfamiliar Problems with Familiar Data Structures

+

Unfamiliar Problems with Familiar Data Structures

After 21 years, this story still makes me wince. Oh, the humiliations I suffer for your enlightenment....

@@ -58,7 +51,7 @@

Maybe you can—but I sure can’t. For example, consider the evolution of my understanding of linked lists.

-

Linked Lists

+

Linked Lists

Linked lists are data structures composed of discrete elements, or nodes, joined together with links. In C, the links are typically pointers. Like all data structures, linked lists have their strengths and their weaknesses. Primary among the strengths are: simplicity; speedy sequential processing; ease and speed of insertion and deletion; the ability to mix nodes of various sizes and types; and the ability to handle variable amounts of data, especially when the total amount of data changes dynamically or is not always known beforehand. Weaknesses include: greater memory requirements than arrays (the pointers take up space); slow non-sequential processing, including finding arbitrary nodes; and an inability to backtrack, unless doubly-linked lists are used. Unfortunately, doubly linked lists need more memory, as well as processing time to maintain the backward links.

@@ -74,9 +67,8 @@

The fundamental problem is that the model of Figure 15.1 unnecessarily complicates link manipulation. In order to delete a node, for example, you must change the preceding node’s NextNode pointer to point to the following node, as shown in Listing 15.1. (Listing 15.2 is the header file LLIST.H, which is #included by all the linked list listings in this chapter.) Easy enough—unless the preceding node happens to be the head pointer, which doesn’t have a NextNode field, because it’s not a node, so Listing 15.1 won’t work. Cumbersome special code and extra information (a pointer to the head of the list) are required to handle the head-pointer case, as shown in Listing 15.3. (I’ll grant you that if you make the next-node pointer the first field in the LinkNode structure, at offset 0, then you could successfully point to the head pointer and pretend it was a LinkNode structure—but that’s an ugly and potentially dangerous trick, and we’ll see a better approach next.)

-


- Figure 15.1
  The basic concept of a linked list.

+


+ Figure 15.1
  The basic concept of a linked list.


@@ -95,10 +87,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/15-02.html b/15-02.html index 736aa58..5e58906 100644 --- a/15-02.html +++ b/15-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - @@ -37,7 +30,7 @@


-

LISTING 15.1 L15-1.C

+

LISTING 15.1 L15-1.C

 /* Deletes the node in a linked list that follows the indicated node.
    Assumes list is headed by a dummy node, so no special testing for
@@ -51,9 +44,9 @@ struct LinkNode *DeleteNodeAfter(struct LinkNode *NodeToDeleteAfter)
          NodeToDeleteAfter->NextNode->NextNode;
    return(NodeToDeleteAfter);
 }
-
+ -

LISTING 15.2 LLIST.H

+

LISTING 15.2 LLIST.H

 /* Linked list header file. */
 #define MAX_TEXT_LENGTH 100   /* longest allowed Text field */
@@ -70,9 +63,9 @@ struct LinkNode *FindNodeBeforeValue(struct LinkNode *, int);
 struct LinkNode *InitLinkedList(void);
 struct LinkNode *InsertNodeSorted(struct LinkNode *,
    struct LinkNode *);
-
+ -

LISTING 15.3 L15-3.C

+

LISTING 15.3 L15-3.C

 /* Deletes the node in the specified linked list that follows the
    indicated node. List is headed by a head-of-list pointer; if the
@@ -92,7 +85,7 @@ struct LinkNode *DeleteNodeAfter(struct LinkNode **HeadOfListPtr,
    }
    return(NodeToDeleteAfter);
 }
-
+

However, it is true that if you’re going to store a variety of types of structures in your linked lists, you should start each node with the LinkNode field. That way, the link pointer is in the same place in every structure, and the same linked list code can handle all of the structure types by casting them to the base link-node structure type. This is a less than elegant approach, but it works. C++ can handle data mixing more cleanly than C, via derivation from a base link-node class.

@@ -100,13 +93,12 @@ struct LinkNode *DeleteNodeAfter(struct LinkNode **HeadOfListPtr,

Similar problems with the head pointer crop up when you’re inserting nodes, and in fact in all link manipulation code. It’s easy to end up working with either pointers to pointers or lots of special-case code, and while those approaches work, they’re inelegant and inefficient.

-

Dummies and Sentinels

+

Dummies and Sentinels

A far better approach is to use a dummy node for the head of the list, as shown in Figure 15.2. I invented this one for myself the next time I encountered linked lists, while designing a seed fill function for MetaWindows, back during my tenure at Metagraphics Corp. But I could have learned it by spending five minutes with Sedgewick’s book.

-


- Figure 15.2
  Using a dummy head and tail node with a linked list.

+


+ Figure 15.2
  Using a dummy head and tail node with a linked list.

@@ -120,7 +112,7 @@ struct LinkNode *DeleteNodeAfter(struct LinkNode **HeadOfListPtr,

Figure 15.3 is a giant step in the right direction, but we can still make a few refinements. The inner loop of any code that scans through such a list has to perform a special test on each node to determine whether the tail has been reached. So, for example, code to find the first node containing a value field greater than or equal to a certain value has to perform two tests in the inner loop, as shown in Listing 15.4.

-

LISTING 15.4 L15-4.C

+

LISTING 15.4 L15-4.C

 /*  Finds the first node in a linked list with a value field greater
     than or equal to a key value, and returns a pointer to the node
@@ -144,13 +136,12 @@ struct LinkNode *FindNodeBeforeValueNotLess(
       return(NodePtr);  /* success; return pointer to node preceding
                            node that was >= */
 }
-
+

Suppose, however, that we make the tail node a sentinel by giving it a value that is guaranteed to terminate the search, as shown in Figure 15.4. The list in Figure 15.4 has a sentinel with a value field of 32,767; since we’re working with integers, that’s the highest possible search value, and is guaranteed to satisfy any search that comes down the pike. The success or failure of the search can then be determined outside the loop, if necessary, by checking for the tail node’s special pointer—but the inside of the loop is streamlined to just one test, as shown in Listing 15.5. Not all linked lists lend themselves to sentinels, but the performance benefits are considerable for those lend themselves to sentinels, but the performance benefits are considerable for those that do.

-


- Figure 15.3
  Representing an empty list.

+


+ Figure 15.3
  Representing an empty list.


@@ -169,10 +160,6 @@ struct LinkNode *FindNodeBeforeValueNotLess(
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/15-03.html b/15-03.html index f36b9be..ee74fef 100644 --- a/15-03.html +++ b/15-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - @@ -37,7 +30,7 @@


-

LISTING 15.5 L15-5.C

+

LISTING 15.5 L15-5.C

 /* Finds the first node in a value-sorted linked list that
    has a Value field greater than or equal to a key value, and
@@ -60,13 +53,12 @@ struct LinkNode *FindNodeBeforeValueNotLess(
       return(NodePtr);  /* success; return pointer to node preceding
                            node that was >= */
 }
-
+ -


- Figure 15.4
  List terminated by a sentinel.

+


+ Figure 15.4
  List terminated by a sentinel.

-

Circular Lists

+

Circular Lists

One minor but elegant refinement yet remains: Use a single node as both the head and the tail of the list. We can do this by connecting the last node back to the first through the head/tail node in a circular fashion, as shown in Figure 15.5. This head/tail node can also, of course, be a sentinel; when it’s necessary to check for the end of the list explicitly, that can be done by comparing the current node pointer to the head pointer. If they’re equal, you’re at the head/tail node.

@@ -78,11 +70,10 @@ struct LinkNode *FindNodeBeforeValueNotLess(

Contrast Figure 15.5 with Figure 15.1, and Listings 15.1, 15.5, 15.6, and 15.7 with Listings 15.3 and 15.4. Yes, linked lists are simple, but not so simple that a little knowledge doesn’t make a substantial difference. Make it a habit to read Knuth or Sedgewick or the like before you write a single line of code.

-


- Figure 15.5
  Representing a circular list.

+


+ Figure 15.5
  Representing a circular list.

-

LISTING 15.6 L15-6.C

+

LISTING 15.6 L15-6.C

 /*  Suite of functions for maintaining a linked list sorted by
     ascending order of the Value field. The list is circular; that
@@ -149,7 +140,7 @@ struct LinkNode *InsertNodeSorted(struct LinkNode *HeadOfListNode,
    NodePtr->NextNode = NodeToInsert;
    return(NodePtr);
 }
-
+


@@ -168,10 +159,6 @@ struct LinkNode *InsertNodeSorted(struct LinkNode *HeadOfListNode,
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/15-04.html b/15-04.html index 6bbb03d..4b5c24e 100644 --- a/15-04.html +++ b/15-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Linked Lists and Unintended Challenges - - @@ -37,7 +30,7 @@


-

LISTING 15.7 L15-7.ASM

+

LISTING 15.7 L15-7.ASM

 ; C near-callable assembly function for inserting a new node in a
 ; linked list sorted by ascending order of the Value field. The list
@@ -98,9 +91,9 @@ SearchLoop:
         ret
 _InsertNodeSorted endp
         end
-
+ -

LISTING 15.8 L15-8.C

+

LISTING 15.8 L15-8.C

 /* Sample linked list program. Tested with Borland C++. */
 #include <stdlib.h>
@@ -183,9 +176,9 @@ void main()
       }
    }
 }
-
+ -

Hi/Lo in 24 Bytes

+

Hi/Lo in 24 Bytes

In one of my PC TECHNIQUES “Pushing the Envelope” columns, I passed along one of David Stafford’s fiendish programming puzzles: Write a C-callable function to find the greatest or smallest unsigned int. Not a big deal—except that David had already done it in 24 bytes, so the challenge was to do it in 24 bytes or less.

@@ -195,7 +188,7 @@ void main()

Yes, a 24-byte hi/lo function is possible, anatomically improbable as it might seem. Which I guess goes to show that when one of David’s puzzles seems less than impossible, odds are you’re missing something. Listing 15.9 is David’s 24-byte solution, from which a lot may be learned if one reads closely enough.

-

LISTING 15.9 L15-9.ASM

+

LISTING 15.9 L15-9.ASM

 ; Find the greatest or smallest unsigned int.
 ; C callable (small model); 24 bytes.
@@ -224,7 +217,7 @@ around:         ja      save
                 jnz     top
 
                 ret
-
+

Before I end this chapter, let me say that I get a lot of feedback from my readers, and it’s much appreciated. Keep those cards, letters, and email messages coming. And if any of you know Jeannie Schweigert, have her drop me a line and let me know how she’s doing these days....

@@ -245,10 +238,6 @@ around: ja save
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-01.html b/16-01.html index 47a2218..3ae504e 100644 --- a/16-01.html +++ b/16-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -37,10 +30,10 @@


-

Chapter 16
+

Chapter 16
There Ain’t No Such Thing as the Fastest Code

-

Lessons Learned in the Pursuit of the Ultimate Word Counter

+

Lessons Learned in the Pursuit of the Ultimate Word Counter

I remember reading an overview of C++ development tools for Windows in a past issue of PC Week. In the lower left corner was the familiar box listing the 10 leading concerns of corporate buyers when it comes to C++. Boiled down, the list looked like this, in order of descending importance to buyers:

@@ -68,7 +61,7 @@

Is something missing here? You bet your maximum gluteus something’s missing—nowhere on that list is there so much as one word about how fast the compiled code runs! I’m not saying that performance is everything, but optimization isn’t even down there at number 10, below online help! Ye gods and little fishes! We are talking here about people who would take a bus from LA to New York instead of a plane because it had a cleaner bathroom; who would choose a painting from a Holiday Inn over a Matisse because it had a fancier frame; who would buy a Yugo instead of—well, hell, anything—because it had a nice owner’s manual and particularly attractive keys. We are talking about people who are focusing on means, and have forgotten about ends. We are talking about people with no programming souls.

-

Counting Words in a Hurry

+

Counting Words in a Hurry

What are we to make of this? At the very least, we can safely guess that very few corporate buyers ever enter optimization contests. Most of my readers do, however; in fact, far more than I thought ever would, but that gladdens me to no end. I issued my first optimization challenge in a “Pushing the Envelope” column in PC TECHNIQUES back in 1991, and was deluged by respondents who, one might also gather, do not live by PC Week.

@@ -128,7 +121,7 @@
-

LISTING 16.1 L16-1.C

+

LISTING 16.1 L16-1.C

  /* Word-counting program. Tested with Borland C++ in C
     compilation mode and the small model. */
@@ -203,7 +196,7 @@
     return(0);
  }
  
-
+


@@ -222,10 +215,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-02.html b/16-02.html index 11429f0..abbba4b 100644 --- a/16-02.html +++ b/16-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -39,7 +32,7 @@

Listing 16.2 is Listing 16.1 modified to call a function that scans each block for words, and Listing 16.3 contains an assembly function that counts words. Used together, Listings 16.2 and 16.3 are just about twice as fast as Listing 16.1, a good return for a little assembly language. Listing 16.3 is a pretty straightforward translation from C to assembly; the new code makes good use of registers, but the key code—determining whether each byte is a character or not—is still done with the same multiple-sequential-tests approach used by the code that the C compiler generates.

-

LISTING 16.2 L16-2.C

+

LISTING 16.2 L16-2.C

  /* Word-counting program incorporating assembly language. Tested
     with Borland C++ in C compilation mode & the small model. */
@@ -99,9 +92,9 @@
     printf(“\nTotal words in file: %lu\n”, WordCount);
     return(0);
  }
-
+ -

LISTING 16.3 L16-3.ASM

+

LISTING 16.3 L16-3.ASM

  ; Assembly subroutine for Listing 16.2. Scans through Buffer, of
  ; length BufferLength, counting words and updating WordCount as
@@ -183,9 +176,9 @@
          ret
  _ScanBuffer     endp
          end
-
+ -

Which Way to Go from Here?

+

Which Way to Go from Here?

We could rearrange the tests in light of the nature of the data being scanned; for example, we could perform the tests more efficiently by taking advantage of the knowledge that if a byte is less than ‘0,’ it’s either an apostrophe or not a character at all. However, that sort of fine-tuning is typically good for speedups of only 10 to 20 percent, and I’ve intentionally refrained from implementing this in Listing 16.3 to avoid pointing you down the wrong path; what we need is a different tack altogether. Ponder this. What we really want to know is nothing more than whether a byte is a character, not what sort of character it is. For each byte value, we want a yes/no status, and nothing else—and that description practically begs for a lookup table. Listing 16.4 uses a lookup table approach to boost performance another 50 percent, to three times the performance of the original C code. On a 20 MHz 386, this represents a change from 4.6 to 1.6 seconds, which could be significant—who likes to wait? On an 8088, the improvement in word-counting a large file could easily be 10 or 20 seconds, which is definitely significant.

@@ -206,10 +199,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-03.html b/16-03.html index 3f01b19..4807d1a 100644 --- a/16-03.html +++ b/16-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -37,7 +30,7 @@


-

LISTING 16.4 L16-4.ASM

+

LISTING 16.4 L16-4.ASM

  ; Assembly subroutine for Listing 16.2. Scans through Buffer, of
  ; length BufferLength, counting words and updating WordCount as
@@ -130,7 +123,7 @@
  _ScanBuffer     endp
          end
  
-
+

Listing 16.4 features several interesting tricks. First, it uses LODSB and XLAT in succession, a very neat way to get a pointed-to byte, advance the pointer, and look up the value indexed by the byte in a table, all with just two instruction bytes. (Interestingly, Listing 16.4 would probably run quite a bit better still on an 8088, where LODSB and XLAT have a greater advantage over conventional instructions. On the 486 and Pentium, however, LODSB and XLAT lose much of their appeal, and should be replaced with MOV instructions.) Better yet, LODSB and XLAT don’t alter the flags, so the Zero flag status set before LODSB is still around to be tested after XLAT .

@@ -146,7 +139,7 @@ -

Challenges and Hazards

+

Challenges and Hazards

The challenge I put to the readers of PC TECHNIQUES was to write a faster module to replace Listing 16.4. The author of the code that counted the words in my secret test file fastest on my 20 MHz cached 386 would be the winner and receive Numerous Valuable Prizes.

@@ -171,10 +164,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-04.html b/16-04.html index fdae835..3a12c77 100644 --- a/16-04.html +++ b/16-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -39,18 +32,18 @@

Truth to tell, I didn’t expect a three-times speedup; around two times was what I had in mind. Which just goes to show that any code can be made faster than you’d expect, if you think about it long enough and from many different perspectives. (The most potent word-counting technique seems to be a 64K lookup table that allows handling two bytes simultaneously. This is not the sort of technique one comes up with by brute-force optimization.) Thinking (or, worse yet, boasting) that your code is the fastest possible is rollescating on a tightrope in a hurricane; you’re due for a fall, if you catch my drift. Case in point: Terje Mathisen’s word-counting program.

-

Blinding Yourself to a Better Approach

+

Blinding Yourself to a Better Approach

Not so long ago, Terje Mathisen, who I introduced earlier in this book, wrote a very fast word-counting program, and posted it on Bix. When I say it was fast, I mean fast; this code was optimized like nobody’s business. We’re talking top-quality code here.

When the topic of optimizing came up in one of the Bix conferences, Terje’s program was mentioned, and he posted the following message: “I challenge BIXens (and especially mabrash!) to speed it up significantly. I would consider 5 percent a good result.” The clear implication was, “That code is as fast as it can possibly be.”

-

Naturally, it wasn’t; there ain’t no such thing as the fastest code (TANSTATFC? I agree, it doesn’t have the ring of TANSTAAFL). I pored over Terje’s 386 native-mode code, and found the critical inner loop, which was indeed as tight as one could imagine, consisting of just a few 386 native-mode instructions. However, one of the instructions was this:

+

Naturally, it wasn’t; there ain’t no such thing as the fastest code (TANSTATFC? I agree, it doesn’t have the ring of TANSTAAFL). I pored over Terje’s 386 native-mode code, and found the critical inner loop, which was indeed as tight as one could imagine, consisting of just a few 386 native-mode instructions. However, one of the instructions was this:

  
  CMP   DH,[EBX+EAX]
  
-
+

Harmless enough, save for two things. First, EBX happened to be zero at this point (a leftover from an earlier version of the code, as it turned out), so it was superfluous as a memory-addressing component; this made it possible to use base-only addressing ([EAX]) rather than base+index addressing ([EBX+EAX]), which saves a cycle on the 386. Second: Changing the instruction to CMP [EAX],DH saved 2 cycles—just enough, by good fortune, to speed up the whole program by 5 percent.

@@ -64,7 +57,7 @@

(Granted, CMP [mem],reg is 1 cycle slower than CMP reg,[mem] on the 286, and they’re both the same on the 8088; in this case, though, the code was specific to the 386. In case you’re curious, both forms take 2 cycles on the 486; quite a lot faster, eh?)

-

Watch Out for Luggable Assumptions!

+

Watch Out for Luggable Assumptions!

The first lesson to be learned here is not to lug assumptions that may no longer be valid from the 8088/286 world into the wonderful new world of 386 native-mode programming. The second lesson is that after you’ve slaved over your code for a while, you’re in no shape to see its flaws, or to be able to get the new perspectives needed to speed it up. I’ll bet Terje looked at that [EBX+EAX] addressing a hundred times while trying to speed up his code, but he didn’t really see what it did; instead, he saw what it was supposed to do. Mental shortcuts like this are what enable us to deal with the complexities of assembly language without overloading after about 20 instructions, but they can be a major problem when looking over familiar code.

@@ -82,7 +75,7 @@

By the way, Terje’s WC50 program is a full-fledged counting program; it counts characters, words, and lines, can handle multiple files, and lets you specify the characters that separate words, should you so desire. Source code is provided as part of the archive WC50 comes in. All in all, it’s a nice piece of work, and you might want to take a look at it if you’re interested in really fast assembly code. I wouldn’t call it the fastest word-counting code, though, because I would of course never be so foolish as to call anything the fastest.

-

The Astonishment of Right-Brain Optimization

+

The Astonishment of Right-Brain Optimization

As it happened, the challenge I issued to my PC TECHNIQUES readers was a smashing success, with dozens of good entries. I certainly enjoyed it, even though I did have to look at a lot of tricky assembly code that I didn’t write—hard work under the best of circumstances. It was worth the trouble, though. The winning entry was an astonishing example of what assembly language can do in the right hands; on my 386, it was four times faster at word counting than the nice, tight assembly code I provided as a starting point—and about 13 times faster than the original C implementation. Attention, high-level language chauvinists: Is the speedup getting significant yet? Okay, maybe word counting isn’t the most critical application, but how would you like to have that kind of improvement in your compression software, or in your real-time games—or in Windows graphics?

@@ -105,10 +98,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-05.html b/16-05.html index fa56885..1a70035 100644 --- a/16-05.html +++ b/16-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -165,7 +158,7 @@ -

LISTING 16.5 QSCAN3.ASM

+

LISTING 16.5 QSCAN3.ASM

  ;  QSCAN3.ASM
  ;  David Stafford
@@ -345,7 +338,7 @@ jumping.
                  .fardata WordTable
  include         qscan3.inc                ;built by MAKETAB
                  end
-
+


@@ -364,10 +357,6 @@ jumping.
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-06.html b/16-06.html index 2d637ab..57ca1a6 100644 --- a/16-06.html +++ b/16-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -37,7 +30,7 @@


-

Levels of Optimization

+

Levels of Optimization

Three levels of optimization were evident in the word-counting entries I received in response to my challenge. I’d briefly describe them as “fine-tuning,” “new perspective,” and “table-driven state machine.” The latter categories produce faster code, but, by the same token, they are harder to design, harder to implement, and more difficult to understand, so they’re suitable for only the most demanding applications. (Heck, I don’t even guarantee that David Stafford’s entry works perfectly, although, knowing him, it probably does; the more complex and cryptic the code, the greater the chance for obscure bugs.)

@@ -49,7 +42,7 @@ -

Optimization Level 1: Good Code

+

Optimization Level 1: Good Code

The first level of optimization involves fine-tuning and clever use of the instruction set. The basic framework is still the same as my code (which in turn is basically the same as that of the original C code), but that framework is implemented more efficiently.

@@ -80,10 +73,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-07.html b/16-07.html index a3d39d0..62f386e 100644 --- a/16-07.html +++ b/16-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -37,7 +30,7 @@


-

Listing 16.6 OPT2.ASM

+

Listing 16.6 OPT2.ASM

  ;
  ;          Opt2         Final optimization word count
@@ -143,9 +136,9 @@
                 ret
  _ScanBuffer    endp
                 end
-
+ -

Level 2: A New Perspective

+

Level 2: A New Perspective

The second level of optimization is one of breaking out of the mode of thinking established by my original code. Some entrants clearly did exactly that. They stepped back, thought about what the code actually needed to do, rather than just improving how it already worked, and implemented code that sprang from that new perspective.

@@ -176,10 +169,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/16-08.html b/16-08.html index 1424fbe..e793435 100644 --- a/16-08.html +++ b/16-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: There Ain't No Such Thing as the Fastest Code - - @@ -37,7 +30,7 @@


-

Listing 16.7 L16-7.ASM

+

Listing 16.7 L16-7.ASM

  ScanLoop:
          lodsw           ;get the next 2 bytes (AL = first, AH = 2nd)
@@ -53,13 +46,13 @@
          dec     dx
          jnz     ScanLoop
  
-
+

John later divides the transition count by two to get the word count. (Food for thought: It’s also possible to use CMP and ADC to detect words without branching.)

John’s approach makes it clear that word-counting is nothing more than a fairly simple state machine. The interesting part, of course, is building the fastest state machine.

-

Level 3: Breakthrough

+

Level 3: Breakthrough

The boundaries between the levels of optimization are not sharply defined. In a sense, level 3 optimization is just like levels 1 and 2, but more so. At level 3, one takes whatever level 2 perspective seems most promising, and implements it as efficiently as possible on the x86. Even more than at level 2, at level 3 this means breaking out of familiar patterns of thinking.

@@ -67,7 +60,7 @@

The key concept at level 3 is the use of a massive (64K) lookup table that processes byte sequences directly into word-count actions. With such a table, it’s possible to look up the appropriate action for two bytes simultaneously in just a few instructions; next, I’m going to look at the inspired and highly unusual way that David’s code, shown in Listing 16.5, does exactly that. (Before assembling Listing 16.5, you must run the C code in Listing 16.8, to generate an include file defining the 64K lookup table. When you assemble Listing 16.5, TASM will report a “location counter overflow” warning; ignore it.)

-

LISTING 16.8 MAKETAB.C

+

LISTING 16.8 MAKETAB.C

  //  MAKETAB.C — Build QSCAN3.INC for QSCAN3.ASM
   
@@ -102,27 +95,25 @@
    fclose( t );
    }
  
-
+

David’s approach is simplicity itself, although his implementation arguably is not. Consider any three sequential bytes in the buffer. Those three bytes define two potential places where a word might be counted, as shown in Figure 16.1. Given the separator/non-separator states of the three bytes, you can instantly determine whether to count a word or not; you count a word if and only if somewhere in the sequence there is a non-separator followed by a separator. Note that a maximum of one word can be counted per three-byte sequence.

The trick, then, is to identify the separator/not statuses of each set of three bytes and turn them into a 1 (count word) or 0 (don’t count word), as quickly as possible. Assuming that the separator/not status for the first byte is in the Carry flag, this is easily accomplished by a lookup in a 64K table, based on the Carry flag and the other two bytes, as shown in Figure 16.2. (Remember that we’re counting 7-bit ASCII here, so the high bit is ignored.) Thus, David is able to add the word/not status for each pair of bytes to the main word count simply by getting the two bytes, working in the carry status from the last byte, and using the resulting value to index into the 64K table, adding in the 1 or 0 value found in that table. A sequence of MOV/ADC/ADD suffices to perform all word-counting tasks for a pair of bytes. Three instructions, no branches—pretty nearly perfect code.

-


- Figure 16.1
  The two potential word count locations.

+


+ Figure 16.1
  The two potential word count locations.

One detail remains to be attended to: setting the Carry flag for next time if the last byte was a non-separator. David does this in a bizarre and incredibly effective way: He presets the high bit of the count, and sets the high bit in the lookup table for those entries looked up by non-separators. When a non-separator’s lookup entry is added to the count, it will produce a carry, as desired. The high bit of the count is masked off before being added to the total count, so David is essentially using different parts of the count variables for different purposes (counting, and setting the Carry flag).

-


- Figure 16.2
  Looking up a word count status.

+


+ Figure 16.2
  Looking up a word count status.

There are a number of other interesting details in David’s code, including the unrolling of the loop 64 times, so that 256 bytes in a row are processed without a single branch. Unfortunately, I lack the space to discuss Listing 16.5 any further. Perhaps that’s not so unfortunate, after all; I’d hate to deny you the pleasure of discovering the wonders of this rather remarkable code yourself. I will say one more thing, though. The cycle count for David’s inner loop is 6.5 cycles per byte processed, and the actual measured time for his routine, overhead and all, is 7.9 cycles/byte. The original C code clocked in at around 100 cycles/byte.

Enough said, I trust.

-

Enough Word Counting Already!

+

Enough Word Counting Already!

Before I finish up this chapter, I’d like to mention that Terje Mathisen’s WC word-counting program, which I’ve mentioned previously and which is available, with source, on Bix, is in the ballpark with David’s code for performance. What’s more, Terje’s program handles 8-bit ASCII, counts lines as well as words, and supports user-definable separator sets. It’s wonderful code, well worth a look; it also happens to be a great word-counting utility. By the way, Terje builds his 64K table on the fly, at program initialization; this allows for customized tables, shrinks the size of the EXE, and, according to Terje’s calculations, takes less time than loading the table off disk as part of the EXE.

@@ -145,10 +136,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/17-01.html b/17-01.html index 84100d9..2b8b55b 100644 --- a/17-01.html +++ b/17-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -37,16 +30,16 @@


-

Chapter 17
+

Chapter 17
The Game of Life

-

The Triumph of Algorithmic Optimization in a Cellular Automata Game

+

The Triumph of Algorithmic Optimization in a Cellular Automata Game

I’ve spent a lot of my life discussing assembly language optimization, which I consider to be an important and underappreciated topic. However, I’d like to take this opportunity to point out that there is much, much more to optimization than assembly language. Assembly is essential for absolute maximum performance, but it’s not the only ingredient; necessary but not sufficient, if you catch my drift—and not even necessary, if you’re looking for improved but not maximum performance. You’ve heard it a thousand times: Optimize your algorithm first. Devise new approaches. Or, as Knuth said, Premature optimization is the root of all evil.

This is, of course, old hat, stuff you know like the back of your hand. Or is it? As Jeff Duntemann pointed out to me the other day, performance programmers are made, not born. While I’m merrily gallivanting around in this book optimizing 486 pipelining and turning simple tasks into horribly complicated and terrifyingly fast state machines, many of you are still developing your basic optimization skills. I don’t want to shortchange those of you in the latter category, so in this chapter, we’ll discuss some high-level language optimizations that can be applied by mere mortals within a reasonable period of time. We’re going to examine a complete optimization process, from start to finish, and what we will find is that it’s possible to get a 50-times speed-up without using one byte of assembly! It’s all a matter of perspective—how you look at your code and data.

-

Conway’s Game

+

Conway’s Game

The program that we’re going to optimize is Conway’s famous Game of Life, long-ago favorite of the hackers at MIT’s AI Lab. If you’ve never seen it, let me assure you: Life is neat, and more than a little hypnotic. Fractals have been the hot graphics topic in recent years, but for eye-catching dazzle, Life is hard to beat.

@@ -54,7 +47,7 @@

First, I’ll describe the ground rules of Life, implement a very straightforward version in C++, and then speed that version up by about eight times without using any drastically different approaches or any assembly. This may be a little tame for some of you, but be patient; for after that, we’ll haul out the big guns and move into the 30 to 40 times speed-up range. Then in the next chapter, I’ll show you how several programmers really floored it in taking me up on my second Optimization Challenge, which involved the Game of Life.

-

The Rules of the Game

+

The Rules of the Game

The Game of Life is ridiculously simple. There is a cellmap, consisting of a rectangular matrix of cells, each of which may initially be either on or off. Each cell has eight neighbors: two horizontally, two vertically, and four diagonally. For each succeeding generation of cells, the game logic determines whether each cell will be on or off according to the following rules:

@@ -85,10 +78,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/17-02.html b/17-02.html index f10e948..fa518f6 100644 --- a/17-02.html +++ b/17-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -37,7 +30,7 @@


-

LISTING 17.1 L17-1.CPP

+

LISTING 17.1 L17-1.CPP

 /* C++ Game of Life implementation for any mode for which mode set
    and draw pixel functions can be provided.
@@ -233,9 +226,9 @@ void cellmap::next_generation(cellmap& next_map)
       }
    }
 }
-
+ -

LISTING 17.2 L17-2.CPP

+

LISTING 17.2 L17-2.CPP

 /* VGA mode 13h functions for Game of Life.
    Tested with Borland C++. */
@@ -293,7 +286,7 @@ void show_text(int x, int y, char *text)
    gotoxy(TEXT_X_OFFSET + x, y);
    puts(text);
 }
-
+


@@ -312,10 +305,6 @@ void show_text(int x, int y, char *text)
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/17-03.html b/17-03.html index 35b050f..f9cf8ec 100644 --- a/17-03.html +++ b/17-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -37,7 +30,7 @@


-

Where Does the Time Go?

+

Where Does the Time Go?

How slow is Listing 17.1? Table 17.1 shows that even on a 486, Listing 17.1 does fewer than three 96x96 generations per second. (The times in Table 17.1 are for 1,000 generations of a 96x96 cell map with seed=1, LIMIT_18_HZ=0, WRAP_EDGES=1, and magnifier=2, running on a 33 MHz 486.) Since my target is 18 generations per second with a 200x200 cellmap on a 20 MHz 386, Listing 17.1 is too slow by a rather wide margin—75 times too slow, in fact. You might say we have a little optimizing to do.

@@ -173,7 +166,7 @@ -

The Hazards and Advantages of Abstraction

+

The Hazards and Advantages of Abstraction

How can we speed up cell_state() and next_generation()? I’ll tell you how not to do it: By writing those member functions in assembly. It’s tempting to say that cell_state() is taking all the time, so we need to speed it up with assembly, but what we really need to do is figure out why cell_state() is taking all the time, then address that aspect of the program directly.

@@ -206,10 +199,6 @@
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/17-04.html b/17-04.html index a1bee60..7e9e487 100644 --- a/17-04.html +++ b/17-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -39,19 +32,17 @@

There’s a kicker here, though, and that’s the counting of neighbors for cells at the edge of the cellmap. When cellmap wrapping is enabled (so that the cellmap becomes essentially a toroid, with each edge joined seamlessly to the opposite edge, as opposed to having a border of off-cells), neighbors that reside on the other edge of the cellmap can’t be accessed by the standard fixed offset, as shown in Figure 17.1. So, in general, we could improve performance by hard-wiring our neighbor-counting for the bit-per-cell cellmap format, but it seems we’d need a lot of conditional code to handle wrapping, and that would slow things back down again.

-


- Figure 17.1
  Edge-wrapping complications.

+


+ Figure 17.1
  Edge-wrapping complications.

When a problem doesn’t lend itself well to optimization, make it a practice to see if you can change the problem definition to one that allows for greater efficiency. In this case, we’ll change the problem by putting padding bytes around the edge of the cellmap, and duplicating each edge of the cellmap in the padding bytes at the opposite side, as shown in Figure 17.2. That way, a hard-wired neighbor count will find exactly what it should—the opposite edge—without any special code at all.

But doesn’t that extra copying of the edges take time? Sure, but only a little; we can build it into the cellmap copying function, and then frankly we won’t even notice it. Avoiding tens or hundreds of thousands of calls to cell_state(), on the other hand, will be very noticeable. Listing 17.3 shows the alterations to Listing 17.1 required to implement a hard-wired neighbor-counting function. This is a minor change, in truth, implemented in about half an hour and not making the code significantly larger—but Listing 17.3 is 3.6 times faster than Listing 17.1, as shown in Table 17.1. We’re up to about 10 generations per second on a 486; not where we want to be, but it is a vast improvement.

-


- Figure 17.2
  The “padding cells” solution.

+


+ Figure 17.2
  The “padding cells” solution.

-

LISTING 17.3 L17-3.CPP

+

LISTING 17.3 L17-3.CPP

 /* cellmap class definition, constructor, copy_cells(), set_cell(),
    clear_cell(), cell_state(), count_neighbors(), and
@@ -214,7 +205,7 @@ void cellmap::next_generation(cellmap& next_map)
       }
    }
 }
-
+


@@ -233,10 +224,6 @@ void cellmap::next_generation(cellmap& next_map)
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/17-05.html b/17-05.html index 01912dd..128a7a7 100644 --- a/17-05.html +++ b/17-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -43,11 +36,11 @@

Not hardly.

-

Heavy-Duty C++ Optimization

+

Heavy-Duty C++ Optimization

Before we get to assembly, we still have to perform C++ optimization, then see if we can find an alternative approach that better fits the application. It would actually have made much more sense if we had looked for a new approach as our first optimization step, but I decided it would be better to cover straightforward C++ optimizations at this point, and the mind-bending stuff a little later. Right now, let’s look at some C++ optimizations; Listing 17.4 is a C++-optimized version of Listing 17.3.

-

LISTING 17.4 L17-4.CPP

+

LISTING 17.4 L17-4.CPP

 /* next_generation(), implemented using fast, all-in-one hard-wired
    neighbor count/update/draw function. Otherwise, the same as
@@ -129,7 +122,7 @@ neighbor_count++;
       row_cell_ptr += width_in_bytes;  // point to start of next row
    }
 }
-
+

Listing 17.4 and Listing 17.3 are functionally the same; the only difference lies in how next_generation() is implemented. (Only next_generation() is shown in Listing 17.4; the program is otherwise identical to Listing 17.3.) Listing 17.4 applies the following optimizations to next_generation():

@@ -164,10 +157,6 @@ neighbor_count++;
Graphics Programming Black Book © 2001 Michael Abrash -
- - - - + diff --git a/17-06.html b/17-06.html index ed8d766..5db9fd1 100644 --- a/17-06.html +++ b/17-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -49,7 +42,7 @@
  • Cells change state relatively infrequently.
  • -

    Bringing In the Right Brain

    +

    Bringing In the Right Brain

    In the previous section, we saw how a C++ program could be sped up about eight times simply by rearranging the data and code in straightforward ways. Now we’re going to see how right-brain non-linear optimization can speed things up by another four times—and make the code simpler.

    @@ -57,7 +50,7 @@

    I have two objectives to achieve in the remainder of this chapter. First, I want to show that optimization consists of many levels, from assembly language up to conceptual design, and that assembly language kicks in pretty late in the optimization process. Second, I want to encourage you to saturate your brain with everything you know about any particular optimization problem, then make space for your right brain to solve the problem.

    -

    Re-Examining the Task

    +

    Re-Examining the Task

    Earlier in this chapter, we looked at a straightforward Game of Life implementation, then increased performance considerably by making the implementation a little less abstract and a little less general. We made a small change to the cellmap format, adding padding bytes off the edges so that pointer arithmetic would always work, but the major optimizations were moving the critical code into a single loop and using pointers rather than member functions whenever possible. In other words, we took what we already knew and made it more efficient.

    @@ -79,11 +72,10 @@

    Know your data.

    -


    - Figure 17.3
      New cell format.

    +


    + Figure 17.3
      New cell format.

    -

    Acting on What We Know

    +

    Acting on What We Know

    Once we’ve changed the cellmap format to store neighbor counts as well as states, with a byte for each cell, we can get another performance boost by again examining what we know about our data. I said earlier that most cells are off during any given generation. This means that most cells have no neighbors that are on. Since the cell map representation for an off-cell that has no neighbors is a zero byte, we can skip over scads of unchanged cells at a pop simply by scanning for non-zero bytes. This is much faster than explicitly testing cell states and neighbor counts, and lends itself beautifully to assembly language implementation as REPZ SCASB or (with a little cleverness) REPZ SCASW. (Unfortunately, there’s no C library function that can scan memory for the next byte that’s non-zero.)

    @@ -106,10 +98,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/17-07.html b/17-07.html index 016e37b..c59dd1f 100644 --- a/17-07.html +++ b/17-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -37,7 +30,7 @@


    -

    LISTING 17.5 L17-5.CPP

    +

    LISTING 17.5 L17-5.CPP

     /* C++ Game of Life implementation for any mode for which mode set
        and draw pixel functions can be provided. The cellmap stores the
    @@ -312,7 +305,7 @@ void cellmap::init()
           }
        } while (—init_length);
     }
    -
    +


    @@ -331,10 +324,6 @@ void cellmap::init()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/17-08.html b/17-08.html index 4cffb76..fe608b2 100644 --- a/17-08.html +++ b/17-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Game of Life - - @@ -55,7 +48,7 @@

    No doubt we could get another two to five times improvement with good assembly code—but that’s dwarfed by a 30-times improvement, so optimization at a conceptual level must come first.

    -

    The Challenge That Ate My Life

    +

    The Challenge That Ate My Life

    The most recent optimization challenge I laid my community of readers was to write the fastest possible Game of Life generation engine. By “engine” I meant that I didn’t care about time spent in input or output, only time consumed by the call to next-generation. The time spent updating the cellmap was what I wanted people to concentrate on.

    @@ -96,10 +89,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/18-01.html b/18-01.html index 7b509c1..d1a5594 100644 --- a/18-01.html +++ b/18-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - @@ -37,10 +30,10 @@


    -

    Chapter 18
    +

    Chapter 18
    It’s a plain Wonderful Life

    -

    Optimization beyond the Pale

    +

    Optimization beyond the Pale

    When I was in high school, my gym teacher had us run a race around the soccer field, or rather, around a course marked with cones that roughly outlined the shape of the field. I quickly settled into second place behind Dwight Chamberlin. We cruised around the field, and when we came to the far corner, Dwight cut across the corner, inside a cone placed awkwardly far out from the others. I followed, and everyone else cut inside the cone too—except the pear-shaped kid bringing up the rear, who plodded his way around every single cone on his way to finishing about half a lap behind. When the laggard finally crossed the finish line, the coach named him the winner, to my considerable irritation. After all, the object was to see who could run the fastest, wasn’t it?

    @@ -56,7 +49,7 @@ -

    Breaking the Rules

    +

    Breaking the Rules

    The other reason for the anecdote has to do with the way my second Optimization Challenge worked itself out. If you’ll recall from the last chapter, the challenge I made to the readers of PC TECHNIQUES was to devise the fastest possible version of the Game of Life cellular automata simulation game. I gave an example, laid out the rules, and stood aside. Good thing, too. Apres moi, le deluge....

    @@ -87,10 +80,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/18-02.html b/18-02.html index 59bcf74..9524e99 100644 --- a/18-02.html +++ b/18-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - @@ -37,7 +30,7 @@


    -

    Table-Driven Magic

    +

    Table-Driven Magic

    David Stafford won my first Optimization Challenge by means of a huge look-up table and an incredible state machine driven by that table. The table didn’t cause David’s entry to exceed the line limit because David’s submission included code to generate the table on the fly as part of the build process. David has done himself one better this time with his QLIFE program; not only does his build process generate a 64K table, but it also generates virtually all his code, consisting of 17,000-plus lines of assembly language spanning another 64K. What David has done is write the equivalent of a bitblt compiler for the Game of Life; one might in fact call it a Life compiler. What David’s code generates is still a general-purpose program; it takes arbitrary seed values, and can run for an arbitrary number of generations, so it’s not as if David simply hardwired the instructions to draw each successive screen. However, it’s a general-purpose program that is exquisitely tailored to the task it needs to perform.

    @@ -120,10 +113,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/18-03.html b/18-03.html index 61d31f5..30d69ad 100644 --- a/18-03.html +++ b/18-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - @@ -37,15 +30,15 @@


    -

    LISTING 18.1 BUILD.BAT

    +

    LISTING 18.1 BUILD.BAT

     bcc -v -D%1=%2;%2=%3;%3=%4;%4=%5;%5=%6;%6=%7;%7=%8;%8 lcomp.c
     lcomp > qlife.asm
     tasmx /mx /kh30000 qlife
     bcc -v -D%1=%2;%2=%3;%3=%4;%4=%5;%5=%6;%6=%7;%7=%8;%8 qlife.obj main.c video.c
    -
    + -

    LISTING 18.2 LCOMP.C

    +

    LISTING 18.2 LCOMP.C

     // LCOMP.C
     //
    @@ -540,7 +533,7 @@ void main( void )
     
       printf( “LIFE ends\nend\n” );
       }
    -
    +


    @@ -559,10 +552,6 @@ void main( void )
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/18-04.html b/18-04.html index 63ceabd..7bdd8b2 100644 --- a/18-04.html +++ b/18-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - @@ -37,7 +30,7 @@


    -

    LISTING 18.3 MAIN.C

    +

    LISTING 18.3 MAIN.C

     // MAIN.C
     //
    @@ -129,9 +122,9 @@ void main( void )
       printf( “Time: %f generations/second\n”,
               (double)generation / (double)end_time * 18.2 );
       }
    -
    + -

    LISTING 18.4 VIDEO.C

    +

    LISTING 18.4 VIDEO.C

     /* VGA mode 13h functions for Game of Life.
        Tested with Borland C++. */
    @@ -169,9 +162,9 @@ void show_text(int x, int y, char *text)
        gotoxy(TEXT_X_OFFSET + x, y);
        puts(text);
     }
    -
    + -

    LISTING 18.5 LIFE.H

    +

    LISTING 18.5 LIFE.H

     void far NextGen( void );
     
    @@ -190,9 +183,9 @@ extern unsigned short far ChangeList1[];
     #define WRAPRIGHT   (LEFT  * (WIDTH - 1))
     #define WRAPUP      (DOWN  * (HEIGHT - 1))
     #define WRAPDOWN    (UP    * (HEIGHT - 1))
    -
    + -

    Keeping Track of Change with a Change List

    +

    Keeping Track of Change with a Change List

    In my earlier optimizations to the Game of Life, described in the last chapter, I noted that most cells in a Life cellmap are dead, and in most cases all the neighbors are dead as well. This observation enabled me to get a major speed-up by scanning the cellmap for the few non-zero bytes (cells that were either alive or have neighbors that are alive). Although that was a big improvement, it still required my code to touch every cell to check its state. David has improved on this by maintaining a change list; that is, a list of pointers to cells that change in the current generation. Only those cells and their neighbors need to be checked or touched in any way in order to create the next generation, saving a great many instructions and also a great many cache misses due to the fact that cellmaps are too big to fit into the 486’s internal cache. During a given generation, David runs down the list of cells that changed from the previous generation to make the changes for this generation, and in the process generates the change list for the next generation.

    @@ -202,9 +195,8 @@ extern unsigned short far ChangeList1[];

    “Since every cell has from zero to eight neighbors, you may be wondering how I can manage to keep track of them with only three bits. Each cell really has only a maximum of seven neighbors since we only need to keep track of neighbors outside of the current cell word. That is, if cell ‘B’ changes state then we don’t need to reflect this in the neighbor counts of cells ‘A’ and ‘C.’ Updating is made a little faster. [In other words, when David picks up a word representing three cells, each of the three cells has at least one of the other cells in that word as a neighbor, and the state of that neighbor is stored right in that word, as shown in Figure 18.1. Therefore, the neighbor count for a given cell never needs to reflect more than seven neighbors, because at least one of the eight neighbors’ states is already encoded in the word.]

    -


    - Figure 18.1
      Cell triplet storage.

    +


    + Figure 18.1
      Cell triplet storage.


    @@ -223,10 +215,6 @@ extern unsigned short far ChangeList1[];
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/18-05.html b/18-05.html index 475d779..5357767 100644 --- a/18-05.html +++ b/18-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: It's a Wonderful Life - - @@ -45,7 +38,7 @@

    Segment usage in David’s assembly code is summarized in Listing 18.6.

    -

    LISTING 18.6 QLIFE Assembly Segment Usage

    +

    LISTING 18.6 QLIFE Assembly Segment Usage

     CS : 64K code (table of routines on 256 byte boundaries)
     DS : DGROUP (1st pass) / 64K cell life/death classification table (second pass)
    @@ -53,18 +46,18 @@ ES : Change list
     SS : DGROUP; the life cell grid and row/column table
     FS : Video segment
     GS : Unused
    -
    + -

    A Layperson’s Overview of QLIFE

    +

    A Layperson’s Overview of QLIFE

    Most likely, you’re scratching your head right now in bemusement. I don’t blame you; I felt the same way myself at first. It’s actually pretty simple, though, once you have the hang of it. Basically, David runs down the change list, visiting every cell that’s due to change in this generation, setting it to the new state, drawing it in the new state, and adjusting the counts of all its neighbors. David has a separate assembly routine for every possible change of state for a cell triplet, and he jumps to the proper routine by taking the cell triplet word, masking off the lower 9 bits, and jumping to the address where the appropriate code to perform that particular change of state resides. He does this for every entry in the change list. When this is completed, the current generation has been drawn and updated.

    -

    Now David runs down the change list again to generate the change list for the next generation. In this case, for every changed cell triplet, David looks at that triplet and all affected neighbors to see which will change in the next generation. He tests for this condition by using each potentially changed cell triplet word as an index into the aforementioned lookup table of new states. If the current state matches the appropriate state for the next generation, then there’s nothing to do and the cell is not added to the change list. If the states don’t match, then the cell is added to the change list, and the appropriate state for the next generation is set in the cell triplet. David checks the minimum possible number of cells for change by branching to code that checks only the relevant cells around each cell triplet in the current change list; that branching is accomplished by taking the cell triplet word, masking off the lower 9 bits, setting bit 8 to a 1-bit, and branching to the routine at that address. As with everything in this amazing program, this represents the least possible work to accomplish the desired result—just three instructions:

    +

    Now David runs down the change list again to generate the change list for the next generation. In this case, for every changed cell triplet, David looks at that triplet and all affected neighbors to see which will change in the next generation. He tests for this condition by using each potentially changed cell triplet word as an index into the aforementioned lookup table of new states. If the current state matches the appropriate state for the next generation, then there’s nothing to do and the cell is not added to the change list. If the states don’t match, then the cell is added to the change list, and the appropriate state for the next generation is set in the cell triplet. David checks the minimum possible number of cells for change by branching to code that checks only the relevant cells around each cell triplet in the current change list; that branching is accomplished by taking the cell triplet word, masking off the lower 9 bits, setting bit 8 to a 1-bit, and branching to the routine at that address. As with everything in this amazing program, this represents the least possible work to accomplish the desired result—just three instructions:

     mov dh,[bp+1]
     or dh,1
     jmp dx
    -
    +

    These suffice to select the proper, minimum-work code to process the next cell triplet that has changed, and all potentially affected neighbors. For all the size of David’s code, it has an astonishing economy of effort, as execution glides through the change list without a wasted instruction.

    @@ -89,10 +82,6 @@ jmp dx
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/19-01.html b/19-01.html index e2d31d1..2a31923 100644 --- a/19-01.html +++ b/19-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - @@ -37,10 +30,10 @@


    -

    Chapter 19
    +

    Chapter 19
    Pentium: Not the Same Old Song

    -

    Learning a Whole Different Set of Optimization Rules

    +

    Learning a Whole Different Set of Optimization Rules

    I can still remember the day I did my first 8088 programming. I had just moved over from the distantly related Z80, so the 8088 wasn’t totally alien, but it was nonetheless an incredibly exciting processor. The 8088’s instruction set was vastly more powerful and varied than the Z80’s, and as someone who thrives on puzzles of all sorts, from crosswords to Freecell to jigsaws to assembly language optimization, I was delighted to find that the 8088 made the optimization universe an order of magnitude more complicated—and correspondingly more interesting.

    @@ -48,7 +41,7 @@

    Happily, the 486 traveled to the beat of a different drum. The 486 had some interesting internal pipeline hazards, as well as an internal cache that made cycle counting more meaningful than ever before, and careful code massaging sometimes yielded startling results. Nonetheless, the 486 was still too simple to mark a return to the golden age of optimization.

    -

    The Return of Optimization as Art

    +

    The Return of Optimization as Art

    Then the Pentium came around, and filled our code with optimization hazards, and life was good again. The Pentium has two execution pipelines and enough rules and exceptions to those rules to bring joy to the heart of the hardest-core assembly junkie. For a change, Intel documented most of the Pentium optimization rules and spread the word about them, so we don’t have to go through as much spelunking of the Pentium as with its predecessors. They’ve done this, I suspect, largely because more than any previous x86 processor, the Pentium’s performance is highly dependent on properly optimized code.

    @@ -60,7 +53,7 @@

    Gimme a “P”....

    -

    The Pentium: An Overview

    +

    The Pentium: An Overview

    Architecturally, the Pentium is vastly different in many ways from the 486, but most of those differences are transparent to programmers. After all, the whole idea behind the Pentium is that it runs the same code as previous x86 processors, but faster; otherwise, Intel could have made a faster, cheaper RISC processor. Still, knowledge of the Pentium’s architecture is useful for understanding exactly how code will perform, and a few of the architectural differences are most decidedly not transparent to performance programmers.

    @@ -85,10 +78,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/19-02.html b/19-02.html index 064962e..a6dbcd0 100644 --- a/19-02.html +++ b/19-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - @@ -47,7 +40,7 @@

    (And yes, self-modifying code still works; as with all Pentium changes, the dual caches introduce no incompatibilities with 386/486 code.) Also, because the code and data caches are separate, code can’t be driven out of the cache in a tight loop that accesses a lot of data, unlike the 486. In addition, the Pentium expands the 486’s 32-byte prefetch queue to 128 bytes. In conjunction with the branch prediction feature (described next), which allows the Pentium to prefetch properly at most branches, this larger prefetch queue means that the Pentium’s two pipes should be better fed than those of any previous x86 processor.

    -

    Crossing Cache Lines

    +

    Crossing Cache Lines

    There are three other characteristics of the Pentium that make for a healthy supply of instruction bytes. One is that the Pentium can prefetch instructions across cache lines. Unlike the 486, where there is a 3-cycle penalty for branching to an instruction that spans a cache line, there’s no such penalty on the Pentium. The second is that the cache line size (the number of bytes fetched from the external cache or main memory on a cache miss) on the Pentium is 32 bytes, twice the size of the 486’s cache line, so a cache miss causes a longer run of instructions to be placed in the cache than on the 486. The third is that the Pentium’s external bus is twice as wide as the 486’s, at 64 bits, and runs twice as fast, at 66 MHz, so the Pentium can fetch both instruction and data bytes from the external cache four times as fast as the 486.

    @@ -63,11 +56,11 @@

    One change in the Pentium that you definitely do have to worry about is superscalar execution. Utilization of the V-pipe can range from near zero percent to 100 percent, depending on the code being executed, and careful rearrangement of code can have amazing effects. Maxing out V-pipe use is not a trivial task; I’ll spend all of the next chapter discussing it so as to have time to cover it properly. In the meantime, two good references for superscalar programming and other Pentium information are Intel’s Pentium Processor User’s Manual: Volume 3: Architecture and Programming Manual (ISBN 1-55512-195-0; Intel order number 241430-001), and the article “Optimizing Pentium Code” by Mike Schmidt, in Dr. Dobb’s Journal for January 1994.

    -

    Cache Organization

    +

    Cache Organization

    There are two other interesting changes in the Pentium’s cache organization. First, the cache is two-way set-associative, whereas the 486 is four-way set-associative. The details of this don’t matter, but simply put, this, combined with the 32-byte cache line size, means that the Pentium has somewhat coarser granularity in both space and time than the 486 in terms of packing bytes into the cache, although the total cache space is now bigger. There’s nothing you can do about this, but it may make it a little harder to get a loop’s working set into the cache. Second, the internal cache can now be configured (by the BIOS or OS; you won’t have to worry about it) for write-back rather than write-through operation. This means that writes to the internal data cache don’t necessarily get propagated to the external bus until other demands for cache space force the data out of the cache, making repeated writes to memory variables such as loop counters cheaper on average than on the 486, although not as cheap as registers.

    -

    As a final note on Pentium architecture for this chapter, the pipeline stalls (what Intel calls AGIs, for Address Generation Interlocks) that I discussed earlier in this book (see Chapter 12) are still present in the Pentium. In fact, they’re there in spades on the Pentium; the two pipelines mean that an AGI can now slow down execution of an instruction that’s three instructions away from the AGI (because four instructions can execute in two cycles). So, for example, the code sequence

    +

    As a final note on Pentium architecture for this chapter, the pipeline stalls (what Intel calls AGIs, for Address Generation Interlocks) that I discussed earlier in this book (see Chapter 12) are still present in the Pentium. In fact, they’re there in spades on the Pentium; the two pipelines mean that an AGI can now slow down execution of an instruction that’s three instructions away from the AGI (because four instructions can execute in two cycles). So, for example, the code sequence

     add edx,4     ;U-pipe cycle 1
     mov ecx,[ebx] ;V-pipe cycle 1
    @@ -77,15 +70,15 @@ mov [edx],ecx ;V-pipe cycle 3
                   ; (would have been
                   ; V-pipe cycle 2)
     
    -
    + -

    takes three cycles rather than the two cycles it should take, because EDX was modified on cycle 1 and an attempt was made to use it on cycle two, before the AGI had time to clear—even though there are two instructions between the instructions that are actually involved in the AGI. Rearranging the code like

    +

    takes three cycles rather than the two cycles it should take, because EDX was modified on cycle 1 and an attempt was made to use it on cycle two, before the AGI had time to clear—even though there are two instructions between the instructions that are actually involved in the AGI. Rearranging the code like

     mov ecx,[ebx]   ;U-pipe cycle 1
     add ebx,4       ;V-pipe cycle 1
     mov [edx+4],ecx ;U-pipe cycle 2
     add edx,4       ;V-pipe cycle 2
    -
    +

    makes it functionally identical, but cuts the cycles to 2—a 50 percent improvement. Clearly, avoiding AGIs becomes a much more challenging and rewarding game in a superscalar world, one to which I’ll return in the next chapter.

    @@ -106,10 +99,6 @@ add edx,4 ;V-pipe cycle 2
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/19-03.html b/19-03.html index 52cb95d..315e096 100644 --- a/19-03.html +++ b/19-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - @@ -37,16 +30,16 @@


    -

    Faster Addressing and More

    +

    Faster Addressing and More

    I’ll spend the rest of this chapter covering a variety of Pentium optimization tips. For starters, effective address calculations (that is, the addition and scaling required to calculate a memory operand’s address, as for example in MOV EAX,[EBX+ECX*2+4]) never take any extra cycles on the Pentium (other than possibly an AGI cycle), even for the use of base+index addressing (as in MOV [ESI+EDI],EAX) or scaling (*2, *4, or *8, as in INC ARRAY[ESI*4]). On the 486, both of the latter cases cause a 1-cycle penalty. The faster effective address calculations have the side effect of making LEA very attractive as an arithmetic instruction. LEA can add any two registers, one of which can be multiplied by one, two, four, or eight, plus a constant value, and can store the result in any register—all in one cycle, apart from AGIs. Not only that, but as we’ll see in the next chapter, LEA can go through either pipe, whereas SHL can only go through the U-pipe, so LEA is often a superior choice for multiplication by three, four, five, eight, or nine. (ADD is the best choice for multiplication by two.) If you use LEA for arithmetic, do remember that unlike ADD and SHL, it doesn’t modify any flags.

    -

    As on the 486, memory operands should not cross any more alignment boundaries than absolutely necessary. Word operands should be word-aligned, dword operands should be dword-aligned, and qword operands (double-precision variables) should be qword-aligned. Spanning a dword boundary, as in

    +

    As on the 486, memory operands should not cross any more alignment boundaries than absolutely necessary. Word operands should be word-aligned, dword operands should be dword-aligned, and qword operands (double-precision variables) should be qword-aligned. Spanning a dword boundary, as in

     mov ebx,3
      :
     mov eax,[ebx]
    -
    +

    costs three cycles. On the other hand, as noted above, branch targets can now span cache lines with impunity, so on the Pentium there’s no good argument for the paragraph (that is, 16-byte) alignment that Intel recommends for 486 jump targets. The 32-byte alignment might make for slightly more efficient Pentium cache usage, but would make code much bigger overall.

    @@ -89,10 +82,6 @@ mov eax,[ebx]
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/19-04.html b/19-04.html index 2e98227..a47fce3 100644 --- a/19-04.html +++ b/19-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium: Not the Same Old Song - - @@ -37,7 +30,7 @@


    -

    Branch Prediction

    +

    Branch Prediction

    One brand-spanking-new feature of the Pentium is branch prediction, whereby the Pentium tries to guess, based on past history, which way (or, for conditional jumps, whether or not), your code will jump at each branch, and prefetches along the likelier path. If the guess is correct, the branch or fall-through takes only 1 cycle—2 cycles less than a branch and the same as a fall-through on the 486; if the guess is wrong, the branch or fall-through takes 4 or 5 cycles (if it executes in the U- or V-pipe, respectively)—1 or 2 cycles more than a branch and 3 or 4 cycles more than a fall-through on the 486.

    @@ -63,17 +56,17 @@ -

    Miscellaneous Pentium Topics

    +

    Miscellaneous Pentium Topics

    The Pentium has all the instructions of the 486, plus a few new ones. One much-needed instruction that has finally made it into the instruction set is CPUID, which allows your code to determine what processor it’s running on. CPUID is 15 years late, but at least it’s finally here. Another new instruction is CMPXCHG8B, which does a compare and conditional exchange on a qword. CMPXCHG8B doesn’t seem to me to be a particularly useful instruction, but I’m sure Intel wouldn’t have added it without a reason; if you know of a use for it, please pass it along to me.

    -

    486 versus Pentium Optimization

    +

    486 versus Pentium Optimization

    Many Pentium optimizations help, or at least don’t hurt, on the 486. Many, but not all—and many do hurt on the 386. As I discuss various Pentium optimizations, I will attempt to note the effects on the 486 as well, but doing this in complete detail would double the sizes of these discussions and make them hard to follow. In general, I’d recommend reserving Pentium optimization for your most critical code, and even there, it’s a good idea to have at least two code paths, one for the 386 and one for the 486/Pentium. It’s also a good idea to time your code on a 486 before and after Pentium-optimizing it, to make sure you haven’t hurt performance on what will be, after all, by far the most important processor over the next couple of years.

    With that in mind, is optimizing for the Pentium even worthwhile today? That depends on your application and its market—but if you want absolutely the best possible performance for your DOS and Windows apps on the fastest hardware, Pentium optimization can make your code scream.

    -

    Going Superscalar

    +

    Going Superscalar

    In the next chapter, we’ll look into the single biggest element of Pentium performance, cranking up the Pentium’s second execution pipe. This is the area in which compiler technology is most touted for the Pentium, the two thoughts apparently being that (1) most existing code is in C, so recompiling to use the second pipe better is an automatic win, and (2) it’s so complicated to optimize Pentium code that only a compiler can do it well. The first point is a reasonable one, but it does suffer from one flaw for large programs, in that Pentium-optimized code is larger than 486- or 386-optimized code, for reasons that will become apparent in the next chapter. Larger code means more cache misses and more page faults; and while most of the code in any program is not critical to performance, compilers optimize code indiscriminately.

    @@ -98,10 +91,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/20-01.html b/20-01.html index 51a0f20..62044f1 100644 --- a/20-01.html +++ b/20-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - @@ -37,10 +30,10 @@


    -

    Chapter 20
    +

    Chapter 20
    Pentium Rules

    -

    How Your Carbon-Based Optimizer Can Put the “Super” in Superscalar

    +

    How Your Carbon-Based Optimizer Can Put the “Super” in Superscalar

    At the 1983 West Coast Computer Faire, my friend Dan Illowsky, Andy Greenberg (co-author of Wizardry, at that time the best-selling computer game ever), and I had an animated discussion about starting a company in the then-budding world of microcomputer software. One hot new software category at the time was educational software, and one of the hottest new educational software companies was Spinnaker Software. Andy used Spinnaker as an example of a company that had been aimed at a good market and started up properly, and was succeeding as a result. Dan didn’t buy this; his point was that Spinnaker had been given a bundle of money to get off the ground, and was growing only by spending a lot of that money in order to move its products. “Heck,” said Dan, “I could get that kind of market share too if I gave away a fifty-dollar bill with each of my games.”

    @@ -52,7 +45,7 @@

    Similarly, there’s no such thing as inherently fast code, only fast code in context. At the moment, the context is the Pentium, and the truth is that a sizable number of the x86 optimization tricks that you and I have learned over the past ten years are obsolete on the Pentium. True, the Pentium contains what amounts to about one-and-a-half 486s, but, as we’ll see shortly, that doesn’t mean that optimized Pentium code looks much like optimized 486 code, or that fast 486 code runs particularly well on a Pentium. (Fast Pentium code, on the other hand, does tend to run well on the 486; the only major downsides are that it’s larger, and that the FXCH instruction, which is largely free on the Pentium, is expensive on the 486.) So discard your x86 preconceptions as we delve into superscalar optimization for this one-of-a-kind processor.

    -

    An Instruction in Every Pipe

    +

    An Instruction in Every Pipe

    In the last chapter, we took a quick tour of the Pentium’s architecture, and started to look into the Pentium’s optimization rules. Now we’re ready to get to the key rules, those having to do with the Pentium’s most unique and powerful feature, the ability to execute more than one instruction per cycle. This is known as superscalar execution, and has heretofore been the sole province of fast RISC CPUs. The Pentium has two integer execution units, called the U-pipe and the V-pipe, which can execute two separate instructions simultaneously, potentially doubling performance—but only under the proper conditions. (There is also a separate floating-point execution unit that I won’t have the space to cover in this book.) Your job, as a performance programmer, is to understand the conditions needed for superscalar performance and make sure they’re met, and that’s what this and the next chapters are all about.

    @@ -60,9 +53,8 @@

    The U-pipe is the more capable of the two pipes, able to execute any instruction in the Pentium’s instruction set. (A number of instructions actually use both pipes at once. Logically, though, you can think of such instructions as U-pipe instructions, and of the Pentium optimization model as one in which the U-pipe is able to execute all instructions and is always active, with the objective being to keep the V-pipe also working as much of the time as possible.) The U-pipe is generally similar to a full 486 in terms of both capabilities and instruction cycle counts. The V-pipe is a 486 subset, able to execute simple instructions such as MOV and ADD, but unable to handle MUL, DIV, string instructions, any sort of rotation or shift, or even ADC or SBB.

    -


    - Figure 20.1
      The Pentium’s two pipes.

    +


    + Figure 20.1
      The Pentium’s two pipes.

    Getting two instructions executing simultaneously in the two pipes is trickier than it sounds, not only because the V-pipe can handle only a relatively small subset of the Pentium’s instruction set, but also because those instructions that the V-pipe can handle are able to pair only with certain U-pipe instructions. For example, MOVSD uses both pipes, so no instruction can be executed in parallel with MOVSD.

    @@ -95,10 +87,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/20-02.html b/20-02.html index cc2f45a..e77e7a3 100644 --- a/20-02.html +++ b/20-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - @@ -37,13 +30,12 @@


    -

    V-Pipe-Capable Instructions

    +

    V-Pipe-Capable Instructions

    Any instruction can go through the U-pipe, and, for practical purposes, the U-pipe is always executing instructions. (The exceptions are when the U-pipe execution unit is waiting for instruction or data bytes after a cache miss, and when a U-pipe instruction finishes before a paired V-pipe instruction, as I’ll discuss below.) Only the instructions shown in Table 20.1 can go through the V-pipe. In addition, the V-pipe can execute a separate instruction only when one of the instructions listed in Table 20.2 is executing in the U-pipe; superscalar execution is not possible while any instruction not listed in Table 20.2 is executing in the U-pipe. So, for example, if you use SHR EDX,CL, which takes 4 cycles to execute, no other instructions can execute during those 4 cycles; if, on the other hand, you use SHR EDX,10, it will take 1 cycle to execute in the U-pipe, and another instruction can potentially execute concurrently in the V-pipe. (As you can see, similar instruction sequences can have vastly different performance characteristics on the Pentium.)

    Basically, after the current instruction or pair of instructions is finished (that is, once neither the U- nor V-pipe is executing anything), the Pentium sends the next instruction through the U-pipe. If the instruction after the one in the U-pipe is an instruction the V-pipe can handle, if the instruction in the U-pipe is pairable, and if register contention doesn’t occur, then the V-pipe starts executing that instruction, as shown in Figure 20.2. Otherwise, the second instruction waits until the first instruction is done, then executes in the U-pipe, possibly pairing with the next instruction in line if all pairing conditions are met.


    -
     MOV      reg,reg          (1 cycle)
              mem,reg          (1 cycle)
    @@ -82,7 +74,7 @@ JMP/CALL near      (1 cycle if predicted correctly;
                         3 cycles otherwise)
     
      Can’t execute in V-pipe if address contains a displacement
    -
    +

    Table 20.1 Instructions that can execute in the V-pipe.


    @@ -91,7 +83,6 @@ JMP/CALL near (1 cycle if predicted correctly;

    Besides, almost all operations can be performed by combinations of pairable instructions. For example, PUSH [mem] is not on either list, but both MOV reg,[mem] and PUSH reg are, and those two instructions can be used to push a value stored in memory. In fact, given the proper instruction stream, the discrete instructions can perform this operation effectively in just 1 cycle (taking one-half of each of 2 cycles, for 2*0.5 = 1 cycle total execution time), as shown in Figure 20.3—a full cycle faster than PUSH [mem], which takes 2 cycles.


    -
     MOV      reg,reg           (1 cycle)
              mem,reg           (1 cycle)
    @@ -128,7 +119,7 @@ ROL/ROR/RCL/RCR  reg,1           (1 cycle)
     
        Can’t pair if address contains a displacement
     †† Includes shift-by-1 forms of instructions
    -
    +

    Table 20.2 Instructions that, when executed in the U-pipe, allow V-pipe-executable instructions to execute simultaneously (pair) in the V-pipe.


    @@ -141,36 +132,34 @@ ROL/ROR/RCL/RCR reg,1 (1 cycle) -


    - Figure 20.2
      Instruction flow through the two pipes.

    +


    + Figure 20.2
      Instruction flow through the two pipes.

    -

    One downside of this “RISCification” (turning complex instructions into simple, RISC-like ones) of Pentium-optimized code is that it makes for substantially larger code. For example,

    +

    One downside of this “RISCification” (turning complex instructions into simple, RISC-like ones) of Pentium-optimized code is that it makes for substantially larger code. For example,

     push dword ptr [esi]
    -
    + -

    is one byte smaller than this sequence:

    +

    is one byte smaller than this sequence:

     mov eax,[esi]
     push eax
    -
    + -


    - Figure 20.3
      Pushing a value from memory effectively in one cycle.

    +


    + Figure 20.3
      Pushing a value from memory effectively in one cycle.

    -

    A more telling example is the following

    +

    A more telling example is the following

     add [MemVar],eax
    -
    + -

    versus the equivalent:

    +

    versus the equivalent:

     mov  edx,[MemVar]
     add  edx,eax
     mov  [MemVar],edx
    -
    +

    The single complex instruction takes 3 cycles and is 6 bytes long; with proper sequencing, interleaving the simple instructions with other instructions that don’t use EDX or Mem Var, the three-instruction sequence can be reduced to 1.5 cycles, but it is 14 bytes long.

    @@ -199,10 +188,6 @@ mov [MemVar],edx
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/20-03.html b/20-03.html index dd02c9c..7ff97de 100644 --- a/20-03.html +++ b/20-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - @@ -37,7 +30,7 @@


    -

    Lockstep Execution

    +

    Lockstep Execution

    You may wonder why anyone would bother breaking ADD [MemVar],EAX into three instructions, given that this instruction can go through either pipe with equal ease. The answer is that while the memory-accessing instructions other than MOV, PUSH, and POP listed in Table 20.1 (that is, INC/DEC [mem], ADD/SUB/XOR/AND/OR/CMP/ADC/SBB reg,[mem], and ADD/SUB/XOR/AND/OR/CMP/ADC/SBB [mem],reg/immed) can be paired, they do not provide the 100 percent overlap that we seek. If you look at Tables 20.1 and 20.2, you will see that instructions taking from 1 to 3 cycles can pair. However, any pair of instructions goes through the two pipes in lockstep. This means, for example, that if ADD [EBX],EDX is going through the U-pipe, and INC EAX is going through the V-pipe, the V-pipe will be idle for 2 of the 3 cycles that the U-pipe takes to execute its instruction, as shown in Figure 20.4. Out of the theoretical 6 cycles of work that can be done during this time, we actually get only 4 cycles of work, or 67 percent utilization. Even though these instructions pair, then, this sequence fails to make maximum use of the Pentium’s horsepower.

    @@ -51,9 +44,8 @@ -


    - Figure 20.4
      Lockstep execution and idle time in the V-pipe.

    +


    + Figure 20.4
      Lockstep execution and idle time in the V-pipe.

    Here’s why. The Pentium is fully capable of handling instructions that use memory operands in either pipe, or, if necessary, in both pipes at once. Each pipe has its own write FIFO, which buffers the last few writes and takes care of writing the data out while the Pentium continues processing. The Pentium also has a write-back internal data cache, so data that is frequently changed doesn’t have to be written to external memory (which is much slower than the cache) very often. This combination means that unless you write large blocks of data at a high speed, the Pentium should be able to keep up with both pipes’ memory writes without stalling execution.

    @@ -61,35 +53,34 @@

    Normally, you won’t pay close attention to which of the eight dword banks your paired memory accesses fall in—that’s just too much work—but you might want to watch out for simultaneously read addresses that have the same values for address

    -


    - Figure 20.5
      The Pentium’s eight bank data cache.

    +


    + Figure 20.5
      The Pentium’s eight bank data cache.

    -

    bits 2, 3, and 4 (fall in the same bank) in tight loops, and you should also avoid sequences like

    +

    bits 2, 3, and 4 (fall in the same bank) in tight loops, and you should also avoid sequences like

     mov  bl,[esi]
     mov  bh,[esi+1]
    -
    + -

    because both operands will generally be in the same bank. An alternative is to place another instruction between the two instructions that access the same bank, as in this sequence:

    +

    because both operands will generally be in the same bank. An alternative is to place another instruction between the two instructions that access the same bank, as in this sequence:

     mov  bl,[esi]
     mov  edi,edx
     mov  bh,[esi+1]
    -
    + -

    By the way, the reason a code sequence that takes two instructions to load a single word is attractive in a 32-bit segment is because it takes only one cycle when the two instructions can be paired with other instructions; by contrast, the obvious way of loading BX

    +

    By the way, the reason a code sequence that takes two instructions to load a single word is attractive in a 32-bit segment is because it takes only one cycle when the two instructions can be paired with other instructions; by contrast, the obvious way of loading BX

     mov bx,[esi]
    -
    +

    takes 1.5 to two cycles because the size prefix can’t pair, as described below. This is yet another example of how different Pentium optimization can be from everything we’ve learned about its predecessors.

    -

    The problem with pairing non-single-cycle instructions arises when a pipe executes an instruction other than MOV that has an explicit memory operand. (I’ll call these complex memory instructions. They’re the only pairable instructions, other than branches, that take more than one cycle.) We’ve already seen that, because instructions go through the pipes in lockstep, if one pipe executes a complex memory instruction such as ADD EAX,[EBX] while the other pipe executes a single-cycle instruction, the pipe with the faster instruction will sit idle for part of the time, wasting cycles. You might think that if both pipes execute complex instructions of the same length, then neither would lie idle, but that turns out to not always be the case. Two two-cycle instructions (instructions with register destination operands) can indeed pair and execute in two cycles, so it’s okay to pair two instructions such as these:

    +

    The problem with pairing non-single-cycle instructions arises when a pipe executes an instruction other than MOV that has an explicit memory operand. (I’ll call these complex memory instructions. They’re the only pairable instructions, other than branches, that take more than one cycle.) We’ve already seen that, because instructions go through the pipes in lockstep, if one pipe executes a complex memory instruction such as ADD EAX,[EBX] while the other pipe executes a single-cycle instruction, the pipe with the faster instruction will sit idle for part of the time, wasting cycles. You might think that if both pipes execute complex instructions of the same length, then neither would lie idle, but that turns out to not always be the case. Two two-cycle instructions (instructions with register destination operands) can indeed pair and execute in two cycles, so it’s okay to pair two instructions such as these:

     add esi,[SourceSkip]        ;U-pipe cycles 1 and 2
     add edi,[DestinationSkip]   ;V-pipe cycles 1 and 2
    -
    +


    @@ -108,10 +99,6 @@ add edi,[DestinationSkip] ;V-pipe cycles 1 and 2
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/20-04.html b/20-04.html index 8f5dc99..ddeb1d8 100644 --- a/20-04.html +++ b/20-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pentium Rules - - @@ -39,35 +32,33 @@

    However, this beneficial pairing does not extend to non-MOV instructions with explicit memory destination operands, such as ADD [EBX],EAX. The Pentium executes only one such memory instruction at a time; if two memory-destination complex instructions get paired, first the U-pipe instruction is executed, and then the V-pipe instruction, with only one cycle of overlap, as shown in Figure 20.6. I don’t know for sure, but I’d guess that this is to guarantee that the two pipes will never perform out-of-order access to any given memory location. Thus, even though AND [EBX],AL pairs with AND [ECX],DL, the two instructions take 5 cycles in all to execute, and 4 cycles of idle time—2 in the U-pipe and 2 in the V-pipe, out of 10 cycles in all—are incurred in the process.

    -


    - Figure 20.6
      Non-overlapped lockstep execution.

    +


    + Figure 20.6
      Non-overlapped lockstep execution.

    -


    - Figure 20.7
      Interleaving simple instructions for maximum performance.

    +


    + Figure 20.7
      Interleaving simple instructions for maximum performance.

    The solution is to break the instructions into simple instructions and interleave them, as shown in Figure 20.7, which accomplishes the same task in 3 cycles, with no idle cycles whatsoever. Figure 20.7 is a good example of what optimized Pentium code generally looks like: mostly one-cycle instructions, mixed together so that at least two operations are in progress at once. It’s not the easiest code to read or write, but it’s the only way to get both pipes running at capacity.

    -

    Superscalar Notes

    +

    Superscalar Notes

    -

    You may well ask why it’s necessary to interleave operations, as is done in Figure 20.7. It seems simpler just to turn

    +

    You may well ask why it’s necessary to interleave operations, as is done in Figure 20.7. It seems simpler just to turn

     and [ebx],al
    -
    + -

    into

    +

    into

     mov   dl,[ebx]
     and   dl,al
     mov   [ebx],dl
    -
    +

    and be done with it. The problem here is one of dependency. Before the Pentium can execute AND DL,AL,, it must first know what is in DL, and it can’t know that until it loads DL from the address pointed to by EBX. Therefore, AND DL,AL can’t happen until the cycle after MOV DL,[EBX] executes. Likewise, the result can’t be stored until the cycle after AND DL,AL has finished. This means that these instructions, as written, can’t possibly pair, so the sequence takes the same three cycles as AND [EBX],AL. (Now it should be clear why AND [EBX], AL takes 3 cycles.) Consequently, it’s necessary to interleave these instructions with instructions that use other registers, so this set of operations can execute in one pipe while the other, unrelated set executes in the other pipe, as is done in Figure 20.7.

    What we’ve just seen is the read-after-write form of the superscalar hazard known as register contention. I’ll return to the subject of register contention in the next chapter; in the remainder of this chapter I’d like to cover a few short items about superscalar execution.

    -

    Register Starvation

    +

    Register Starvation

    The above examples should make it pretty clear that effective superscalar programming puts a lot of strain on the Pentium’s relatively small register set. There are only seven general-purpose registers (I strongly suggest using EBP in critical loops), and it does not help to have to sacrifice one of those registers for temporary storage on each complex memory operation; in pre-superscalar days, we used to employ those handy CISC memory instructions to do all that stuff without using any extra registers.

    @@ -85,9 +76,8 @@ mov [ebx],dl

    It should be excruciatingly clear by this point that you must time your Pentium-optimized code if you’re to have any hope of knowing if your optimizations are working as well as you think they are; there are just too many details involved for you to be sure your optimizations are working properly without checking. My most basic optimization rule has always been to grab the Zen timer and measure actual performance—and nowhere is this more true than on the Pentium. Don’t believe it until you measure it!

    -


    - Figure 20.8
      Prefix delays.

    +


    + Figure 20.8
      Prefix delays.


    @@ -106,10 +96,6 @@ mov [ebx],dl
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/21-01.html b/21-01.html index f06a15e..25215c9 100644 --- a/21-01.html +++ b/21-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - @@ -37,10 +30,10 @@


    -

    Chapter 21
    +

    Chapter 21
    Unleashing the Pentium’s V-Pipe

    -

    Focusing on Keeping Both Pentium Pipes Full

    +

    Focusing on Keeping Both Pentium Pipes Full

    The other day, my daughter suggested that we each draw the prettiest picture we could, then see whose was prettier. I won’t comment on who won, except to note that apparently a bolt of lightning zipping toward a moose with antlers that bear an unfortunate resemblance to a propeller beanie isn’t going to win me any scholarships to art school, if you catch my drift. Anyway, my drawing happened to feature the word “chartreuse” (because it rhymed with “moose” and “Zeus”—hence the lightning; more than that I am not at liberty to divulge), and she wanted to know if the moose was actually chartreuse. I had to admit that I didn’t know, so we went to the dictionary, whereupon we learned that chartreuse is a pale apple-green color. Then she brought up the Windows Control Panel, pointed to the selection of predefined colors, and asked, “Which of those is chartreuse?”—and I realized that I still didn’t know.

    @@ -48,24 +41,23 @@

    In the last chapter, we explored the dual-execution-pipe nature of the Pentium, and learned which instructions could pair (execute simultaneously) in which pipes. Now we’re ready to look at AGIs and register contention—two hazards that can prevent otherwise properly written code from taking full advantage of the Pentium’s two pipes, and can thereby keep your code from pushing the Pentium to maximum performance.

    -

    Address Generation Interlocks

    +

    Address Generation Interlocks

    The Pentium is advertised as having a five-stage pipeline for each of its execution units. All this means is that at any given time, up to five instructions are in various stages of execution in each pipe; this overlapping of execution is done for speed, so each instruction doesn’t have to wait until the previous one has finished. The only way that the Pentium’s pipelining directly affects the way you program is in the areas of AGIs and register dependencies.

    -

    AGIs are Address Generation Interlocks, a fancy way of saying that if a register is used to address memory, as is EBX in this instruction

    +

    AGIs are Address Generation Interlocks, a fancy way of saying that if a register is used to address memory, as is EBX in this instruction

     mov [ebx],eax
    -
    +

    and the value of the register is not set far enough ahead for the Pentium to perform the addressing calculations before the instruction needs the address, then the Pentium will stall the pipe in which the instruction is executing until the value becomes available and the addressing calculations have been performed. Remember, also, that instructions execute in lockstep on the Pentium, so if one pipe stalls for a cycle, making its instruction take one cycle longer, that extends by one cycle the time until the other pipe can begin its next instruction, as well.

    The rule for AGIs is simple: If you modify any part of a register during a cycle, you cannot use that register to address memory during either that cycle or the next cycle. If you try to do this, the Pentium will simply stall the instruction that tries to use that register to address memory until two cycles after the register was modified. This was true on the 486 as well, but the Pentium’s new twist is that since more than one instruction can execute in a single cycle, an AGI can stall an instruction that’s as many as three instructions away from the changing of the addressing register, as shown in Figure 21.1, and an AGI can also cause a stall that costs as many as three instructions, as shown in Figure 21.2. This means that AGIs are both much easier to cause and potentially more expensive than on the 486, and you must keep a sharp eye out for them. It also means that it’s often worth calculating a memory pointer several instructions ahead of its actual use. Unfortunately, this tends to extend the lifetimes of pointer registers to span a greater number of instructions, making the Pentium’s relatively small register set seem even smaller.

    -


    - Figure 21.1
      An AGI can stall up to three instructions later.

    +


    + Figure 21.1
      An AGI can stall up to three instructions later.

    -

    As an example of a sort of AGI that’s new to the Pentium, consider the following test for a NULL pointer, followed by the use of the pointer if it’s not NULL:

    +

    As an example of a sort of AGI that’s new to the Pentium, consider the following test for a NULL pointer, followed by the use of the pointer if it’s not NULL:

     push ebx          ;U-pipe cycle 1
     mov  ebx,[Ptr]    ;V-pipe cycle 1
    @@ -75,13 +67,12 @@ mov  eax,[ebx]    ;U-pipe cycle 3 AGI stall
     mov  edx,[ebp-8]  ;V-pipe cycle 3 lockstep idle
                       ;U-pipe cycle 4 mov eax,[ebx]
                       ;V-pipe cycle 4 mov edx,[ebp-8]
    -
    +

    This commonplace code loses a U-pipe cycle to the AGI caused by AND EBX,EBX, followed by the attempt two instructions later to use EBX to point to memory. The code loses a V-pipe cycle as well, because lockstep execution won’t let the next V-pipe instruction execute until the paired U-pipe instruction that suffered the AGI finishes. The solution is to use TEST EBX,EBX instead of AND; TEST can’t modify EBX, so no AGI occurs. Sure, AND EBX,EBX doesn’t modify EBX either, but the Pentium doesn’t know that, so it has to insert the AGI.

    -


    - Figure 21.2
      An AGI can cost as many as 3 cycles.

    +


    + Figure 21.2
      An AGI can cost as many as 3 cycles.


    @@ -100,10 +91,6 @@ mov edx,[ebp-8] ;V-pipe cycle 3 lockstep idle
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/21-02.html b/21-02.html index 75ede03..b92c840 100644 --- a/21-02.html +++ b/21-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - @@ -37,29 +30,29 @@


    -

    As on the 486, you should keep a careful eye out for AGIs involving the stack pointer. Implicit modifiers of ESP, such as PUSH and POP, are special-cased so you don’t have to worry about AGIs. However, if you explicitly modify ESP with this instruction

    +

    As on the 486, you should keep a careful eye out for AGIs involving the stack pointer. Implicit modifiers of ESP, such as PUSH and POP, are special-cased so you don’t have to worry about AGIs. However, if you explicitly modify ESP with this instruction

     sub esp,100h
    -
    + -

    for example, or with the popular

    +

    for example, or with the popular

     mov esp,ebp
    -
    + -

    you can then get AGIs if you attempt to use ESP to address memory, either explicitly with instructions like this one

    +

    you can then get AGIs if you attempt to use ESP to address memory, either explicitly with instructions like this one

     moveax,[esp+20h]
    -
    +

    or via PUSH, POP, or other instructions that implicitly use ESP as an addressing register.

    -

    On the 486, any instruction that had both a constant value and an addressing displacement, such as

    +

    On the 486, any instruction that had both a constant value and an addressing displacement, such as

     mov dword ptr [ebp+16],1
    -
    + -

    suffered a 1-cycle penalty, taking a total of 2 cycles. Such instructions take only one cycle on the Pentium, but they cannot pair, so they’re still the most expensive sort of MOV. Knowing this can speed up something as simple as zeroing two memory variables, as in

    +

    suffered a 1-cycle penalty, taking a total of 2 cycles. Such instructions take only one cycle on the Pentium, but they cannot pair, so they’re still the most expensive sort of MOV. Knowing this can speed up something as simple as zeroing two memory variables, as in

     sub eax,eax        ;U-pipe 1
                        ;any V-pipe pairable
    @@ -67,49 +60,49 @@ sub eax,eax        ;U-pipe 1
                        ; or SUB could be in V-pipe
     mov [MemVar1],eax  ;U-pipe 2
     mov [MemVar2],eax  ;V-pipe 2
    -
    + -

    which should never be slower and should potentially be 0.5 cycles faster, and six bytes smaller than this sequence:

    +

    which should never be slower and should potentially be 0.5 cycles faster, and six bytes smaller than this sequence:

     mov [MemVar1],0 ;U-pipe 1
     mov [MemVar2],0 ;U-pipe 2
    -
    +

    Note, however, that my experiments thus far indicate that the two writes in the first case don’t actually pair (possibly because the memory variables have never been read into the internal cache), so you might want to insert an instruction between the two MOVs—and, of course, this is yet another reason why you should always measure your code’s actual performance.

    -

    Register Contention

    +

    Register Contention

    -

    Finally, we come to the last major component of superscalar optimization: register contention. The basic premise here is simple: You can’t use the same register in two inherently sequential ways in a single cycle. For example, you can’t execute

    +

    Finally, we come to the last major component of superscalar optimization: register contention. The basic premise here is simple: You can’t use the same register in two inherently sequential ways in a single cycle. For example, you can’t execute

     inc eax     ;U-pipe cycle 1
                 ;V-pipe idle cycle 1
                 ; due to dependency
     and ebx,eax ;U-pipe cycle 2
    -
    +

    in a single cycle; AND EBX,EAX can’t execute until the value in EAX is known, and that can’t happen until INC EAX is done. Consequently, the V-pipe idles while INC EAX executes in the U-pipe. We saw this in the last chapter when we discussed splitting instructions into simple instructions, and it is by far the most common sort of register contention, known as read-after-write register contention. Read-after-write register contention is the primary reason we have to interleave independent operations in order to get maximum V-pipe usage.

    -

    The other sort of register contention is known as write-after-write. Write-after-write register contention happens when two instructions try to write to the same register on the same cycle. While that may not seem like a particularly useful operation in general, it can happen when subregisters are being set, as in the following

    +

    The other sort of register contention is known as write-after-write. Write-after-write register contention happens when two instructions try to write to the same register on the same cycle. While that may not seem like a particularly useful operation in general, it can happen when subregisters are being set, as in the following

     sub eax,eax   ;U-pipe cycle 1
                   ;V-pipe idle cycle 1
                   ; due to register contention
     mov al,[Var]  ;U-pipe cycle 2
    -
    +

    where an attempt is made to set both EAX and its AL subregister on the same cycle. Write-after-write contention implies that the two instructions comprising the above substitute for MOVZX should have at least one unrelated instruction between them when SUB EAX,EAX executes in the V-pipe.

    -

    Exceptions to Register Contention

    +

    Exceptions to Register Contention

    -

    Intel has special-cased some very useful exceptions to register contention. Happily, write-after-read operations do not cause contention. Such operations, as in

    +

    Intel has special-cased some very useful exceptions to register contention. Happily, write-after-read operations do not cause contention. Such operations, as in

     mov eax,edx ;U-pipe cycle 1
     sub edx,edxX ;V-pipe cycle 1
    -
    +

    are free of charge.

    -

    Also, stack-related instructions that modify ESP only implicitly (without ESP as part of any explicit operand) do not cause AGIs, and neither do they cause register contention with other instructions that use ESP only implicitly; such instructions include PUSH reg/immed, POP reg, and CALL. (However, these instructions do cause register contention on ESP—but not AGIs—with instructions that use ESP explicitly, such as MOV EAX,[ESP+4].) Without this special case, the following sequence would hardly use the V-pipe at all:

    +

    Also, stack-related instructions that modify ESP only implicitly (without ESP as part of any explicit operand) do not cause AGIs, and neither do they cause register contention with other instructions that use ESP only implicitly; such instructions include PUSH reg/immed, POP reg, and CALL. (However, these instructions do cause register contention on ESP—but not AGIs—with instructions that use ESP explicitly, such as MOV EAX,[ESP+4].) Without this special case, the following sequence would hardly use the V-pipe at all:

     mov  eax,[MemVar] ;U-pipe cycle 1
     push esi          ;V-pipe cycle 1
    @@ -117,24 +110,24 @@ push eax          ;U-pipe cycle 2
     push edi          ;V-pipe cycle 2
     push ebx          ;U-pipe cycle 3
     call FooTilde     ;V-pipe cycle 3
    -
    +

    But in fact, all the instructions pair, even though ESP is modified five times in the space of six instructions.

    -

    The final register-contention special case is both remarkable and remarkably important. There is exactly one sort of instruction that can pair only in the V-pipe: branches. Any near call or conditional or unconditional near jump can execute in the V-pipe paired with any pairable U-pipe instruction, as illustrated by this sequence:

    +

    The final register-contention special case is both remarkable and remarkably important. There is exactly one sort of instruction that can pair only in the V-pipe: branches. Any near call or conditional or unconditional near jump can execute in the V-pipe paired with any pairable U-pipe instruction, as illustrated by this sequence:

     LoopTop:
        mov [esi],eax ;U-pipe cycle 1
        add esi,4     ;V-pipe cycle 1
        dec ecx       ;U-pipe cycle 2
        jnz LoopTop   ;V-pipe cycle 2
    -
    +

    Branches can’t pair in the U-pipe; a branch that executes in the U-pipe runs alone, with the V-pipe idle. If a call or jump is correctly predicted by the Pentium’s branch prediction circuitry (as discussed in the last chapter), it executes in a single cycle, pairing if it runs in the V-pipe; if mispredicted, conditional jumps take 4 cycles in the U-pipe and 5 cycles in the V-pipe, and mispredicted calls and unconditional jumps take 3 cycles in either pipe. Note that RET can’t pair.

    -

    Who’s in First?

    +

    Who’s in First?

    -

    One of the trickiest things about superscalar optimization is that a given instruction stream can execute at a different speed depending on the pipe where it starts execution, because which instruction goes through the U-pipe first determines which of the following instructions will be able to pair. If we take the last example and add one more instruction, the other instructions will go through different pipes than previously, and cause the loop as a whole to take 50 percent longer, even though we only added 25 percent more cycles:

    +

    One of the trickiest things about superscalar optimization is that a given instruction stream can execute at a different speed depending on the pipe where it starts execution, because which instruction goes through the U-pipe first determines which of the following instructions will be able to pair. If we take the last example and add one more instruction, the other instructions will go through different pipes than previously, and cause the loop as a whole to take 50 percent longer, even though we only added 25 percent more cycles:

     LoopTop:
        inc edx           ;U-pipe cycle 1
    @@ -145,7 +138,7 @@ LoopTop:
                          ;V-pipe idle cycle 3
                          ; because JNZ can’t
                          ; pair in the U-pipe
    -
    +


    @@ -164,10 +157,6 @@ LoopTop:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/21-03.html b/21-03.html index 1103a32..473c8a3 100644 --- a/21-03.html +++ b/21-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - @@ -39,11 +32,11 @@

    It’s actually not hard to figure out which instructions go through which pipes; just back up until you find an instruction that can’t pair or can only go through the U-pipe, and work forward from there, given the knowledge that that instruction executes in the U-pipe. The easiest thing to look for is branches. All branch target instructions execute in the U-pipe, as do all instructions after conditional branches that fall through. Instructions with prefix bytes are generally good U-pipe markers, although they’re expensive instructions that should be avoided whenever possible, and have at least one aberration with regard to pipe usage, as discussed below. Shifts, rotates, ADC, SBB, and all other instructions not listed in Table 20.1 in the last chapter are likewise U-pipe markers.

    -

    Pentium Optimization in Action

    +

    Pentium Optimization in Action

    Now, let’s take a look at one of the simplest, tightest pieces of code imaginable, and see what our new Pentium perspective reveals. Listing 21.1 shows a loop implementing the TCP/IP checksum, a 16-bit checksum that wraps carries around to the low bit so that the result is endian-independent. This makes it easy to perform checksums on blocks of data regardless of the endian characteristics of the machines on which those blocks are generated and received. (Thanks to fellow performance enthusiast Terje Mathisen for suggesting this checksum as fertile ground for Pentium optimization, in the ibm.pc/fast.code forum on Bix.) The loop in Listing 21.1 consists of exactly five instructions; it’s hard to imagine that there’s a lot of performance to be wrung from this snippet, right?

    -

    LISTING 21.1 L21-1.ASM

    +

    LISTING 21.1 L21-1.ASM

     ; Calculates TCP/IP (16-bit carry-wrapping) checksum for buffer
     ;  starting at ESI, of length ECX words.
    @@ -73,7 +66,7 @@ ckloop:
             add     esi,2           ;cycle 5 V-pipe
             dec     ecx             ;cycle 6 U-pipe
             jnz     ckloop          ;cycle 6 V-pipe
    -
    +

    Wrong, wrong, wrong! As detailed in Listing 21.1, this loop should take 6 cycles per checksummed word in 32-bit protected mode, a ridiculously high number for the Pentium. (You’ll see why I say “should take,” not “takes,” shortly.) We should lose 2 cycles in each pipe to the two size prefixes (because the ADDs are 16-bit operations in a 32-bit segment), and another 2 cycles because of register contention that arises when ADC AX,0 has to wait for the result of ADD AX,[ESI]. Then, too, even though DEC and JNZ can pair and the branch prediction for JNZ is presumably correct virtually all the time, they do take a full cycle, and maybe we can do something about that as well.

    @@ -96,10 +89,6 @@ ckloop:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/21-04.html b/21-04.html index b95186a..27e36c3 100644 --- a/21-04.html +++ b/21-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - @@ -41,7 +34,7 @@

    Listing 21.2 shows one interesting alternative that doesn’t really buy us anything. Here, we’ve eliminated all size prefixes by doing byte-sized MOVs and ADDs, but because the size prefix on ADD AX,[ESI], for whatever reason, didn’t cost anything in Listing 21.1, our efforts are to no avail—Listing 21.2 still takes 4 cycles per checksummed word. What’s worth noting about Listing 21.2 is the extent to which the code is broken into simple instructions and reordered so as to avoid size prefixes, register contention, AGIs, and data bank conflicts (the latter because both [ESI] and [ESI+1] are in the same cache data bank, as discussed in the last chapter).

    -

    LISTING 21.2 L21-2.ASM

    +

    LISTING 21.2 L21-2.ASM

     ; Calculates TCP/IP (16-bit carry-wrapping) checksum for buffer
     ;  starting at ESI, of length ECX words.
    @@ -69,11 +62,11 @@ ckloop:
     ckloopend:
             add     ax,dx           ;checksum the last word
             adc     eax,0
    -
    +

    Listing 21.3 is a more sophisticated attempt to speed up the checksum calculation. Here we see a hallmark of Pentium optimization: two operations (the checksumming of the current and next pair of words) interleaved together to allow both pipes to run at near maximum capacity. Another hallmark that’s apparent in Listing 21.3 is that Pentium-optimized code tends to use more registers and require more instructions than 486-optimized code. Again, note the careful mixing of byte-sized reads to avoid AGIs, register contention, and cache bank collisions, in particular the way in which the byte reads of memory are interspersed with the additions to avoid register contention, and the placement of ADD ESI,4 to avoid an AGI.

    -

    LISTING 21.3 L21-3.ASM

    +

    LISTING 21.3 L21-3.ASM

     ; Calculates TCP/IP (16-bit carry-wrapping) checksum for buffer
     ;  starting at ESI, of length ECX words.
    @@ -121,7 +114,7 @@ ckloopend:
             add         ax,dx
             adc         eax,0
     ckloopdone:
    -
    +

    The checksum loop in Listing 21.3 takes longer than the loop in Listing 21.2, at 6 cycles versus 4 cycles for Listing 21.2—but Listing 21.3 does two checksum operations in those 6 cycles, so we’ve cut the time per checksum addition from 4 to 3 cycles. You might think that this small an improvement doesn’t justify the additional complexity of Listing 21.3, but it is a one-third speedup, well worth it if this is a critical loop—and, in general, if it isn’t critical, there’s no point in hand-tuning it. That’s why I haven’t bothered to try to optimize the non-inner-loop code in Listing 21.3; it’s only executed once per checksum, so it’s unlikely that a cycle or two saved there would make any real-world difference.

    @@ -144,10 +137,6 @@ ckloopdone:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/21-05.html b/21-05.html index 8eef28e..a1c6b89 100644 --- a/21-05.html +++ b/21-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Unleashing the Pentium's - - @@ -37,7 +30,7 @@


    -

    LISTING 21.4 L21-4.ASM

    +

    LISTING 21.4 L21-4.ASM

     ; Calculates TCP/IP (16-bit carry-wrapping) checksum for buffer
     ;  starting at ESI, of length ECX words.
    @@ -69,11 +62,11 @@ ckloopend:
             shr     edx,16          ; into a 16-bit checksum
             add     ax,dx
             adc     eax,0
    -
    +

    Listing 21.5 improves upon Listing 21.4 by processing 2 dwords per loop, thereby bringing the time per checksummed word down to exactly 1 cycle. Listing 21.5 basically does nothing but unroll Listing 21.4’s loop one time, demonstrating that the venerable optimization technique of loop unrolling still has some life left in it on the Pentium. The cost for this is, as usual, increased code size and complexity, and the use of more registers.

    -

    LISTING 21.5 L21-5.ASM

    +

    LISTING 21.5 L21-5.ASM

     ; Calculates TCP/IP (16-bit carry-wrapping) checksum for buffer
     ;  starting at ESI, of length ECX words.
    @@ -115,13 +108,13 @@ ckloopdone:
             shr     edx,16          ; into a 16-bit checksum
             add     ax,dx
             adc     eax,0
    -
    +

    Listing 21.5 is undeniably intricate code, and not the sort of thing one would choose to write as a matter of course. On the other hand, it’s five times as fast as the tight, seemingly-speedy loop in Listing 21.1 (and six times as fast as Listing 21.1 would have been if the prefix byte had behaved as expected). That’s an awful lot of speed to wring out of a five-instruction loop, and the TCP/IP checksum is, in fact, used by network software, an area in which a five-times speedup might make a significant difference in overall system performance.

    I don’t claim that Listing 21.5 is the fastest possible way to do a TCP/IP checksum on a Pentium; in fact, it isn’t. Unrolling the loop one more time, together with a trick of Terje’s that uses LEA to advance ESI (neither LEA nor DEC affects the carry flag, allowing Terje to add the carry from the previous loop iteration into the next iteration’s checksum via ADC), produces a version that’s a full 33 percent faster. Nonetheless, Listings 21.1 through 21.5 illustrate many of the techniques and considerations in Pentium optimization. Hand-optimization for the Pentium isn’t simple, and requires careful measurement to check the efficacy of your optimizations, so reserve it for when you really, really need it—but when you need it, you need it bad.

    -

    A Quick Note on the 386 and 486

    +

    A Quick Note on the 386 and 486

    I’ve mentioned that Pentium-optimized code does fine on the 486, but not always so well on the 386. On a 486, Listing 21.1 runs at 9 cycles per checksummed word, and Listing 21.5 runs at 2.5 cycles per checksummed word, a healthy 3.6-times speedup. On a 386, Listing 21.1 runs at 22 cycles per word; Listing 21.5 runs at 7 cycles per word, a 3.1-times speedup. As is often the case, Pentium optimization helped the other processors, but not as much as it helped the Pentium, and less on the 386 than on the 486.

    @@ -142,10 +135,6 @@ ckloopdone:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/22-01.html b/22-01.html index d41257a..aa5e739 100644 --- a/22-01.html +++ b/22-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Zenning and the Flexible Mind - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Zenning and the Flexible Mind - - @@ -37,22 +30,22 @@


    -

    Chapter 22
    +

    Chapter 22
    Zenning and the Flexible Mind

    -

    Taking a Spin through What You’ve Learned

    +

    Taking a Spin through What You’ve Learned

    And so we come to the end of our journey; for now, at least. What follows is a modest bit of optimization, one which originally served to show readers of Zen of Assembly Language that they had learned more than just bits and pieces of knowledge; that they had also begun to learn how to apply the flexible mind—unconventional, broadly integrative thinking—to approaching high-level optimization at the algorithmic and program design levels. You, of course, need no such reassurance, having just spent 21 chapters learning about the flexible mind in many guises, but I think you’ll find this example instructive nonetheless. Try to stay ahead as the level of optimization rises from instruction elimination to instruction substitution to more creative solutions that involve broader understanding and redesign. We’ll start out by compacting individual instructions and bits of code, but by the end we’ll come up with a solution that involves the very structure of the subroutine, with each instruction carefully integrated into a remarkably compact whole. It’s a neat example of how optimization operates at many levels, some much less determininstic than others—and besides, it’s just plain fun.

    Enjoy!

    -

    Zenning

    +

    Zenning

    In Jeff Duntemann’s excellent book Borland Pascal From Square One (Random House, 1993), there’s a small assembly subroutine that’s designed to be called from a Turbo Pascal program in order to fill the screen or a systemscreen buffer with a specified character/attribute pair in text mode. This subroutine involves only 21 instructions and works perfectly well; however, with what we know, we can compact the subroutine tremendously and speed it up a bit as well. To coin a verb, we can “Zen” this already-tight assembly code to an astonishing degree. In the process, I hope you’ll get a feel for how advanced your assembly skills have become.

    Jeff’s original code follows as Listing 22.1 (with some text converted to lowercase in order to match the style of this book), but the comments are mine.

    -

    LISTING 22.1 L22-1.ASM

    +

    LISTING 22.1 L22-1.ASM

     OnStack      struc      ;data that’s stored on the stack after PUSH BP
     OldBP        dw   ?     ;caller’s BP
    @@ -88,7 +81,7 @@ Bye:mov      sp,bp                   ;restore original stack pointer
         pop      bp                      ; and caller’s BP
         ret      EndMrk-RetAddr-2        ;return, clearing the parms from the stack
     ClearS       endp
    -
    +

    The first thing you’ll notice about Listing 22.1 is that ClearS uses a REP STOSW instruction. That means that we’re not going to improve performance by any great amount, no matter how clever we are. While we can eliminate some cycles, the bulk of the work in ClearS is done by that one repeated string instruction, and there’s no way to improve on that.

    @@ -115,10 +108,6 @@ ClearS endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/22-02.html b/22-02.html index d4238c6..943b525 100644 --- a/22-02.html +++ b/22-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Zenning and the Flexible Mind - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Zenning and the Flexible Mind - - @@ -37,7 +30,7 @@


    -

    LISTING 22.2 L22-2.ASM

    +

    LISTING 22.2 L22-2.ASM

     ClearS        proc near
           push    bp                        ;save caller’s BP
    @@ -60,7 +53,7 @@ Bye:
          pop       bp                       ;restore caller’s BP
          ret       EndMrk-RetAddr-2         ;return, clearing the parms from the stack
     ClearS         endp
    -
    +

    (The OnStack structure definition doesn’t change in any of our examples, so I’m not going clutter up this chapter by reproducing it for each new version of ClearS.)

    @@ -68,7 +61,7 @@ ClearS endp

    Well, LES would serve better than two MOV instructions for loading ES and DI as shown in Listing 22.3.

    -

    LISTING 22.3 L22-3.ASM

    +

    LISTING 22.3 L22-3.ASM

     ClearS         proc near
           push     bp                       ;save caller’s BP
    @@ -91,13 +84,13 @@ Bye:
           pop      bp                       ;restore caller’s BP
           ret      EndMrk-RetAddr-2         ;return, clearing the parms from the stack
     ClearS         endp
    -
    +

    That’s good for another three bytes. We’re down to 43 bytes, and counting.

    We can save 3 more bytes by clearing the low and high bytes of AX and BX, respectively, by using SUB reg8,reg8 rather than ANDing 16-bit values as shown in Listing 22.4.

    -

    LISTING 22.4 L22-4.ASM

    +

    LISTING 22.4 L22-4.ASM

     ClearS         proc near
           push     bp                       ;save caller’s BP
    @@ -120,7 +113,7 @@ Bye:
           pop      bp                       ;restore caller’s BP
           ret      EndMrk-RetAddr-2         ;return, clearing the parms from the stack
     ClearS         endp
    -
    +

    Now we’re down to 40 bytes—more than 20 percent smaller than the original code. That’s pretty much it for simple instruction optimizations. Now let’s look for instruction optimizations.

    @@ -145,10 +138,6 @@ ClearS endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/22-03.html b/22-03.html index b1ed376..635cbfa 100644 --- a/22-03.html +++ b/22-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Zenning and the Flexible Mind - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Zenning and the Flexible Mind - - @@ -37,7 +30,7 @@


    -

    LISTING 22.5 L22-5.ASM

    +

    LISTING 22.5 L22-5.ASM

     ClearS         proc near
           push     bp                       ;save caller’s BP
    @@ -56,13 +49,13 @@ Bye:
           pop      bp                       ;restore caller’s BP
           ret      EndMrk-RetAddr-2         ;return, clearing the parms from the stack
     ClearS         endp
    -
    +

    (We could get rid of yet another instruction by having the calling code pack both the attribute and the fill value into the same word, but that’s not part of the specification for this particular routine.)

    Another nifty instruction-rearrangement trick saves 6 more bytes. ClearS checks to see whether the far pointer is null (zero) at the start of the routine...then loads and uses that same far pointer later on. Let’s get that pointer into registers and keep it there; that way we can check to see whether it’s null with a single comparison, and can use it later without having to reload it from memory. This technique is shown in Listing 22.6.

    -

    LISTING 22.6 L22-6.ASM

    +

    LISTING 22.6 L22-6.ASM

     ClearS         proc near
           push     bp                       ;save caller’s BP
    @@ -80,7 +73,7 @@ Bye:
           pop      bp                       ;restore caller’s BP
           ret      EndMrk-RetAddr-2         ;return, clearing the parms from the stack
     ClearS         endp
    -
    +

    Well. Now we’re down to 28 bytes, having reduced the size of this subroutine by nearly 50 percent. Only 13 instructions remain. Realistically, how much smaller can we make this code?

    @@ -96,7 +89,7 @@ ClearS endp

    With that problem dealt with, Listing 22.7 shows the Zenned version of ClearS.

    -

    LISTING 22.7 L22-7.ASM

    +

    LISTING 22.7 L22-7.ASM

     ClearS         procnear
           pop      dx                  ;get the return address
    @@ -114,7 +107,7 @@ ClearS         procnear
     Bye:
           jmp      dx                  ;return to the calling code
     ClearS         endp
    -
    +

    At long last, we’re down to the bare metal. This version of ClearS is just 19 bytes long. That’s just 37 percent as long as the original version, without any change whatsoever in the functionality that ClearS makes available to the calling code. The code is bound to run a bit faster too, given that there are far fewer instruction bytes and fewer memory accesses.

    @@ -137,10 +130,6 @@ ClearS endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/23-01.html b/23-01.html index 0afec1f..e317652 100644 --- a/23-01.html +++ b/23-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - @@ -39,10 +32,10 @@

    Part II

    -

    Chapter 23
    +

    Chapter 23
    Bones and Sinew

    -

    At the Very Heart of Standard PC Graphics

    +

    At the Very Heart of Standard PC Graphics

    The VGA is unparalleled in the history of computer graphics, for it is by far the most widely-used graphics standard ever, the closest we may ever come to a lingua franca of computer graphics. No other graphics standard has even come close to the 50,000,000 or so VGAs in use today, and virtually every PC compatible sold today has full VGA compatibility built in. There are, of course, a variety of graphics accelerators that outperform the standard VGA, and indeed, it is becoming hard to find a plain vanilla VGA anymore—but there is no standard for accelerators, and every accelerator contains a true-blue VGA at its core.

    @@ -52,7 +45,7 @@

    We’ll start our exploration with a quick overview of the VGA, and then we’ll dive right in and get a taste of what the VGA can do.

    -

    The VGA

    +

    The VGA

    The VGA is the baseline adapter for modern IBM PC compatibles, present in virtually every PC sold today or in the last several years. (Note that the VGA is often nothing more than a chip on a motherboard, with some memory, a DAC, and maybe a couple of glue chips; nonetheless, I’ll refer to it as an adapter from now on for simplicity.) It guarantees that every PC is capable of documented resolutions up to 640x480 (with 16 possible colors per pixel) and 320x200 (with 256 colors per pixel), as well as undocumented—but nonetheless thoroughly standard—resolutions up to 360x480 in 256-color mode, as we’ll see in Chapters 31-34 and 47-49. In order for a video adapter to claim VGA compatibility, it must support all the features and code discussed in this book (with a very few minor exceptions that I’ll note)—and my experience is that just about 100 percent of the video hardware currently shipping or shipped since 1990 is in fact VGA compatible. Therefore, VGA code will run on nearly all of the 50,000,000 or so PC compatibles out there, with the exceptions being almost entirely obsolete machines from the 1980s. This makes good VGA code and VGA programming expertise valuable commodities indeed.

    @@ -64,7 +57,7 @@

    Let’s begin.

    -

    An Introduction to VGA Programming

    +

    An Introduction to VGA Programming

    Most discussions of the VGA start out with a traditional “Here’s a block diagram of the VGA” approach, with lists of registers and statistics. I’ll get to that eventually, but you can find it in IBM’s VGA documentation and several other books. Besides, it’s numbing to read specifications and explanations, and the VGA is an exciting adapter, the kind that makes you want to get your hands dirty probing under the hood, to write some nifty code just to see what the board can do. What’s more, the best way to understand the VGA is to see it work, so let’s jump right into a sample of the VGA in action, getting a feel for the VGA’s architecture in the process.

    @@ -87,10 +80,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/23-02.html b/23-02.html index 43cb264..659fe2b 100644 --- a/23-02.html +++ b/23-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - @@ -37,7 +30,7 @@


    -

    At the Core

    +

    At the Core

    A little background is necessary before we’re ready to examine Listing 23.1. The VGA is built around four functional blocks, named the CRT Controller (CRTC), the Sequence Controller (SC), the Attribute Controller (AC), and the Graphics Controller (GC). The single-chip VGA could have been designed to treat the registers for all the blocks as one large set, addressed at one pair of I/O ports, but in the EGA, each of these blocks was a separate chip, and the legacy of EGA compatibility is why each of these blocks has a separate set of registers and is addressed at different I/O ports in the VGA.

    @@ -185,12 +178,12 @@ A speed tip: The setting of each chip’s Index register remains the same until it is reprogrammed. This means that in cases where you are setting the same internal register repeatedly, you can set the Index register to point to that internal register once, then write to the Data register multiple times. For example, the Bit Mask register (GC register 8) is often set repeatedly inside a loop when drawing lines. The standard code for this is: - +
          MOV     DX,03CEH   ;point to GC Index register
          MOV     AL,8       ;internal index of Bit Mask register
          OUT     DX,AX      ;AH contains Bit Mask register setting
    -
    + @@ -198,13 +191,13 @@ -
    Alternatively, the GC Index register could initially be set to point to the Bit Mask register with
    +
          MOV     DX,03CEH  ;point to GC Index register
          MOV     AL,8      ;internal index of Bit Mask register
          OUT     DX,AL     ;set GC Index register
          INC     DX        ;point to GC Data register
    -
    + @@ -212,10 +205,10 @@ -
    and then the Bit Mask register could be set repeatedly with the byte-size OUT instruction
    +
          OUT     DX,AL    ;AL contains Bit Mask register setting
    -
    + @@ -242,10 +235,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/23-03.html b/23-03.html index fec7ca1..d6f8e3b 100644 --- a/23-03.html +++ b/23-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - @@ -37,7 +30,7 @@


    -

    Linear Planes and True VGA Modes

    +

    Linear Planes and True VGA Modes

    The VGA’s memory is organized as four 64K planes. Each of these planes is a linear bitmap; that is, each byte from a given plane controls eight adjacent pixels on the screen, the next byte controls the next eight pixels, and so on to the end of the scan line. The next byte then controls the first eight pixels of the next scan line, and so on to the end of the screen.

    @@ -49,13 +42,12 @@

    The VGA BIOS supports several graphics modes (modes 4, 5, and 6) in which VGA memory appears not to be organized as four linear planes. These modes exist for CGA compatibility only, and are not true VGA graphics modes; use them when you need CGA-type operation and ignore them the rest of the time. The VGA’s special features are most powerful in true VGA modes, and it is on the 16-color true-VGA modes (modes 0DH (320x200), 0EH (640x200), 10H (640x350), and 12H (640x480)) that I will concentrate in this part of the book. There is also a 256-color mode, mode 13H, that appears to be a single linear plane, but, as we will see in Chapters 31-34 and 47-49 of this book, that’s a polite fiction—and discarding that fiction gives us an opportunity to unleash the power of the VGA’s hardware for vastly better performance. VGA text modes, which feature soft fonts, are another matter entirely, upon which we’ll touch from time to time.

    -


    - Figure 23.1
      Video data from memory to pixel.

    +


    + Figure 23.1
      Video data from memory to pixel.

    With that background out of the way, we can get on to the sample VGA program shown in Listing 23.1. I suggest you run the program before continuing, since the explanations will mean far more to you if you’ve seen the features in action.

    -

    LISTING 23.1 L23-1.ASM

    +

    LISTING 23.1 L23-1.ASM

     ; Sample VGA program.
     ; Animates four balls bouncing around a playfield by using
    @@ -612,7 +604,7 @@ DrawBorderBlock endp
     AdjustPanning   endp
     cseg    ends
             end     start
    -
    +


    @@ -631,10 +623,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/23-04.html b/23-04.html index a0c958d..dc96bf1 100644 --- a/23-04.html +++ b/23-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - @@ -37,13 +30,12 @@


    -

    Smooth Panning

    +

    Smooth Panning

    The first thing you’ll notice upon running the sample program is the remarkable smoothness with which the display pans from side-to-side and up-and-down. That the display can pan at all is made possible by two VGA features: 256K of display memory and the virtual screen capability. Even the most memory-hungry of the VGA modes, mode 12H (640x480), uses only 37.5K per plane, for a total of 150K out of the total 256K of VGA memory. The medium-resolution mode, mode 10H (640x350), requires only 28K per plane, for a total of 112K. Consequently, there is room in VGA memory to store more than two full screens of video data in mode 10H (which the sample program uses), and there is room in all modes to store a larger virtual screen than is actually displayed. In the sample program, memory is organized as two virtual screens, each with a resolution of 672x384, as shown in Figure 23.2. The area of the virtual screen actually displayed at any given time is selected by setting the display memory address at which to begin fetching video data; this is set by way of the start address registers (Start Address High, CRTC register 0CH, and Start Address Low, CRTC register 0DH). Together these registers make up a 16-bit display memory address at which the CRTC begins fetching data at the beginning of each video frame. Increasing the start address causes higher-memory areas of the virtual screen to be displayed. For example, the Start Address High register could be set to 80H and the Start Address Low register could be set to 00H in order to cause the display screen to reflect memory starting at offset 8000H in each plane, rather than at the default offset of 0.

    -


    - Figure 23.2
      Video memory organization for Listing 23.1.

    +


    + Figure 23.2
      Video memory organization for Listing 23.1.

    The logical height of the virtual screen is defined by the amount of VGA memory available. As the VGA scans display memory for video data, it progresses from the start address toward higher memory one scan line at a time, until the frame is completed. Consequently, if the start address is increased, lines farther toward the bottom of the virtual screen are displayed; in effect, the virtual screen appears to scroll up on the physical screen.

    @@ -82,10 +74,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/23-05.html b/23-05.html index cb89408..d9df83d 100644 --- a/23-05.html +++ b/23-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - @@ -37,7 +30,7 @@


    -

    Color Plane Manipulation

    +

    Color Plane Manipulation

    The VGA provides a considerable amount of hardware assistance for manipulating the four display memory planes. Two features illustrated by the sample program are the ability to control which planes are written to by a CPU write and the ability to copy four bytes—one from each plane—with a single CPU read and a single CPU write.

    @@ -57,7 +50,7 @@

    Don’t worry if you’re not catching everything in this chapter on the first pass; the VGA is a complicated beast, and learning about it is an iterative process. We’ll be going over these features again, in different contexts, over the course of the rest of this book.

    -

    Page Flipping

    +

    Page Flipping

    When animated graphics are drawn directly on the screen, with no intermediate frame-composition stage, the image typically flickers and/or ripples, an unavoidable result of modifying display memory at the same time that it is being scanned for video data. The display memory of the VGA makes it possible to perform page flipping, which eliminates such problems. The basic premise of page flipping is that one area of display memory is displayed while another is being modified. The modifications never affect an area of memory as it is providing video data, so no undesirable side effects occur. Once the modification is complete, the modified buffer is selected for display, causing the screen to change to the new image in a single frame’s time, typically 1/60th or 1/70th of a second. The other buffer is then available for modification.

    @@ -86,10 +79,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/23-06.html b/23-06.html index 48e3691..ff8c7d7 100644 --- a/23-06.html +++ b/23-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bones and Sinew - - @@ -53,17 +46,17 @@

    To see the program run in 640x200 16-color mode, comment out the EQU line for MEDRES_VIDEO_MODE.

    -

    The Hazards of VGA Clones

    +

    The Hazards of VGA Clones

    Earlier, I said that any VGA that doesn’t support the features and functionality covered in this book can’t properly be called VGA compatible. I also noted that there are some exceptions, however, and we’ve just come to the most prominent one. You see, all VGAs really are compatible with the IBM VGA’s functionality when it comes to drawing pixels into display memory; all the write modes and read modes and set/reset capabilities and everything else involved with manipulating display memory really does work in the same way on all VGAs and VGA clones. That compatibility isn’t as airtight when it comes to scanning pixels out of display memory and onto the screen in certain infrequently-used ways, however.

    The areas of incompatibility of which I’m aware are illustrated by the sample program, and may in fact have caused you to see some glitches when you ran Listing 23.1. The problem, which arises only on certain VGAs, is that some settings of the Row Offset register cause some pixels to be dropped or displaced to the wrong place on the screen; often, this happens only in conjunction with certain start address settings. (In my experience, only VRAM (Video RAM)-based VGAs exhibit this problem, no doubt due to the way that pixel data is fetched from VRAM in large blocks.) Panning and large virtual bitmaps can be made to work reliably, by careful selection of virtual bitmap sizes and start addresses, but it’s difficult; that’s one of the reasons that most commercial software does not use these features, although a number of games do. The upshot is that if you’re going to use oversized virtual bitmaps and pan around them, you should take great care to test your software on a wide variety of VRAM- and DRAM-based VGAs.

    -

    Just the Beginning

    +

    Just the Beginning

    That pretty well covers the important points of the sample VGA program in Listing 23.1. There are many VGA features we didn’t even touch on, but the object was to give you a feel for the variety of features available on the VGA, to convey the flexibility and complexity of the VGA’s resources, and in general to give you an initial sense of what VGA programming is like. Starting with the next chapter, we’ll begin to explore the VGA systematically, on a more detailed basis.

    -

    The Macro Assembler

    +

    The Macro Assembler

    The code in this book is written in both C and assembly. I think C is a good development environment, but I believe that often the best code (although not necessarily the easiest to write or the most reliable) is written in assembly. This is especially true of graphics code for the x86 family, given segments, the string instructions, and the asymmetric and limited register set, and for real-time programming of a complex board like the VGA, there’s really no other choice for the lowest-level code.

    @@ -88,10 +81,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/24-01.html b/24-01.html index 8a01f29..3fb6bb3 100644 --- a/24-01.html +++ b/24-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Parallel Processing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Parallel Processing - - @@ -37,14 +30,14 @@


    -

    Chapter 24
    +

    Chapter 24
    Parallel Processing with the VGA

    -

    Taking on Graphics Memory Four Bytes at a Time

    +

    Taking on Graphics Memory Four Bytes at a Time

    This heading refers to the ability of the VGA chip to manipulate up to four bytes of display memory at once. In particular, the VGA provides four ALUs (Arithmetic Logic Units) to assist the CPU during display memory writes, and this hardware is a tremendous resource in the task of manipulating the VGA’s sizable frame buffer. The ALUs are actually only one part of the surprisingly complex data flow architecture of the VGA, but since they’re involved in almost all memory access operations, they’re a good place to begin.

    -

    VGA Programming: ALUs and Latches

    +

    VGA Programming: ALUs and Latches

    I’m going to begin our detailed tour of the VGA at the heart of the flow of data through the VGA: the four ALUs built into the VGA’s Graphics Controller (GC) circuitry. The ALUs (one for each display memory plane) are capable of ORing, ANDing, and XORing CPU data and display memory data together, as well as masking off some or all of the bits in the data from affecting the final result. All the ALUs perform the same logical operation at any given time, but each ALU operates on a different display memory byte.

    @@ -54,9 +47,8 @@

    Each ALU logically combines the byte written by the CPU and the byte stored in the matching latch, according to the settings of bits 3 and 4 of the Data Rotate register (and the Bit Mask register as well, which I’ll cover next time), and then writes the result to display memory. It is most important to understand that neither ALU operand comes directly from display memory. The temptation is to think of the ALUs as combining CPU data and the contents of the display memory address being written to, but they actually combine CPU data and the contents of the last display memory location read, which need not be the location being modified. The most common application of the ALUs is indeed to modify a given display memory location, but doing so requires a read from that location to load the latches before the write that modifies it. Omission of the read results in a write operation that logically combines CPU data n with whatever data happens to be in the latches from the last read, which is normally undesirable.

    -


    - Figure 24.1
      VGA ALU data flow.

    +


    + Figure 24.1
      VGA ALU data flow.

    Occasionally, however, the independence of the latches from the display memory location being written to can be used to great advantage. The latches can be used to perform 4-byte-at-a-time (one byte from each plane) block copying; in this application, the latches are loaded with a read from the source area and written unmodified to the destination area. The latches can be written unmodified in one of two ways: By selecting write mode 1 (for an example of this, see the last chapter), or by setting the Bit Mask register to 0 so only the latched bits are written.

    @@ -83,10 +75,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/24-02.html b/24-02.html index 52499b2..fe593de 100644 --- a/24-02.html +++ b/24-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Parallel Processing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Parallel Processing - - @@ -37,7 +30,7 @@


    -

    LISTING 24.1 L24-1.ASM

    +

    LISTING 24.1 L24-1.ASM

     ; Program to illustrate operation of ALUs and latches of the VGA’s
     ;  Graphics Controller.  Draws a variety of patterns against
    @@ -291,7 +284,7 @@ DrawVerticalBox endp
     cseg    ends
             end     start
     
    -
    +


    @@ -310,10 +303,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/24-03.html b/24-03.html index 2d86f21..2ce14ca 100644 --- a/24-03.html +++ b/24-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Parallel Processing - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Parallel Processing - - @@ -41,7 +34,7 @@

    Logical functions 1 through 3 cause the CPU data to be ANDed, ORed, and XORed with the latched data, respectively. Of these, XOR is the most useful, since exclusive-ORing is a traditional way to perform animation. The uses of the AND and OR logical functions are less obvious. AND can be used to mask a blank area into display memory, or to mask off those portions of a drawing operation that don’t overlap an existing display memory image. OR could conceivably be used to force an image into display memory over an existing image. To be honest, I haven’t encountered any particularly valuable applications for AND and OR, but they’re the sort of building-block features that could come in handy in just the right context, so keep them in mind.

    -

    Notes on the ALU/Latch Demo Program

    +

    Notes on the ALU/Latch Demo Program

    VGA settings such as the logical function select should be restored to their default condition before the BIOS is called to output text or draw pixels. The VGA BIOS does not guarantee that it will set most VGA registers except on mode sets, and there are so many compatible BIOSes around that the code of the IBM BIOS is not a reliable guide. For instance, when the BIOS is called to draw text, it’s likely that the result will be illegible if the Bit Mask register is not in its default state. Similarly, a mode set should generally be performed before exiting a program that tinkers with VGA settings.

    @@ -53,16 +46,16 @@

    All text in the sample program is drawn by VGA BIOS function 13H, the write string function. This function is also present in the AT’s BIOS, but not in the XT’s or PC’s, and as a result is rarely used; the function is always available if a VGA is installed, however. Text drawn with this function is relatively slow. If speed is important, a program can draw text directly into display memory much faster in any given display mode. The great virtue of the BIOS write string function in the case of the VGA is that it provides an uncomplicated way to get text on the screen reliably in any mode and color, over any background.

    -

    The expression used to load DX in the TEXT_UP macro in the sample program may seem strange, but it’s a convenient way to save a byte of program code and a few cycles of execution time. DX is being loaded with a word value that’s composed of two independent immediate byte values. The obvious way to implement this would be with

    +

    The expression used to load DX in the TEXT_UP macro in the sample program may seem strange, but it’s a convenient way to save a byte of program code and a few cycles of execution time. DX is being loaded with a word value that’s composed of two independent immediate byte values. The obvious way to implement this would be with

     MOV DL,VALUE1
     MOV DH,VALUE2
    -
    + -

    which requires four instruction bytes. By shifting the value destined for the high byte into the high byte with MASM’s shift-left operator, SHL (*100H would work also), and then logically combining the values with MASM’s OR operator (or the ADD operator), both halves of DX can be loaded with a single instruction, as in

    +

    which requires four instruction bytes. By shifting the value destined for the high byte into the high byte with MASM’s shift-left operator, SHL (*100H would work also), and then logically combining the values with MASM’s OR operator (or the ADD operator), both halves of DX can be loaded with a single instruction, as in

     MOV DX,(VALUE2 SHL 8) OR VALUE1
    -
    +

    which takes only three bytes and is faster, being a single instruction. (Note, though, that in 32-bit protected mode, there’s a size and performance penalty for 16-bit instructions such as the MOV above; see the first part of this book for details.) As shown, a macro is an ideal place to use this technique; the macro invocation can refer to two separate byte values, making matters easier for the programmer, while the macro itself can combine the values into a single word-sized constant.

    @@ -93,10 +86,6 @@ MOV DX,(VALUE2 SHL 8) OR VALUE1
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/25-01.html b/25-01.html index 7ff1834..850afb7 100644 --- a/25-01.html +++ b/25-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - @@ -37,32 +30,30 @@


    -

    Chapter 25
    +

    Chapter 25
    VGA Data Machinery

    -

    The Barrel Shifter, Bit Mask, and Set/Reset Mechanisms

    +

    The Barrel Shifter, Bit Mask, and Set/Reset Mechanisms

    In the last chapter, we examined a simplified model of data flow within the GC portion of the VGA, featuring the latches and ALUs. Now we’re ready to expand that model to include the barrel shifter, bit mask, and the set/reset capabilities, leaving only the write modes to be explored over the next few chapters.

    -

    VGA Data Rotation

    +

    VGA Data Rotation

    Figure 25.1 shows an expanded model of GC data flow, featuring the barrel shifter and bit mask circuitry. Let’s look at the barrel shifter first. A barrel shifter is circuitry capable of shifting—or rotating, in the VGA’s case—data an arbitrary number of bits in a single operation, as opposed to being able to shift only one bit position at a time. The barrel shifter in the VGA can rotate incoming CPU data up to seven bits to the right (toward the least significant bit), with bit 0 wrapping back to bit 7, after which the VGA continues processing the rotated byte just as it normally processes unrotated CPU data. Thanks to the nature of barrel shifters, this rotation requires no extra processing time over unrotated VGA operations. The number of bits by which CPU data is shifted is controlled by bits 2-0 of GC register 3, the Data Rotate register, which also contains the ALU function select bits (data unmodified, AND, OR, and XOR) that we looked at in the last chapter.

    -


    - Figure 25.1
      Data flow through the Graphics Controller.

    +


    + Figure 25.1
      Data flow through the Graphics Controller.

    The barrel shifter is powerful, but (as sometimes happens in this business) it sounds more useful than it really is. This is because the GC can only rotate CPU data, a task that the CPU itself is perfectly capable of performing. Two OUTs are needed to select a given rotation: one to set the GC Index register, and one to set the Data Rotate register. However, with careful programming it’s sometimes possible to leave the GC Index always pointing to the Data Rotate register, so only one OUT is needed. Even so, it’s often easier and/or faster to simply have the CPU rotate the data of interest CL times than to set the Data Rotate register. (Bear in mind that a single OUT takes from 11 to 31 cycles on a 486—and longer if the VGA is sluggish at responding to OUTs, as many VGAs are.) If only the VGA could rotate latched data, then there would be all sorts of useful applications for rotation, but, sadly, only CPU data can be rotated.

    The drawing of bit-mapped text is one use for the barrel shifter, and I’ll demonstrate that application below. In general, though, don’t knock yourself out trying to figure out how to work data rotation into your programs—it just isn’t all that useful in most cases.

    -

    The Bit Mask

    +

    The Bit Mask

    The VGA has bit mask circuitry for each of the four memory planes. The four bit masks operate in parallel and are all driven by the same mask data for each operation, so they’re generally referred to in the singular, as “the bit mask.” Figure 25.2 illustrates the operation of one bit of the bit mask for one plane. This circuitry occurs eight times in the bit mask for a given plane, once for each bit of the byte written to display memory. Briefly, the bit mask determines on a bit-by-bit basis whether the source for each byte written to display memory is the ALU for that plane or the latch for that plane.

    -


    - Figure 25.2
      Bit mask operation.

    +


    + Figure 25.2
      Bit mask operation.

    The bit mask is controlled by GC register 8, the Bit Mask register. If a given bit of the Bit Mask register is 1, then the corresponding bit of data from the ALUs is written to display memory for all four planes, while if that bit is 0, then the corresponding bit of data from the latches for the four planes is written to display memory unchanged. (In write mode 3, the actual bit mask that’s applied to data written to display memory is the logical AND of the contents of the Bit Mask register and the data written by the CPU, as we’ll see in Chapter 26.)

    @@ -87,10 +78,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/25-02.html b/25-02.html index b7e392a..b361f29 100644 --- a/25-02.html +++ b/25-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - @@ -37,7 +30,7 @@


    -

    LISTING 25.1 L25-1.ASM

    +

    LISTING 25.1 L25-1.ASM

     ; Program to illustrate operation of data rotate and bit mask
     ;  features of Graphics Controller. Draws 8x8 character at
    @@ -269,7 +262,7 @@ SelectFont      endp
     ;
     cseg    ends
             end     start
    -
    +

    The bit mask can be used for much more than bit-aligned fonts. For example, the bit mask is useful for fast pixel drawing, such as that performed when drawing lines, as we’ll see in Chapter 35. It’s also useful for drawing the edges of primitives, such as filled polygons, that potentially involve modifying some but not all of the pixels controlled by a single byte of display memory.

    @@ -292,10 +285,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/25-03.html b/25-03.html index cc9fe63..6025900 100644 --- a/25-03.html +++ b/25-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - @@ -53,15 +46,14 @@

    He’s got a point there.

    -

    The VGA’s Set/Reset Circuitry

    +

    The VGA’s Set/Reset Circuitry

    At last we come to the final aspect of data flow through the GC on write mode 0 writes: the set/reset circuitry. Figure 25.3 shows data flow on a write mode 0 write. The only difference between this figure and Figure 25.1 is that on its way to each plane potentially the rotated CPU data passes through the set/reset circuitry, which may or may not replace the CPU data with set/reset data. Briefly put, the set/reset circuitry enables the programmer to elect to independently replace the CPU data for each plane with either 00 or 0FFH.

    What is the use of such a feature? Well, the standard way to control color is to set the Map Mask register to enable writes to only those planes that need to be set to produce the desired color. For example, the Map Mask register would be set to 09H to draw in high-intensity blue; here, bits 0 and 3 are set to 1, so only the blue plane (plane 0) and the intensity plane (plane 3) are written to.

    -


    - Figure 25.3
      Data flow during a write mode 0 write operation.

    +


    + Figure 25.3
      Data flow during a write mode 0 write operation.

    Remember, though, that planes that are disabled by the Map Mask register are not written to or modified in any way. This means that the above approach works only if the memory being written to is zeroed; if, however, the memory already contains non-zero data, that data will remain in the planes disabled by the Map Mask, and the end result will be that some planes contain the data just written and other planes contain old data. In short, color control using the Map Mask does not force all planes to contain the desired color. In particular, it is not possible to force some planes to zero and other planes to one in a single write with the Map Mask register.

    @@ -84,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/25-04.html b/25-04.html index 1413466..95d619b 100644 --- a/25-04.html +++ b/25-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - @@ -37,7 +30,7 @@


    -

    LISTING 25.2 L25-2.ASM

    +

    LISTING 25.2 L25-2.ASM

     ; Program to illustrate operation of Map Mask register when drawing
     ;  to memory that already contains data.
    @@ -120,9 +113,9 @@ HorzBarLoop:
     start   endp
     cseg    ends
             end     start
    -
    + -

    Setting All Planes to a Single Color

    +

    Setting All Planes to a Single Color

    The set/reset circuitry can be used to force some planes to 0-bits and others to 1-bits during a single write, while letting CPU data go to still other planes, and so provides an efficient way to set all planes to a desired color. The set/reset circuitry works as follows:

    @@ -149,10 +142,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/25-05.html b/25-05.html index 18aea62..08a757f 100644 --- a/25-05.html +++ b/25-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - @@ -37,7 +30,7 @@


    -

    LISTING 25.3 L25-3.ASM

    +

    LISTING 25.3 L25-3.ASM

     ; Program to illustrate operation of set/reset circuitry to force
     ;  setting of memory that already contains data.
    @@ -147,15 +140,15 @@ HorzBarLoop:
     start   endp
     cseg    ends
             end     start
    -
    + -

    Manipulating Planes Individually

    +

    Manipulating Planes Individually

    Listing 25.4 illustrates the use of set/reset to control only some, rather than all, planes. Here, the set/reset circuitry forces plane 2 to 1 and planes 0 and 3 to 0. Because bit 1 of the Enable Set/Reset register is 0, however, set/reset does not affect plane 1; the CPU data goes unchanged to the plane 1 ALU. Consequently, the CPU data can be used to control the value written to plane 1. Given the settings of the other three planes, this means that each bit of CPU data that is 1 generates a brown pixel, and each bit that is 0 generates a red pixel. Writing alternating bytes of 07H and 0E0H, then, creates a vertically striped pattern of brown and red.

    In Listing 25.4, note that the vertical bars are 10 and 6 bytes wide, and do not start on byte boundaries. Although set/reset replaces an entire byte of CPU data for a plane, the combination of set/reset for some planes and CPU data for other planes, as in the example above, can be used to control individual pixels.

    -

    LISTING 25.4 L25-4.ASM

    +

    LISTING 25.4 L25-4.ASM

     ; Program to illustrate operation of set/reset circuitry in conjunction
     ;  with CPU data to modify setting of memory that already contains data.
    @@ -265,7 +258,7 @@ HorzBarLoop:
     start   endp
     cseg    ends
             end     start
    -
    +


    @@ -284,10 +277,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/25-06.html b/25-06.html index a013350..6804722 100644 --- a/25-06.html +++ b/25-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Data Machinery - - @@ -39,7 +32,7 @@

    There is no clearly defined role for the set/reset circuitry, as there is for, say, the bit mask. In many cases, set/reset is largely interchangeable with CPU data, particularly with CPU data written in write mode 2 (write mode 2 operates similarly to the set/reset circuitry, as we’ll see in Chapter 27). The most powerful use of set/reset, in my experience, is in applications such as the example of Listing 25.4, where it is used to force the value written to certain planes while the CPU data is written to other planes. In general, though, think of set/reset as one more tool you have at your disposal in getting the VGA to do what you need done, in this case a tool that lets you force all bits in each plane to either zero or one, or pass CPU data through unchanged, on each write to display memory. As tools go, set/reset is a handy one, and it’ll pop up often in this book.

    -

    Notes on Set/Reset

    +

    Notes on Set/Reset

    The set/reset circuitry is not active in write modes 1 or 2. The Enable Set/Reset register is inactive in write mode 3, but the Set/Reset register provides the primary drawing color in write mode 3, as discussed in the next chapter.

    @@ -51,7 +44,7 @@
    -

    A Brief Note on Word OUTs

    +

    A Brief Note on Word OUTs

    In the early days of the EGA and VGA, there was considerable debate about whether it was safe to do word OUTs (OUT DX,AX) to set Index/Data register pairs in a single instruction. Long ago, there were a few computers with buses that weren’t quite PC-compatatible, in that the two bytes in each word OUT went to the VGA in the wrong order: Data register first, then Index register, with predictably disastrous results. Consequently, I generally wrote my code in those days to use two 8-bit OUTs to set indexed registers. Later on, I made it a habit to use macros that could do either one 16-bit OUT or two 8-bit OUTs, depending on how I chose to assemble the code, and in fact you’ll find both ways of dealing with OUTs sprinkled through the code in this part of the book. Using macros for word OUTs is still not a bad idea in that it does no harm, but in my opinion it’s no longer necessary. Word OUTs are standard now, and it’s been a long time since I’ve heard of them causing any problems.

    @@ -72,10 +65,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/26-01.html b/26-01.html index 8df93be..97550a4 100644 --- a/26-01.html +++ b/26-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Write Mode - - @@ -37,16 +30,16 @@


    -

    Chapter 26
    +

    Chapter 26
    VGA Write Mode 3

    -

    The Write Mode That Grows on You

    +

    The Write Mode That Grows on You

    Over the last three chapters, we’ve covered the VGA’s write path from stem to stern—with one exception. Thus far, we’ve only looked at how writes work in write mode 0, the straightforward, workhorse mode in which each byte that the CPU writes to display memory fans out across the four planes. (Actually, we also took a quick look at write mode 1, in which the latches are always copied unmodified, but since exactly the same result can be achieved by setting the Bit Mask register to 0 in write mode 0, write mode 1 is of little real significance.)

    Write mode 0 is a very useful mode, but some of VGA’s most interesting capabilities involve the two write modes that we have yet to examine: write mode 1, and, especially, write mode 3. We’ll get to write mode 1 in the next chapter, but right now I want to focus on write mode 3, which can be confusing at first, but turns out to be quite a bit more powerful than one might initially think.

    -

    A Mode Born in Strangeness

    +

    A Mode Born in Strangeness

    Write mode 3 is strange indeed, and its use is not immediately obvious. The first time I encountered write mode 3, I understood immediately how it functioned, but could think of very few useful applications for it. As time passed, and as I came to understand the atrocious performance characteristics of OUT instructions, and the importance of text and pattern drawing as well, write mode 3 grew considerably in my estimation. In fact, my esteem for this mode ultimately reached the point where in the last major chunk of 16-color graphics code I wrote, write mode 3 was used more than write mode 0 overall, excluding simple pixel copying. So write mode 3 is well worth using, but to use it you must first understand it. Here’s how it works.

    @@ -73,10 +66,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/26-02.html b/26-02.html index b723c42..f1b767c 100644 --- a/26-02.html +++ b/26-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Write Mode - - @@ -37,7 +30,7 @@


    -

    LISTING 26.1 L26-1.ASM

    +

    LISTING 26.1 L26-1.ASM

     ; Program to illustrate operation of write mode 3 of the VGA.
     ;  Draws 8x8 characters at arbitrary locations without disturbing
    @@ -332,7 +325,7 @@ SelectFont      endp
     ;
     cseg    ends
             end     start
    -
    +


    @@ -351,10 +344,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/26-03.html b/26-03.html index 7ce3f5e..31500bf 100644 --- a/26-03.html +++ b/26-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: VGA Write Mode - - @@ -55,7 +48,7 @@

    The write mode 3 approach used in Listing 26.1 can be efficiently extended to drawing large blocks of text. For example, suppose that we were to draw a line of 8-pixel-wide bit-mapped text 40 characters long. We could then set up the bit mask and data rotation as appropriate for the left portion of each bit-aligned character (the portion of each character to the left of the byte boundary) and then draw the left portions only of all 40 characters in write mode 3. Then the bit mask could be set up for the right portion of each character, and the right portions of all 40 characters could be drawn. The VGA’s fast rotator would be used to do all rotation, and the only OUTs required would be those required to set the bit mask and data rotation. This technique could well outperform single-character bit-mapped text drivers such as the one in Listing 26.1 by a significant margin. Listing 26.2 illustrates one implementation of such an approach. Incidentally, note the use of the 8x14 ROM font in Listing 26.2, rather than the 8x8 ROM font used in Listing 26.1. There is also an 8x16 font stored in ROM, along with the tables used to alter the 8x14 and 8x16 ROM fonts into 9x14 and 9x16 fonts.

    -

    LISTING 26.2 L26-2.ASM

    +

    LISTING 26.2 L26-2.ASM

     ; Program to illustrate high-speed text-drawing operation of
     ;  write mode 3 of the VGA.
    @@ -375,7 +368,7 @@ SelectFont      endp
     ;
     cseg    ends
             end     start
    -
    +

    In this chapter, I’ve tried to give you a feel for how write mode 3 works and what it might be used for, rather than providing polished, optimized, plug-it-in-and-go code. Like the rest of the VGA’s write path, write mode 3 is a resource that can be used in a remarkable variety of ways, and I don’t want to lock you into thinking of it as useful in just one context. Instead, you should take the time to thoroughly understand what write mode 3 does, and then, when you do VGA programming, think about how write mode 3 can best be applied to the task at hand. Because I focused on illustrating the operation of write mode 3, neither listing in this chapter is the fastest way to accomplish the desired result. For example, Listing 26.2 could be made nearly twice as fast by simply having the CPU rotate, mask, and join the bytes from adjacent characters, then draw the combined bytes to display memory in a single operation.

    @@ -383,7 +376,7 @@ cseg ends

    As a final note, consider that non-transparent text could also be accelerated with write mode 3. The latches could be filled with the background (text box) color, set/reset could be set to the foreground (text) color, and write mode 3 could then be used to turn monochrome text bytes written by the CPU into characters on the screen with just one write per byte. There are complications, such as drawing partial bytes, and rotating the bytes to align the characters, which we’ll revisit later on in Chapter 55, while we’re working through the details of the X-Sharp library. Nonetheless, the performance benefit of this approach can be a speedup of as much as four times—all thanks to the decidedly quirky but surprisingly powerful and flexible write mode 3.

    -

    A Note on Preserving Register Bits

    +

    A Note on Preserving Register Bits

    If you take a quick look, you’ll see that the code in Listing 26.1 uses the readable register feature of the VGA to preserve reserved bits and bits other than those being modified. Older adapters such as the CGA and EGA had few readable registers, so it was necessary to set all bits in a register whenever that register was modified. Happily, all VGA registers are readable, which makes it possible to change only those bits of immediate interest, and, in general, I highly recommend doing exactly that, since IBM (or clone manufacturers) may well someday use some of those reserved bits or change the meanings of some of the bits that are currently in use.

    @@ -404,10 +397,6 @@ cseg ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/27-01.html b/27-01.html index 42bfb7c..e393801 100644 --- a/27-01.html +++ b/27-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - @@ -37,10 +30,10 @@


    -

    Chapter 27
    +

    Chapter 27
    Yet Another VGA Write Mode

    -

    Write Mode 2, Chunky Bitmaps,and Text-Graphics Coexistence

    +

    Write Mode 2, Chunky Bitmaps,and Text-Graphics Coexistence

    In the last chapter, we learned about the markedly peculiar write mode 3 of the VGA, after having spent three chapters learning the ins and outs of the VGA’s data path in write mode 0, touching on write mode 1 as well in Chapter 23. In all, the VGA supports four write modes—write modes 0, 1, 2, and 3—and read modes 0 and 1 as well. Which leaves two burning questions: What is write mode 2, and how the heck do you read VGA memory?

    @@ -48,7 +41,7 @@

    Let’s start with the easy stuff, write mode 2, and save the read modes for the next chapter.

    -

    Write Mode 2 and Set/Reset

    +

    Write Mode 2 and Set/Reset

    Remember how set/reset works? Good, because that’s pretty much how write mode 2 works. (You don’t remember? Well, I’ll provide a brief refresher, but I suggest that you go back through Chapters 23 through 25 and come up to speed on the VGA.)

    @@ -58,15 +51,14 @@

    It’s possible that you understand write mode 2 thoroughly at this point; nonetheless, I suspect that some additional explanation of an admittedly non-obvious mode wouldn’t hurt. Let’s follow the CPU byte through the VGA in write mode 2, step by step.

    -

    A Byte’s Progress in Write Mode 2

    +

    A Byte’s Progress in Write Mode 2

    Figure 27.1 shows the write mode 2 data path. The CPU byte comes into the VGA and is split into four separate bits, one for each plane. Bits 7-4 of the CPU byte vanish into the bit bucket, never to be heard from again. Speculation long held that those 4 unused bits indicated that IBM would someday come out with an 8-plane adapter that supported 256 colors. When IBM did finally come out with a 256-color mode (mode 13H of the VGA), it turned out not to be planar at all, and the upper nibble of the CPU byte remains unused in write mode 2 to this day.

    The bit of the CPU byte sent to each plane is expanded to a 0 or 0FFH byte, depending on whether the bit is 0 or 1, respectively. The byte for each plane then becomes the CPU-side input to the respective plane’s ALU. From this point on, the write mode 2 data path is identical to the write mode 0 data path. As discussed in earlier articles, the latch byte for each plane is the other ALU input, and the ALU either ANDs, ORs, or XORs the two bytes together or simply passes the CPU-side byte through. The byte generated by each plane’s ALU then goes through the bit mask circuitry, which selects on a bit-by-bit basis between the ALU byte and the latch byte. Finally, the byte from the bit mask circuitry for each plane is written to that plane if the corresponding bit in the Map Mask register is set to 1.

    -


    - Figure 27.1
      VGA data flow in write mode 2.

    +


    + Figure 27.1
      VGA data flow in write mode 2.

    @@ -99,10 +91,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/27-02.html b/27-02.html index d09cee8..02f8eab 100644 --- a/27-02.html +++ b/27-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - @@ -37,7 +30,7 @@


    -

    Copying Chunky Bitmaps to VGA Memory Using Write Mode 2

    +

    Copying Chunky Bitmaps to VGA Memory Using Write Mode 2

    Let’s take a look at two examples of write mode 2 in action. Listing 27.1 presents a program that uses write mode 2 to copy a graphics image in chunky format to the VGA. In chunky format adjacent bits in a single byte make up each pixel: mode 4 of the CGA, EGA, and VGA is a 2-bit-per-pixel chunky mode, and mode 13H of the VGA is an 8-bit-per-pixel chunky mode. Chunky format is convenient, since all the information about each pixel is contained in a single byte; consequently chunky format is often used to store bitmaps in system memory.

    @@ -47,7 +40,7 @@

    This process is then repeated for the rightmost chunky pixel, if necessary, and repeated again for as many pixels as there are in the image.

    -

    LISTING 27.1 L27-1.ASM

    +

    LISTING 27.1 L27-1.ASM

     ; Program to illustrate one use of write mode 2 of the VGA and EGA by
     ; animating the image of an “A” drawn by copying it from a chunky
    @@ -266,7 +259,7 @@ CheckMoreScanLines:
     DrawFromChunkyBitmap    endp
     Code    ends
             end     Start
    -
    +


    @@ -285,10 +278,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/27-03.html b/27-03.html index 86d423b..df16049 100644 --- a/27-03.html +++ b/27-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - @@ -47,13 +40,13 @@
    -

    Drawing Color-Patterned Lines Using Write Mode 2

    +

    Drawing Color-Patterned Lines Using Write Mode 2

    A more serviceable use of write mode 2 is shown in the program presented in Listing 27.2. The program draws multicolored horizontal, vertical, and diagonal lines, basing the color patterns on passed color tables. Write mode 2 is ideal because in this application color can vary from one pixel to the next, and in write mode 2 all that’s required to set pixel color is a change of the lower nibble of the byte written by the CPU. Set/reset could be used to achieve the same result, but an index/data pair of OUTs would be required to set the Set/Reset register to each new color. Similarly, the Map Mask register could be used in write mode 0 to set pixel color, but in this case not only would an index/data pair of OUTs be required but there would also be no guarantee that data already in display memory wouldn’t interfere with the color of the pixel being drawn, since the Map Mask register allows only selected planes to be drawn to.

    Listing 27.2 is hardly a comprehensive line drawing program. It draws only a few special line cases, and although it is reasonably fast, it is far from the fastest possible code to handle those cases, because it goes through a dot-plot routine and because it draws horizontal lines a pixel rather than a byte at a time. Write mode 2 would, however, serve just as well in a full-blown line drawing routine. For any type of patterned line drawing on the VGA, the basic approach remains the same: Use the bit mask to select the pixel (or pixels) to be altered and use the CPU byte in write mode 2 to select the color in which to draw.

    -

    LISTING 27.2 L27-2.ASM

    +

    LISTING 27.2 L27-2.ASM

     ; Program to illustrate one use of write mode 2 of the VGA and EGA by
     ; drawing lines in color patterns.
    @@ -381,7 +374,7 @@ DotUpInColor    endp
     Start   endp
     Code    ends
             end     Start
    -
    +


    @@ -400,10 +393,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/27-04.html b/27-04.html index 3e52ccb..14175a5 100644 --- a/27-04.html +++ b/27-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - @@ -37,7 +30,7 @@


    -

    When to Use Write Mode 2 and When to Use Set/Reset

    +

    When to Use Write Mode 2 and When to Use Set/Reset

    As indicated earlier, write mode 2 and set/reset are functionally interchangeable. Write mode 2 lends itself to more efficient implementations when the drawing color changes frequently, as in Listing 27.2.

    @@ -45,11 +38,11 @@

    Set/reset is also the mode of choice whenever it is necessary to force the value written to some planes to a fixed value while allowing the CPU byte to modify other planes. This is the mode of operation when set/reset is enabled for some but not all planes.

    -

    Mode 13H—320x200 with 256 Colors

    +

    Mode 13H—320x200 with 256 Colors

    I’m going to take a minute—and I do mean a minute—to discuss the programming model for mode 13H, the VGA’s 320x200 256-color mode. Frankly, there’s just not much to it, especially compared to the convoluted 16-color model that we’ve explored over the last five chapters. Mode 13H offers the simplest programming model in the history of PC graphics: A linear bitmap starting at A000:0000, consisting of 64,000 bytes, each controlling one pixel. The byte at offset 0 controls the upper left pixel on the screen, the byte at offset 319 controls the upper right pixel on the screen, the byte at offset 320 controls the second pixel down at the left of the screen, and the byte at offset 63,999 controls the lower right pixel on the screen. That’s all there is to it; it’s so simple that I’m not going to spend any time on a demo program, especially given that some of the listings later in this book, such as the antialiasing code in Chapter F on the companion CD-ROM, use mode 13H.

    -

    Flipping Pages from Text to Graphics and Back

    +

    Flipping Pages from Text to Graphics and Back

    A while back, I got an interesting letter from Phil Coleman, of La Jolla, who wrote:

    @@ -86,10 +79,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/27-05.html b/27-05.html index 14716eb..6fb2428 100644 --- a/27-05.html +++ b/27-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yet Another VGA Write Mode - - @@ -43,7 +36,7 @@

    As I said, Phil Coleman’s question is an interesting one, and I’ve only touched on the intriguing possibilities arising from the various configurations of display memory in VGA graphics and text modes. Right now, though, we’ve still got the basics of the remarkably complex (but rewarding!) VGA to cover.

    -

    LISTING 27.3 L27-3.ASM

    +

    LISTING 27.3 L27-3.ASM

     ; Program to illustrate flipping from bit-mapped graphics mode to
     ; text mode and back without losing any of the graphics bit-map.
    @@ -236,7 +229,7 @@ FillBitMap:
     Start   endp
     Code    ends
             end     Start
    -
    +


    @@ -255,10 +248,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/28-01.html b/28-01.html index 66bbc86..1b44557 100644 --- a/28-01.html +++ b/28-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - @@ -37,16 +30,16 @@


    -

    Chapter 28
    +

    Chapter 28
    Reading VGA Memory

    -

    Read Modes 0 and 1, and the Color Don’t Care Register

    +

    Read Modes 0 and 1, and the Color Don’t Care Register

    Well, it’s taken five chapters, but we’ve finally covered the data write path and all four write modes of the VGA. Now it’s time to tackle the VGA’s two read modes. While the read modes aren’t as complex as the write modes, they’re nothing to sneeze at. In particular, read mode 1 (also known as color compare mode) is rather unusual and not at all intuitive.

    You may well ask, isn’t anything about programming the VGA straightforward? Well...no. But then, clearing up the mysteries of VGA programming is what this part of the book is all about, so let’s get started.

    -

    Read Mode 0

    +

    Read Mode 0

    Read mode 0 is actually relatively uncomplicated, given that you understand the four-plane nature of the VGA. (If you don’t understand the four-plane nature of the VGA, I strongly urge you to read Chapters 23-27 before continuing with this chapter.) Read mode 0, the read mode counterpart of write mode 0, lets you read from one (and only one) plane of VGA memory at any one time.

    @@ -75,10 +68,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/28-02.html b/28-02.html index 1cbd903..72c0c02 100644 --- a/28-02.html +++ b/28-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - @@ -37,7 +30,7 @@


    -

    LISTING 28.1 L28-1.ASM

    +

    LISTING 28.1 L28-1.ASM

     ; Program to illustrate the use of the Read Map register in read mode 0.
     ; Animates by copying a 16-color image from VGA memory to system memory,
    @@ -274,7 +267,7 @@ GetImageOffset proc near
     GetImageOffset endp
     code  ends
          end  Start
    -
    +


    @@ -293,10 +286,6 @@ code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/28-03.html b/28-03.html index 21c07bd..ff97191 100644 --- a/28-03.html +++ b/28-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - @@ -49,7 +42,7 @@ -

    Read Mode 1

    +

    Read Mode 1

    Read mode 0 is the workhorse read mode, but it’s got an annoying limitation: Whenever you want to determine the color of a given pixel in read mode 0, you have to perform four VGA memory reads, one for each plane, and then interpret the four bytes you’ve read as eight 16-color pixels. That’s a lot of programming. The code is also likely to run slowly, all the more so because a standard IBM VGA takes an average of 1.1 microseconds to complete each memory read, and read mode 0 requires four reads in order to read the four planes, not to mention the even greater amount of time taken by the OUTs required to switch between the planes. (1.1 microseconds may not sound like much, but on a 66-MHz 486, it’s 73 clock cycles! Local-bus VGAs can be a good deal faster, but a read from the fastest local-bus adapter I’ve yet seen would still cost in the neighborhood of 10 486/66 cycles.)

    @@ -76,10 +69,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/28-04.html b/28-04.html index d102b56..632ee75 100644 --- a/28-04.html +++ b/28-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - @@ -37,7 +30,7 @@


    -

    LISTING 28.2 L28-2.ASM

    +

    LISTING 28.2 L28-2.ASM

     ; Program to illustrate use of read mode 1 (color compare mode)
     ; to detect collisions in display memory. Draws a yellow line on a
    @@ -200,9 +193,9 @@ SelectSetResetColorprocnear
     SelectSetResetColorendp
     code ends
     end  Start
    -
    + -

    When all Planes “Don’t Care”

    +

    When all Planes “Don’t Care”

    Still and all, there aren’t all that many uses for basic color compare operations. There is, however, a genuinely odd application of read mode 1 that’s worth knowing about; but in order to understand that, we must first look at the “don’t care” aspect of color compare operation.

    @@ -233,10 +226,6 @@ end Start
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/28-05.html b/28-05.html index f76742f..5501705 100644 --- a/28-05.html +++ b/28-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Reading VGA Memory - - @@ -41,7 +34,7 @@

    By the way, Listing 28.3 illustrates how write mode 3 can make for excellent pixel- and line-drawing code.

    -

    LISTING 28.3 L28-3.ASM

    +

    LISTING 28.3 L28-3.ASM

     ; Program that draws a diagonal line to illustrate the use of a
     ; Color Don't Care register setting of 0FFH to support fast
    @@ -146,7 +139,7 @@ WaitKeyLoop:
     Startendp
     code ends
         end  Start
    -
    +

    I hope I’ve given you a good feel for what color compare mode is and what it might be used for. Color compare mode isn’t particularly easy to understand, but it’s not that complicated in actual operation, and it’s certainly useful at times; take some time to study the sample code and perform a few experiments of your own, and you may well find useful applications for color compare mode in your graphics code.

    @@ -171,10 +164,6 @@ code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/29-01.html b/29-01.html index f12e142..b936741 100644 --- a/29-01.html +++ b/29-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - @@ -37,16 +30,16 @@


    -

    Chapter 29
    +

    Chapter 29
    Saving Screens and Other VGA Mysteries

    -

    Useful Nuggets from the VGA Zen File

    +

    Useful Nuggets from the VGA Zen File

    There are a number of VGA graphics topics that aren’t quite involved enough to warrant their own chapters, yet still cause a fair amount of programmer headscratching—and thus deserve treatment somewhere in this book. This is the place, and during the course of this chapter we’ll touch on saving and restoring 16-color EGA and VGA screens, the 16-out-of-64 colors issue, and techniques involved in reading and writing VGA control registers.

    That’s a lot of ground to cover, so let’s get started!

    -

    Saving and Restoring EGA and VGA Screens

    +

    Saving and Restoring EGA and VGA Screens

    The memory architectures of EGAs and VGAs are similar enough to treat both together in this regard. The basic principle for saving EGA and VGA 16-color graphics screens is astonishingly simple: Write each plane to disk separately. Let’s take a look at how this works in the EGA’s hi-res mode 10H, which provides 16 colors at 640x350.

    @@ -54,15 +47,14 @@

    The program shown later on in Listing 29.1 does just what I’ve described here, putting the screen into mode 10H, putting up some bittext so there is something to save, and creating the 112K file SNAPSHOT.SCR, which contains the visible portion of the mode 10H frame buffer.

    -


    - Figure 29.1
      Saving EGA/VGA display memory.

    +


    + Figure 29.1
      Saving EGA/VGA display memory.

    The only part of Listing 29.1 that’s even remotely tricky is the use of the Read Map register (Graphics Controller register 4) to make each of the four planes of display memory readable in turn. The same code is used to write 28,000 bytes of display memory to disk four times, and 28,000 bytes of memory starting at A000:0000 are written to disk each time; however, a different plane is read each time, thanks to the changing setting of the Read Map register. (If this is unclear, refer back to Figure 29.1; you may also want to reread Chapter 28 to brush up on the operation of the Read Map register in particular and reading EGA and VGA memory in general.)

    Of course, we’ll want the ability to restore what we’ve saved, and Listing 29.2 does this. Listing 29.2 reverses the action of Listing 29.1, selecting mode 10H and then loading 28,000 bytes from SNAPSHOT.SCR into each plane of display memory. The Map Mask register (Sequence Controller register 2) is used to select the plane to be written to. If your computer is slow enough, you can see the colors of the text change as each plane is loaded when Listing 29.2 runs. Note that Listing 29.2 does not itself draw any text, but rather simply loads the bit map saved by Listing 29.1 back into the mode 10H frame buffer.

    -

    LISTING 29.1 L29-1.ASM

    +

    LISTING 29.1 L29-1.ASM

     ; Program to put up a mode 10h EGA graphics screen, then save it
     ; to the file SNAPSHOT.SCR.
    @@ -192,7 +184,7 @@ Done:
     Start          endp
     Code           ends
                    end      Start
    -
    +


    @@ -211,10 +203,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/29-02.html b/29-02.html index 4755757..6b33768 100644 --- a/29-02.html +++ b/29-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - @@ -37,7 +30,7 @@


    -

    LISTING 29.2 L29-2.ASM

    +

    LISTING 29.2 L29-2.ASM

     ; Program to restore a mode 10h EGA graphics screen from
     ; the file SNAPSHOT.SCR.
    @@ -153,7 +146,7 @@ Done:
     Start         endp
     Code          ends
                   end      Start
    -
    +

    If you compare Listings 29.1 and 29.2, you will see that the Map Mask register setting used to load a given plane does not match the Read Map register setting used to read that plane. This is so because while only one plane can ever be read at a time, anywhere from zero to four planes can be written to at once; consequently, Read Map register settings are plane selections from 0 to 3, while Map Mask register settings are plane masks from 0 to 15, where a bit 0 setting of 1 enables writes to plane 0, a bit 1 setting of 1 enables writes to plane 1, and so on. Again, Chapter 28 provides a detailed explanation of the differences between the Read Map and Map Mask registers.

    @@ -176,10 +169,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/29-03.html b/29-03.html index 7cc377d..f0a2b17 100644 --- a/29-03.html +++ b/29-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - @@ -47,10 +40,10 @@

    What’s the solution? Frankly, the solution is to get VGA-specific. A TSR designed for the VGA can simply read out and save the state of the registers of interest, program those registers as needed, save the screen image, and restore the original settings. From a programmer’s perspective, readable registers are certainly near the top of the list of things to like about the VGA! The remaining installed base of EGAs is steadily dwindling, and you may be able to ignore it as a market today, as you couldn’t even a year or two ago.

    -

    If you are going to write a hi-res VGA version of the screen capture program, be sure to account for the increased size of the VGA’s mode 12H bit map. The mode 12H (640x480) screen uses 37.5K per plane of display memory, so for mode 12H the displayed screen size equate in Listings 29.1 and 29.2 should be changed to:

    +

    If you are going to write a hi-res VGA version of the screen capture program, be sure to account for the increased size of the VGA’s mode 12H bit map. The mode 12H (640x480) screen uses 37.5K per plane of display memory, so for mode 12H the displayed screen size equate in Listings 29.1 and 29.2 should be changed to:

     DISPLAYED_SCREEN_SIZEequ(640/8)*480
    -
    +

    Similarly, if you’re capturing a graphics screen that starts at an offset other than 0 in the segment at A000H, you must change the memory offset used by the disk functions to match. You can, if you so desire, read the start offset of the display memory providing the information shown on the screen from the Start Address registers (CRT Controller registers 0CH and 0DH); these registers are readable even on an EGA.

    @@ -66,7 +59,7 @@ DISPLAYED_SCREEN_SIZEequ(640/8)*480 -

    16 Colors out of 64

    +

    16 Colors out of 64

    How does one produce the 64 colors from which the 16 colors displayed by the EGA can be chosen? The answer is simple enough: There’s a BIOS function that lets you select the mapping of the 16 possible pixel values to the 64 possible colors. Let’s lay out a bit of background before proceeding, however.

    @@ -78,9 +71,8 @@ DISPLAYED_SCREEN_SIZEequ(640/8)*480

    Each of the 16 palette registers stores the mapping of one of the 16 possible 4-bit pixel values from memory to one of 64 possible 6-bit pixel values to be sent to the monitor as video data, as shown in Figure 29.2. A 4-bit pixel value of 0 causes the 6-bit value stored in palette register 0 to be sent to the display as the color of that pixel, a pixel value of 1 causes the contents of palette register 1 to be sent to the display, and so on. Since there are only four input bits, it stands to reason that only 16 colors are available at any one time; since there are six output bits, however, those 16 colors can be mapped to any of 64 colors. The mapping for each of the 16 pixel values is controlled by the lower six bits of the corresponding palette register, as shown in Figure 29.3. Secondary red, green, and blue are less-intense versions of red, green, and blue, although their exact effects vary from monitor to monitor. The best way to figure out what the 64 colors look like on your monitor is to see them, and that’s just what the program in Listing 29.3, which we’ll discuss shortly, lets you do.

    -


    - Figure 29.2
      Color translation via the palette registers.

    +


    + Figure 29.2
      Color translation via the palette registers.


    @@ -99,10 +91,6 @@ DISPLAYED_SCREEN_SIZEequ(640/8)*480
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/29-04.html b/29-04.html index 89bfdc7..366ac02 100644 --- a/29-04.html +++ b/29-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - @@ -41,9 +34,8 @@

    Video function 10H is invoked by performing an INT 10H with AH set to 10H. If AL is 0 (subfunction 0), then BL contains the number of the palette register to set, and BH contains the value to set that register to. If AL is 1 (subfunction 1), then BH contains the value to set the overscan (border) color to. Finally, if AL is 2 (subfunction 2), then ES:DX points to a 17-byte array containing the values to set palette registers 0-15 and the overscan register to. (For completeness, although it’s unrelated to the palette registers, there is one more subfunction of video function 10H. If AL = 3 (subfunction 3), bit 0 of BL is set to 1 to cause bit 7 of text attributes to select blinking, or set to 0 to cause bit 7 of text attributes to select highreverse video.)

    -


    - Figure 29.3
      Bit organization within a palette register.

    +


    + Figure 29.3
      Bit organization within a palette register.

    Listing 29.3 uses video function 10H, subfunction 2 to step through all 64 possible colors. This is accomplished by putting up 16 color bars, one for each of the 16 possible 4-bit pixel values, then changing the mapping provided by the palette registers to select a different group of 16 colors from the set of 64 each time a key is pressed. Initially, colors 0-15 are displayed, then 1-16, then 2-17, and so on up to color 3FH wrapping around to colors 0-14, and finally back to colors 0-15. (By the way, at mode set time the 16 palette registers are not set to colors 0-15, but rather to 0H, 1H, 2H, H, 4H, 5H, 14H, 7H, 38H, 39H, 3AH, 3BH, 3CH, 3DH, 3EH, and 3FH, respectively. Bits 6, 5, and 4—secondary red, green, and blue—are all set to 1 in palette registers 8-15 in order to produce high-intensity colors. Palette register 6 is set to 14H to produce brown, rather than the yellow that the expected value of 6H would produce.)

    @@ -51,7 +43,7 @@

    It’s important to understand that in Listing 29.3 the contents of display memory are never changed after initialization. The only change is the mapping from the 4-bit pixel data coming out of display memory to the 6-bit data going to the monitor. For this reason, it’s technically inaccurate to speak of bits in display memory as representing colors; more accurately, they represent attributes in the range 0-15, which are mapped to colors 0-3FH by the palette registers.

    -

    LISTING 29.3 L29-3.ASM

    +

    LISTING 29.3 L29-3.ASM

     ; Program to illustrate the color mapping capabilities of the
     ; EGA’s palette registers.
    @@ -303,7 +295,7 @@ ColorNumbersUpendp
     Start          endp
     Code           ends
                    end       Start
    -
    +


    @@ -322,10 +314,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/29-05.html b/29-05.html index 27c5a2c..b8dd051 100644 --- a/29-05.html +++ b/29-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - @@ -37,7 +30,7 @@


    -

    Overscan

    +

    Overscan

    While we’re at it, I’m going to touch on overscan. Overscan is the color of the border of the display, the rectangular area around the edge of the monitor that’s outside the region displaying active video data but inside the blanking area. The overscan (or border) color can be programmed to any of the 64 possible colors by either setting Attribute Controller register 11H directly or calling video function 10H, subfunction 1.

    @@ -49,11 +42,11 @@ -

    A Bonus Blanker

    +

    A Bonus Blanker

    An interesting bonus: The Attribute Controller provides a very convenient way to blank the screen, in the form of the aforementioned bit 5 of the Attribute Controller Index register (at address 3C0H after the Input Status 1 register—3DAH in color, 3BAH in monochrome—has been read and on every other write to 3C0H thereafter). Whenever bit 5 of the AC Index register is 0, video data is cut off, effectively blanking the screen. Setting bit 5 of the AC Index back to 1 restores video data immediately. Listing 29.4 illustrates this simple but effective form of screen blanking.

    -

    LISTING 29.4 L29-4.ASM

    +

    LISTING 29.4 L29-4.ASM

     ; Program to demonstrate screen blanking via bit 5 of the
     ; Attribute Controller Index register.
    @@ -136,7 +129,7 @@ Done:
     Start              endp
     Code               ends
                 end    Start
    -
    +


    @@ -155,10 +148,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/29-06.html b/29-06.html index 5d6736e..93f4b2c 100644 --- a/29-06.html +++ b/29-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Saving Screens and Other VGA Mysteries - - @@ -39,11 +32,11 @@

    Does that do it for color selection? Yes and no. For the EGA, we’ve covered the whole of color selection—but not so for the VGA. The VGA can emulate everything we’ve discussed, but actually performs one 4-bit to 8-bit translation (except in 256-color modes, where all 256 colors are simultaneously available), followed by yet another translation, this one 8-bit to 18-bit. What’s more, the VGA has the ability to flip instantly through as many as 16 16-color sets. The VGA’s color selection capabilities, which are supported by another set of BIOS functions, can be used to produce stunning color effects, as we’ll see when we cover them starting in Chapter 33.

    -

    Modifying VGA Registers

    +

    Modifying VGA Registers

    EGA registers are not readable. VGA registers are readable. This revelation will not come as news to most of you, but many programmers still insist on setting entire VGA registers even when they’re modifying only selected bits, as if they were programming the EGA. This comes to mind because I recently received a query inquiring why write mode 1 (in which the contents of the latches are copied directly to display memory) didn’t work in Mode X. (I’ll go into Mode X in detail later in this book.) Actually, write mode 1 does work in Mode X; it didn’t work when this particular correspondent enabled it because he did so by writing the value 01H to the Graphics Mode register. As it happens, the write mode field is only one of several fields in that register, as shown in Figure 29.4. In 256-color modes, one of the other fields—bit 6, which enables 256-color pixel formatting—is not 0, and setting it to 0 messes up the screen quite thoroughly.

    -

    The correct way to set a field within a VGA register is, of course, to read the register, mask off the desired field, insert the desired setting, and write the result back to the register. In the case of setting the VGA to write mode 1, do this:

    +

    The correct way to set a field within a VGA register is, of course, to read the register, mask off the desired field, insert the desired setting, and write the result back to the register. In the case of setting the VGA to write mode 1, do this:

     mov   dx,3ceh          ;Graphics controller index
     mov   al,5             ;Graphics mode reg index
    @@ -53,15 +46,14 @@ in    al,dx            ;get current mode setting
     and   al,not 3         ;mask off write mode field
     or    al,1             ;set write mode field to 1
     out   dx,al            ;set write mode 1
    -
    +

    This approach is more of a nuisance than simply setting the whole register, but it’s safer. It’s also slower; for cases where you must set a field repeatedly, it might be worthwhile to read and mask the register once at the start, and save it in a variable, so that the value is readily available in memory and need not be repeatedly read from the port. This approach is especially attractive because INs are much slower than memory accesses on 386 and 486 machines.

    Astute readers may wonder why I didn’t put a delay sequence, such as JMP $+2, between the IN and OUT involving the same register. There are, after all, guidelines from IBM, specifying that a certain period should be allowed to elapse before a second access to an I/O port is attempted, because not all devices can respond as rapidly as a 286 or faster CPU can access a port. My answer is that while I can’t guarantee that a delay isn’t needed, I’ve never found a VGA that required one; I suspect that the delay specification has more to do with motherboard chips such as the timer, the interrupt controller, and the like, and I sure hate to waste the delay time if it’s not necessary. However, I’ve never been able to find anyone with the definitive word on whether delays might ever be needed when accessing VGAs, so if you know the gospel truth, or if you know of a VGA/processor combo that does require delays, please let me know by contacting me through the publisher. You’d be doing a favor for a whole generation of graphics programmers who aren’t sure whether they’re skating on thin ice without those legendary delays.

    -


    - Figure 29.4
      Graphics mode register fields.

    +


    + Figure 29.4
      Graphics mode register fields.


    @@ -80,10 +72,6 @@ out dx,al ;set write mode 1
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-01.html b/30-01.html index 3fc015f..57f5dcb 100644 --- a/30-01.html +++ b/30-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -37,10 +30,10 @@


    -

    Chapter 30
    +

    Chapter 30
    Video Est Omnis Divisa

    -

    The Joys and Galling Problems of Using Split Screens on the EGA and VGA

    +

    The Joys and Galling Problems of Using Split Screens on the EGA and VGA

    The ability to split the screen into two largely independent portions one—displayed above the other on the screen—is one of the more intriguing capabilities of the VGA and EGA. The split screen feature can be used for popups (including popups that slide smoothly onto the screen), or simply to display two separate portions of display memory on a single screen. While it’s possible to accomplish the same effects purely in software without using the split screen, software solutions tend to be slow and hard to implement.

    @@ -48,7 +41,7 @@

    Let’s start with the basic operation of the split screen.

    -

    How the Split Screen Works

    +

    How the Split Screen Works

    The operation of the split screen is simplicity itself. A split screen start scan line value is programmed into two EGA registers or three VGA registers. (More on exactly which registers in a moment.) At the beginning of each frame, the video circuitry begins to scan display memory for video data starting at the address specified by the start address registers, just as it normally would. When the video circuitry encounters the specified split screen start scan line in the course of scanning video data onto the screen, it completes that scan line normally, then resets the internal pointer which addresses the next byte of display memory to be read for video data to zero. Display memory from address zero onward is then scanned for video data in the usual way, progressing toward the high end of memory. At the end of the frame, the pointer to the next byte of display memory to scan is reloaded from the start address registers, and the whole process starts over.

    @@ -56,9 +49,8 @@

    If both the start address and the split screen start scan line are set to 0, the data at offset zero in display memory is displayed as both the first scan line on the screen and the second scan line. There is no way to make the split screen cover the entire screen—it always comes up at least one scan line short.

    -


    - Figure 30.1
      Display memory and the split screen.

    +


    + Figure 30.1
      Display memory and the split screen.

    So, where is the split screen start scan line stored? The answer varies a bit, depending on whether you’re talking about the EGA or the VGA. On the EGA, the split screen start scan line is a 9-bit value, with bits 7-0 stored in the Line Compare register (CRTC register 18H) and bit 8 stored in bit 4 of the Overflow register (CRTC register 7). Other bits in the Overflow register serve as the high bits of other values, such as the vertical total and the vertical blanking start. Since EGA registers are—alas!—not readable, you must know the correct settings for the other bits in the Overflow registers to use the split screen on an EGA. Fortunately, there are only two standard Overflow register settings on the EGA: 11H for 200-scan-line modes and 1FH for 350-scan-line modes.

    @@ -66,7 +58,7 @@

    Turning the split screen on involves nothing more than setting all bits of the split screen start scan line to the scan line after which you want the split screen to start appearing. (Of course, you’ll probably want to change the start address before using the split screen; otherwise, you’ll just end up displaying the memory at offset zero twice: once in the normal screen and once in the split screen.) Turning off the split screen is a simple matter of setting the split screen start scan line to a value equal to or greater than the last scan line displayed; the safest such approach is to set all bits of the split screen start scan line to 1. (That is, in fact, the split screen start scan line value programmed by the BIOS during a mode set.)

    -

    The Split Screen in Action

    +

    The Split Screen in Action

    All of these points are illustrated by Listing 30.1. Listing 30.1 fills display memory starting at offset zero (the split screen area of memory) with text identifying the split screen, fills display memory starting at offset 8000H with a graphics pattern, and sets the start address to 8000H. At this point, the normal screen is being displayed (the split screen start scan line is still set to the BIOS default setting, with all bits equal to 1, so the split screen is off), with the pixels based on the contents of display memory at offset 8000H. The contents of display memory between offset 0 and offset 7FFFH are not visible at all.

    @@ -91,10 +83,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-02.html b/30-02.html index 0335958..a94d1ac 100644 --- a/30-02.html +++ b/30-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -39,7 +32,7 @@

    Finally, after another keypress, Listing 30.1 halts.

    -

    LISTING 30.1 L30-1.ASM

    +

    LISTING 30.1 L30-1.ASM

     ; Demonstrates the VGA/EGA split screen in action.
     ;
    @@ -413,7 +406,7 @@ SplitScreenDown    endp
     ;*********************************************************************
     Code    ends
             end    Start
    -
    +


    @@ -432,10 +425,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-03.html b/30-03.html index f7098db..49208c4 100644 --- a/30-03.html +++ b/30-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -37,7 +30,7 @@


    -

    VGA and EGA Split-Screen Operation Don’t Mix

    +

    VGA and EGA Split-Screen Operation Don’t Mix

    You must set the IS_VGA equate at the start of Listing 30.1 correctly for the adapter the code will run on in order for the program to perform properly. This equate determines how the upper bits of the split screen start scan line are set by SetSplitScreenRow. If IS_VGA is 0 (specifying an EGA target), then bit 8 of the split screen start scan line is set by programming the entire Overflow register to 1FH; this is hard-wired for the 350-scan-line modes of the EGA. If IS_VGA is 1 (specifying a VGA target), then bits 8 and 9 of the split screen start scan line are set by reading the registers they reside in, changing only the split-screen-related bits, and writing the modified settings back to their respective registers.

    @@ -45,7 +38,7 @@

    By the way, Listing 30.1 operates in mode 10H because that’s the highest-resolution mode the VGA and EGA share. That’s not the only mode the split screen works in, however. In fact, it works in all modes, as we’ll see later.

    -

    Setting the Split-Screen-Related Registers

    +

    Setting the Split-Screen-Related Registers

    Setting the split-screen-related registers is not as simple a matter as merely outputting the right values to the right registers; timing is also important. The split screen start scan line value is checked against the number of each scan line as that scan line is displayed, which means that the split screen start scan line potentially takes effect the moment it is set. In other words, if the screen is displaying scan line 15 and you set the split screen start to 16, that change will be picked up immediately and the split screen will start after the next scan line. This is markedly different from changes to the start address, which take effect only at the start of the next frame.

    @@ -63,7 +56,7 @@

    One interesting effect of setting the split screen registers at the start of vertical sync is that it has the effect of synchronizing the program to the display adapter’s frame rate. No matter how fast the computer running Listing 30.1 may be, the split screen will move at a maximum rate of once per frame. This is handy for regulating execution speed over a wide variety of hardware performance ranges; however, be aware that the VGA supports 70 Hz frame rates in all non-480-scan-line modes, while the VGA in 480-scan-line-modes and the EGA in all color modes support 60 Hz frame rates.

    -

    The Problem with the EGA Split Screen

    +

    The Problem with the EGA Split Screen

    I mentioned earlier that the EGA’s split screen is a little buggy. How? you may well ask, particularly given that Listing 30.1 illustrates that the EGA split screen seems pretty functional.

    @@ -98,10 +91,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-04.html b/30-04.html index 0e7bd98..86162cb 100644 --- a/30-04.html +++ b/30-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -37,7 +30,7 @@


    -

    Split Screen and Panning

    +

    Split Screen and Panning

    Back in Chapter 23, I presented a program that performed smooth horizontal panning. Smooth horizontal panning consists of two parts: byte-by-byte (8-pixel) panning by changing the start address and pixel-by-pixel intrabyte panning by setting the Pel Panning register (AC register 13H) to adjust alignment by 0 to 7 pixels. (IBM prefers its own jargon and uses the word “pel” instead of “pixel” in much of their documentation, hence “pel panning.” Then there’s DASD, a.k.a. Direct Access Storage Device—IBM-speak for hard disk.)

    @@ -55,7 +48,7 @@

    On the VGA, there is recourse. A VGA-only bit, bit 5 of the AC Mode Control register (AC register 10H), turns off pel panning in the split screen. In other words, when this bit is set to 1, pel panning is reset to zero before the first line of the split screen, and remains zero until the end of the frame. This doesn’t allow you to pan the split screen horizontally, mind you—there’s no way to do that—but it does let you pan the normal screen while the split screen stays rock-solid. This can be used to produce an attractive “streaming tape” effect in the normal screen while the split screen is used to display non-moving information.

    -

    The Split Screen and Horizontal Panning: An Example

    +

    The Split Screen and Horizontal Panning: An Example

    Listing 30.2 illustrates the interaction of horizontal smooth panning with the split screen, as well as the suppression of pel panning in the split screen. Listing 30.2 creates a virtual screen 1024 pixels across by setting the Offset register (CRTC register 13H) to 64, sets the normal screen to scan video data beginning far enough up in display memory to leave room for the split screen starting at offset zero, turns on the split screen, and fills in the normal screen and split screen with distinctive patterns. Next, Listing 30.2 pans the normal screen horizontally without setting bit 5 of the AC Mode Control register to 1. As you’d expect, the split screen jerks about quite horribly. After a key press, Listing 30.2 sets bit 5 of the Mode Control register and pans the normal screen again. This time, the split screen doesn’t budge an inch—if the code is running on a VGA.

    @@ -78,10 +71,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-05.html b/30-05.html index c80518f..01672c1 100644 --- a/30-05.html +++ b/30-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -37,7 +30,7 @@


    -

    LISTING 30.2 L30-2.ASM

    +

    LISTING 30.2 L30-2.ASM

     ; Demonstrates the interaction of the split screen and
     ; horizontal pel panning. On a VGA, first pans right in the top
    @@ -457,7 +450,7 @@ PanRight     endp
     ;*********************************************************************
     Codeends
     endStart
    -
    +


    @@ -476,10 +469,6 @@ endStart
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-06.html b/30-06.html index 9a967e3..0886161 100644 --- a/30-06.html +++ b/30-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -37,7 +30,7 @@


    -

    Notes on Setting and Reading Registers

    +

    Notes on Setting and Reading Registers

    There are a few interesting points regarding setting and reading registers to be made about Listing 30.2. First, bit 5 of the AC Index register should be set to 1 whenever palette RAM is not being set (which is to say, all the time in your code, because palette RAM should normally be set via the BIOS). When bit 5 is 0, video data from display memory is no longer sent to palette RAM, and the screen becomes a solid color—not normally a desirable state of affairs.

    @@ -80,10 +73,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/30-07.html b/30-07.html index 3c37966..dd4d51c 100644 --- a/30-07.html +++ b/30-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Video Est Omnis Divisa - - @@ -37,7 +30,7 @@


    -

    Split Screens in Other Modes

    +

    Split Screens in Other Modes

    So far we’ve only discussed the split screen in mode 10H. What about other modes? Generally, the split screen works in any mode; the basic rule is that when a scan line on the screen matches the split screen scan line, the internal display memory pointer is reset to zero. I’ve found this to be true even in oddball modes, such as line-doubled CGA modes and the 320x200 256-color mode (which is really a 320x400 mode with each line repeated. For split-screen purposes, the VGA and EGA seem to count purely in scan lines, not in rows or doubled scan lines or the like. However, I have run into small anomalies in those modes on clones, and I haven’t tested all modes (nor, lord knows, all clones!) so be careful when using the split screen in modes other than modes 0DH-12H, and test your code on a variety of hardware.

    @@ -45,7 +38,7 @@

    What of the split screen in text mode? It works fine; in fact, it not only resets the internal memory pointer to zero, but also resets the text scan line counter—which marks which line within the font you’re on—to zero, so the split screen starts out with a full row of text. There’s only one trick with text mode: When split screen pel panning suppression is on, the pel panning setting is forced to 0 for the rest of the frame. Unfortunately, 0 is not the “no-panning” setting for 9-dot-wide text; 8 is. The result is that when you turn on split screen pel panning suppression, the text in the split screen won’t pan with the normal screen, as intended, but will also display the undesirable characteristic of moving one pixel to the left. Whether this causes any noticeable on-screen effects depends on the text displayed by a particular application; for example, there should be no problem if the split screen has a border of blanks on the left side.

    -

    How Safe?

    +

    How Safe?

    So, how safe is it to use the split screen? My opinion is that it’s perfectly safe, although I’d welcome input from people with extensive split screen experience—and the effects are striking enough that the split screen is well worth using in certain applications.

    @@ -70,10 +63,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/31-01.html b/31-01.html index 470b83e..e6a65ec 100644 --- a/31-01.html +++ b/31-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - @@ -37,10 +30,10 @@


    -

    Chapter 31
    +

    Chapter 31
    Higher 256-Color Resolution on the VGA

    -

    When Is 320x200 Really 320x400?

    +

    When Is 320x200 Really 320x400?

    One of the more appealing features of the VGA is its ability to display 256 simultaneous colors. Unfortunately, one of the less appealing features of the VGA is the limited resolution (320x200) of the one 256-color mode the IBM-standard BIOS supports. (There are, of course, higher resolution 256-color modes in the legion of SuperVGAs, but they are by no means a standard, and differences between seemingly identical modes from different manufacturers can be vexing.) More colors can often compensate for less resolution, but the resolution difference between the 640x480 16-color mode and the 320x200 256-color mode is so great that many programmers must regretfully decide that they simply can’t afford to use the 256-color mode.

    @@ -50,7 +43,7 @@

    So. Let’s get started.

    -

    Why 320x200? Only IBM Knows for Sure

    +

    Why 320x200? Only IBM Knows for Sure

    The first question, of course, is, “How can it be possible to get higher 256-color resolutions out of the VGA?” After all, there were no unused higher resolutions to be found in the CGA, Hercules card, or EGA.

    @@ -60,7 +53,7 @@

    On the other hand, the smaller display memory size of the MCGA also limits the number of colors supported in 640x480 mode to 2, rather than the 16 supported by the VGA. In this case, though, IBM simply created two modes and made both available on the VGA: mode 11H for 640x480 2-color graphics and mode 12H for 640x480 16-color graphics. The same could have been done for 256-color graphics—but wasn’t. Why? I don’t know. Maybe IBM just didn’t like the odd aspect ratio of a 320x400 graphics mode. Maybe they didn’t want to have to worry about how to map in more than 64K of display memory. Heck, maybe they made a mistake in designing the chip. Whatever the reason, mode 13H is really a 400-scan-line mode masquerading as a 200-scan-line mode, and we can readily end that masquerade.

    -

    320x400 256-Color Mode

    +

    320x400 256-Color Mode

    Okay, what’s so great about 320x400 256-color mode? Two things: easy, safe mode sets and page flipping.

    @@ -70,7 +63,7 @@

    That’s why I like 320x400 256-color mode. The next step is to understand how display memory is organized in 320x400 mode, and that’s not so simple.

    -

    Display Memory Organization in 320x400 Mode

    +

    Display Memory Organization in 320x400 Mode

    First, let’s look at why display memory must be organized differently in 320x400 256-color mode than in mode 13H. The designers of the VGA intentionally limited the maximum size of the bitmap in mode 13H to 64K, thereby limiting resolution to 320x200. This was accomplished in hardware, so there is no way to extend the bitmap organization of mode 13H to 320x400 mode.

    @@ -91,10 +84,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/31-02.html b/31-02.html index 4094fe8..b9808ca 100644 --- a/31-02.html +++ b/31-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - @@ -43,9 +36,8 @@

    Let’s look at this another way. Ideally, we’d like one long bitmap, with each pixel at the address that’s just after the address of the pixel to the left. Well, that’s true in this case too, if you consider the number of the plane that the pixel is in to be part of the pixel’s address. View the pixel numbers on the screen as increasing from left to right and from the end of one scan line to the start of the next. Then the pixel number, n, of the pixel at display memory address address in plane plane is:

    -


    - Figure 31.1
      Bitmap organization in 320x400 256-color mode in 320x400 256-color mode.

    +


    + Figure 31.1
      Bitmap organization in 320x400 256-color mode in 320x400 256-color mode.

    n = (address * 4) + plane

    @@ -63,7 +55,7 @@

    Our next task is to convert standard mode 13H into 320x400 mode. That’s accomplished by undoing some of the mode bits that are set up especially for mode 13H, so that from a programming perspective the VGA reverts to a straightforward planar model of memory. That means taking the VGA out of chain 4 mode and doubleword mode, turning off the double display of each scan line, making sure chain mode, odd/even mode, and word mode are turned off, and selecting byte mode for video data display. All that’s done in the Set320x400Mode subroutine in Listing 31.1, which we’ll discuss next.

    -

    Reading and Writing Pixels

    +

    Reading and Writing Pixels

    The basic graphics functions in any mode are functions to read and write single pixels. Any more complex function can be built on these primitives, although that’s rarely the speediest solution. What’s more, once you understand the operation of the read and write pixel functions, you’ve got all the knowledge you need to create functions that perform more complex graphics functions. Consequently, we’ll start our exploration of 320x400 mode with pixel-at-a-time line drawing.

    @@ -86,10 +78,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/31-03.html b/31-03.html index 2a1eeac..9dd0470 100644 --- a/31-03.html +++ b/31-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - @@ -37,7 +30,7 @@


    -

    LISTING 31.1 L31-1.ASM

    +

    LISTING 31.1 L31-1.ASM

     ; Program to demonstrate pixel drawing in 320x400 256-color
     ; mode on the VGA. Draws 8 lines to form an octagon, a pixel
    @@ -366,7 +359,7 @@ GetNextKey   endp
     Code   ends
     ;
     end    Start
    -
    +


    @@ -385,10 +378,6 @@ end Start
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/31-04.html b/31-04.html index 6a3a6d9..c33020c 100644 --- a/31-04.html +++ b/31-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - @@ -67,7 +60,7 @@

    It’s all a bit complicated, but as I say, you should be able to design an adequately fast—and often very fast—version for 320x400 mode of whatever graphics function you need. If you’re not all that concerned with speed, WritePixel and ReadPixel should meet your needs.

    -

    Two 256-Color Pages

    +

    Two 256-Color Pages

    Listing 31.2 demonstrates the two pages of 320x400 256-color mode by drawing slanting color bars in page 0, then drawing color bars slanting the other way in page 1 and flipping to page 1 on the next key press. (Note that page 1 is accessed starting at offset 8000H in display memory, and is—unsurprisingly—displayed by setting the start address to 8000H.) Finally, Listing 31.2 draws vertical color bars in page 0 and flips back to page 0 when another key is pressed.

    @@ -90,10 +83,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/31-05.html b/31-05.html index e7b6c17..9fdc4f0 100644 --- a/31-05.html +++ b/31-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Higher 256-Color Resolution on the VGA - - @@ -37,7 +30,7 @@


    -

    LISTING 31.2 L31-2.ASM

    +

    LISTING 31.2 L31-2.ASM

     ; Program to demonstrate the two pages available in 320x400
     ; 256-color modes on a VGA.  Draws diagonal color bars in all
    @@ -298,11 +291,11 @@ GetNextKey      endp
     Codeends
     ;
     endStart
    -
    +

    When you run Listing 31.2, note the extremely smooth edges and fine gradations of color, especially in the screens with slanting color bars. The displays produced by Listing 31.2 make it clear that 320x400 256-color mode can produce effects that are simply not possible in any 16-color mode.

    -

    Something to Think About

    +

    Something to Think About

    You can, if you wish, use the display memory organization of 320x400 mode in 320x200 mode by modifying Set320x400Mode to leave the maximum scan line setting at 1 in the mode set. (The version of Set320x400Mode in Listings 31.1 and 31.2 forces the maximum scan line to 0, doubling the effective resolution of the screen.) Why would you want to do that? For one thing, you could then choose from not two but four 320x200 256-color display pages, starting at offsets 0, 4000H, 8000H, and 0C000H in display memory. For another, having only half as many pixels per screen can as much as double drawing speeds; that’s one reason that many games run at 320x200, and even then often limit the active display drawing area to only a portion of the screen.

    @@ -323,10 +316,6 @@ endStart
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/32-01.html b/32-01.html index 6759abe..6d081d8 100644 --- a/32-01.html +++ b/32-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - @@ -37,10 +30,10 @@


    -

    Chapter 32
    +

    Chapter 32
    Be It Resolved: 360x480

    -

    Taking 256-Color Modes About as Far as the Standard VGA Can Take Them

    +

    Taking 256-Color Modes About as Far as the Standard VGA Can Take Them

    In the last chapter, we learned how to coax 320x400 256-color resolution out of a standard VGA. At the time, I noted that the VGA was actually capable of supporting 256-color resolutions as high as 360x480, but didn’t pursue the topic further, preferring to concentrate on the versatile and easy-to-set 320x400 256-color mode instead.

    @@ -48,7 +41,7 @@

    In this chapter, I’m going to combine John’s mode set code with appropriately modified versions of the dot-plot code from Chapter 31 and the line-drawing code that we’ll develop in Chapter 35. Together, those routines will make a pretty nifty demo of the capabilities of 360x480 256-color mode.

    -

    Extended 256-Color Modes: What’s Not to Like?

    +

    Extended 256-Color Modes: What’s Not to Like?

    When last we left 256-color programming, we had found that the standard 256-color mode, mode 13H, which officially offers 320x200 resolution, actually displays 400, not 200, scan lines, with line-doubling used to reduce the effective resolution to 320x200. By tweaking a few of the VGA’s mode registers, we converted mode 13H to a true 320x400 256-color mode. As an added bonus, that 320x400 mode supports two graphics pages, a distinct improvement over the single graphics page supported by mode 13H. (We also learned how to get four graphics pages at 320x200 resolution, should that be needed.)

    @@ -64,7 +57,7 @@

    In other words, 360x480 256-color mode is worth considering—so let’s have a look.

    -

    360x480 256-Color Mode

    +

    360x480 256-Color Mode

    I’m going to start by showing you 360x480 256-color mode in action, after which we’ll look at how it works. I suspect that once you see what this mode looks like, you’ll be more than eager to learn how to use it.

    @@ -89,10 +82,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/32-02.html b/32-02.html index dbab7d3..30ac73b 100644 --- a/32-02.html +++ b/32-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - @@ -37,7 +30,7 @@


    -

    LISTING 32.1 L32-1.ASM

    +

    LISTING 32.1 L32-1.ASM

     ; Borland C/C++ tiny/small/medium model-callable assembler
     ; subroutines to:
    @@ -250,7 +243,7 @@ _Read360x480Dotprocnear
     _Read360x480Dot  endp
     _TEX   Tends
            end
    -
    +


    @@ -269,10 +262,6 @@ _TEX Tends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/32-03.html b/32-03.html index dbb7fd4..b961507 100644 --- a/32-03.html +++ b/32-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - @@ -37,7 +30,7 @@


    -

    LISTING 32.2 L32-2.C

    +

    LISTING 32.2 L32-2.C

      * Sample program to illustrate VGA line drawing in 360x480
      * 256-color mode.
    @@ -239,7 +232,7 @@ void main()
        _AX = TEXT_MODE;
        geninterrupt(BIOS_VIDEO_INT);
     }
    -
    +

    The first thing you’ll notice when you run this code is that the speed of 360x480 256-color mode is pretty good, especially considering that most of the program is im-plemented in C.

    @@ -268,10 +261,6 @@ void main()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/32-04.html b/32-04.html index 2ab4bf5..4332f2d 100644 --- a/32-04.html +++ b/32-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - @@ -43,19 +36,19 @@

    Now that we’ve seen the wonders of which our new mode is capable, let’s take the time to understand how it works.

    -

    How 360x480 256-Color Mode Works

    +

    How 360x480 256-Color Mode Works

    In describing 360x480 256-color mode, I’m going to assume that you’re familiar with the discussion of 320x400 256-color mode in the last chapter. If not, go back to that chapter and read it; the two modes have a great deal in common, and I’m not going to bore you by repeating myself when the goods are just a few page flips (the paper kind) away.

    360x480 256-color mode is essentially 320x400 256-color mode, but stretched in both dimensions. Let’s look at the vertical stretching first, since that’s the simpler of the two.

    -

    480 Scan Lines per Screen: A Little Slower, But No Big Deal

    +

    480 Scan Lines per Screen: A Little Slower, But No Big Deal

    There’s nothing unusual about 480 scan lines; standard modes 11H and 12H support that vertical resolution. The number of scan lines has nothing to do with either the number of colors or the horizontal resolution, so converting 320x400 256mode to 320x480 256-color mode is a simple matter of reprogramming the VGA’s vertical control registers—which control the scan lines displayed, the vertical sync pulse, vertical blanking, and the total number of scan lines—to the 480-scansettings, and setting the polarities of the horizontal and vertical sync pulses to tell the monitor to adjust to a 480-line screen.

    Switching to 480 scan lines has the effect of slowing the screen refresh rate. The VGA always displays at 70 Hz except in 480-scan-line modes; there, due to the time required to scan the extra lines, the refresh rate slows to 60 Hz. (VGA monitors always scan at the same rate horizontally; that is, the distance across the screen covered by the electron beam in a given period of time is the same in all modes. Consequently, adding extra lines per frame requires extra time.) 60 Hz isn’t bad—that’s the only refresh rate the EGA ever supported, and the EGA was the industry standard in its time—but it does tend to flicker a little more and so is a little harder on the eyes than 70 Hz.

    -

    360 Pixels per Scan Line: No Mean Feat

    +

    360 Pixels per Scan Line: No Mean Feat

    Converting from 320 to 360 pixels per scan line is more difficult than converting from 400 to 480 scan lines per screen. None of the VGA’s graphics modes supports 360 pixels across the screen, or anything like it; the standard choices are 320 and 640 pixels across. However, the VGA does support the horizontal resolution we seek—360 pixels—in 40-column text mode.

    @@ -78,10 +71,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/32-05.html b/32-05.html index 92fdc2f..7c95d3d 100644 --- a/32-05.html +++ b/32-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Be It Resolved: 360x480 - - @@ -47,7 +40,7 @@

    Once all that’s done, the VGA is in 360x480 mode, awaiting our every high-resolution 256-color graphics whim.

    -

    Accessing Display Memory in 360x480 256-Color Mode

    +

    Accessing Display Memory in 360x480 256-Color Mode

    Setting up for 360x480 256-color mode proved to be quite a task. Is drawing in this mode going to be as difficult?

    @@ -67,9 +60,8 @@

    That’s really all there is to drawing in 360x480 256-color mode. From a programming perspective, this mode is no more complicated than 320x400 256-color mode once the mode set is completed, and should be capable of good performance given some clever coding. It’s not particular straightforward to implement bitblt, block move, or fast line-drawing code for any of the extended 256-color modes, but it can be done—and it’s worth the trouble. Even the small taste we’ve gotten of the capabilities of these modes shows that they put the traditional CGA, EGA, and generally even VGA modes to shame.

    -


    - Figure 32.1
      Pixel organization in 360x480 256-color mode.

    +


    + Figure 32.1
      Pixel organization in 360x480 256-color mode.

    There’s more and better to come, though; in later chapters, we’ll return to high-resolution 256-color programming in a big way, by exploring the tremendous potential of these modes for real time 2-D and 3-D animation.

    @@ -90,10 +82,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/33-01.html b/33-01.html index 3007def..d3b7f02 100644 --- a/33-01.html +++ b/33-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - @@ -37,10 +30,10 @@


    -

    Chapter 33
    +

    Chapter 33
    Yogi Bear and Eurythmics Confront VGA Colors

    -

    The Basics of VGA Color Generation

    +

    The Basics of VGA Color Generation

    Kevin Mangis wants to know about the VGA’s 4-bit to 8-bit to 18-bit color translation. Mansur Loloyan would like to find out how to generate a look-up table containing 256 colors and how to change the default color palette. And surely they are only the tip of the iceberg; hordes of screaming programmers from every corner of the planet are no doubt tearing the place up looking for a discussion of VGA color, and venting their frustration at my mailbox. Let’s have it, they’ve said, clearly and in considerable numbers. As Eurythmics might say, who is this humble writer to disagree?

    @@ -52,23 +45,22 @@

    I would be remiss if I didn’t point you in the direction of two more articles, these in the July 1990 issue of Dr. Dobb’s Journal. “Super VGA Programming,” by Chris Howard, provides a good deal of useful information about SuperVGA chipsets, modes, and programming. “Circles and the Digital Differential Analyzer,” by Tim Paterson, is a good article about fast circle drawing, a topic we’ll tackle soon. All in all, the dog days of 1990 were good times for graphics.

    -

    VGA Color Basics

    +

    VGA Color Basics

    Briefly put, the VGA color translation circuitry takes in one 4- or 8-bit pixel value at a time and translates it into three 6-bit values, one each of red, green, and blue, that are converted to corresponding analog levels and sent to the monitor. Seems simple enough, doesn’t it? Unfortunately, nothing is ever that simple on the VGA, and color translation is no exception.

    -

    The Palette RAM

    +

    The Palette RAM

    The color path in the VGA involves two stages, as shown in Figure 33.1. The first stage fetches a 4-bit pixel from display memory and feeds it into the EGA-compatible palette RAM (so called because it is functionally equivalent to the palette RAM color translation circuitry of the EGA), which translates it into a 6-bit value and sends it on to the DAC. The translation involves nothing more complex than the 4-bit value of a pixel being used as the address of one of the 16 palette RAM registers; a pixel value of 0 selects the contents of palette RAM register 0, a pixel value of 1 selects register 1, and so on. Each palette RAM register stores 6 bits, so each time a palette RAM register is selected by an incoming 4-bit pixel value, 6 bits of information are sent out by the palette RAM. (The operation of the palette RAM was described back in Chapter 29.)

    The process is much the same in text mode, except that in text mode each 4-bit pixel value is generated based on the character’s font pattern and attribute. In 256-color mode, which we’ll get to eventually, the palette RAM is not a factor from the programmer’s perspective and should be left alone.

    -

    The DAC

    +

    The DAC

    Once the EGA-compatible palette RAM has fulfilled its karma and performed 4-bit to 6-bit translation on a pixel, the resulting value is sent to the DAC (Digital/Analog Converter). The DAC performs an 8-bit to 18-bit conversion in much the same manner as the palette RAM, converts the 18-bit result to analog red, green, and blue signals (6 bits for each signal), and sends the three analog signals to the monitor. The DAC is a separate chip, external to the VGA chip, but it’s an integral part of the VGA standard and is present on every VGA.

    -


    - Figure 33.1
      The VGA color generation path.

    +


    + Figure 33.1
      The VGA color generation path.

    (I’d like to take a moment to point out that you can’t speak of “color” at any point in the color translation process until the output stage of the DAC. The 4-bit pixel values in memory, 6-bit values in the palette RAM, and 8-bit values sent to the DAC are all attributes, not colors, because they’re subject to translation by a later stage. For example, a pixel with a 4-bit value of 0 isn’t black, it’s attribute 0. It will be translated to 3FH if palette RAM register 0 is set to 3FH, but that’s not the color white, just another attribute. The value 3FH coming into the DAC isn’t white either, and if the value stored in DAC register 63 is red=7, green=0, and blue=0, the actual color displayed for that pixel that was 0 in display memory will be dim red. It isn’t color until the DAC says it’s color.)

    @@ -91,10 +83,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/33-02.html b/33-02.html index 5c1be2d..47f236f 100644 --- a/33-02.html +++ b/33-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - @@ -37,7 +30,7 @@


    -

    Color Paging with the Color Select Register

    +

    Color Paging with the Color Select Register

    “Wait a minute,” you say bemusedly. “Aren’t you missing some bits between the palette RAM and the DAC?” Indeed I am. The palette RAM puts out 6 bits at a time, and the DAC takes in 8 bits at a time. The two missing bits—bits 6 and 7 going into the DAC—are supplied by bits 2 and 3 of the Color Select register (Attribute Controller register 14H). This has intriguing implications. In 16-color modes, pixel data can select only one of 16 attributes, which the EGA palette RAM translates into one of 64 attributes. Normally, those 64 attributes look up colors from registers 0 through 63 in the DAC, because bits 2 and 3 of the Color Select register are both zero. By changing the Color Select register, however, one of three other 64 color sets can be selected instantly. I’ll refer to the process of flipping through color sets in this manner as color paging.

    @@ -45,7 +38,7 @@

    Why is it a good idea to set the palette RAM to a pass-through state? It’s a good idea because the palette RAM is programmed by the BIOS to EGA-compatible settings and the first 64 DAC registers are programmed to emulate the 64 colors that an EGA can display during mode sets for 16-color modes. This is done for compatibility with EGA programs, and it’s useless if you’re going to tinker with the VGA’s colors. As a VGA programmer, you want to take a 4-bit pixel value and turn it into an 18-bit RGB value; you can do that without any help from the palette RAM, and setting the palette RAM to pass-through values effectively takes it out of the circuit and simplifies life something wonderful. The palette RAM exists solely for EGA compatibility, and serves no useful purpose that I know of for VGA-only color programming.

    -

    256-Color Mode

    +

    256-Color Mode

    So far I’ve spoken only of 16-color modes; what of 256-color modes?

    @@ -53,7 +46,7 @@

    On the other hand, feel free to alter the DAC settings to your heart’s content in 256-color mode, all the more so because this is the only mode in which all 256 DAC settings can be displayed simultaneously. By the way, the Color Select register and bit 7 of the Attribute Controller Mode register are ignored in 256-color mode; all 8 bits sent from the VGA chip to the DAC come from display memory. Therefore, there is no color paging in 256-color mode. Of course, that makes sense given that all 256 DAC registers are simultaneously in use in 256-color mode.

    -

    Setting the Palette RAM

    +

    Setting the Palette RAM

    The palette RAM can be programmed either directly or through BIOS interrupt 10H, function 10H. I strongly recommend using the BIOS interrupt; a clone BIOS may mask incompatibilities with genuine IBM silicon. Such incompatibilities could include anything from flicker to trashing the palette RAM; or they may not exist at all, but why find out the hard way? My policy is to use the BIOS unless there’s a clear reason not to do so, and there’s no such reason that I know of in this case.

    @@ -80,10 +73,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/33-03.html b/33-03.html index dfeb426..790c21c 100644 --- a/33-03.html +++ b/33-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - @@ -37,7 +30,7 @@


    -

    Setting the DAC

    +

    Setting the DAC

    Like the palette RAM, the DAC registers can be set either directly or through the BIOS. Again, the BIOS should be used whenever possible, but there are a few complications here. My experience is that varying degrees of flicker and screen bounce occur on many VGAs when a large block of DAC registers is set through the BIOS. That’s not a problem when the DAC is loaded just once and then left that way, as is the case in Listing 33.1, which we’ll get to shortly, but it can be a serious problem when the color set is changed rapidly (“cycled”) to produce on-screen effects such as rippling colors. My (limited) experience is that it’s necessary to program the DAC directly in order to cycle colors cleanly, although input from readers who have worked extensively with VGA color is welcome.

    @@ -47,7 +40,7 @@

    A block of sequential DAC registers ranging in size from one register up to all 256 can be set via subfunction 12H (AL=12H) of interrupt 10H, function 10H (AH=10H). In this case, BX contains the number of the first register to set, CX contains the number of registers to set, and ES:DX contains the address of a table of color entries to which DAC registers BX through BX+CX-1 are to be set. The color entry for each DAC register consists of three bytes; the first byte is a 6-bit red component, the second byte is a 6-bit green component, and the third byte is a 6-bit blue component, as illustrated by Listing 33.1.

    -

    If You Can’t Call the BIOS, Who Ya Gonna Call?

    +

    If You Can’t Call the BIOS, Who Ya Gonna Call?

    Although the palette RAM and DAC registers should be set through the BIOS whenever possible, there are times when the BIOS is not the best choice or even a choice at all; for example, a protected-mode program may not have access to the BIOS. Also, as mentioned earlier, it may be necessary to program the DAC directly when performing color cycling. Therefore, I’ll briefly describe how to set the palette RAM and DAC registers directly; in Chapter A on the companion CD-ROM I’ll discuss programming the DAC directly in more detail.

    @@ -63,7 +56,7 @@

    In the meantime, if you can use the BIOS to set the DAC, do so; then you won’t have to worry about the real and potential complications of setting the DAC directly.

    -

    An Example of Setting the DAC

    +

    An Example of Setting the DAC

    This chapter has gotten about as big as a chapter really ought to be; the VGA color saga will continue in the next few. Quickly, then, Listing 33.1 is a simple example of setting the DAC that gives you a taste of the spectacular effects that color translation makes possible. There’s nothing particularly complex about Listing 33.1; it just selects 256-color mode, fills the screen with one-pixel-wide concentric diamonds drawn with sequential attributes, and sets the DAC to produce a smooth gradient of each of the three primary colors and of a mix of red and blue. Run the program; I suspect you’ll be surprised at the stunning display this short program produces. Clever color manipulation is perhaps the easiest way to produce truly eye-catching effects on the PC.

    @@ -84,10 +77,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/33-04.html b/33-04.html index 9a6fd18..f9ff64e 100644 --- a/33-04.html +++ b/33-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Yogi Bear and Eurythmics Confront VGA Colors - - @@ -37,7 +30,7 @@


    -

    LISTING 33.1 L33-1.ASM

    +

    LISTING 33.1 L33-1.ASM

     ; Program to demonstrate use of the DAC registers by selecting a
     ; smoothly contiguous set of 256 colors, then filling the screen
    @@ -214,7 +207,7 @@ FillVertLoop:
            ret;
     
             endStart
    -
    +

    Note the jagged lines at the corners of the screen when you run Listing 33.1. This shows how coarse the 320x200 resolution of mode 13H actually is. Now look at how smoothly the colors blend together in the rest of the screen. This is an excellent example of how careful color selection can boost perceived resolution, as for example when drawing antialiased lines, as discussed in Chapter 42.

    @@ -239,10 +232,6 @@ FillVertLoop:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/34-01.html b/34-01.html index 3751de2..6f8cbe4 100644 --- a/34-01.html +++ b/34-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - @@ -37,16 +30,16 @@


    -

    Chapter 34
    +

    Chapter 34
    Changing Colors without Writing Pixels

    -

    Special Effects through Realtime Manipulation of DAC Colors

    +

    Special Effects through Realtime Manipulation of DAC Colors

    Sometimes, strange as it may seem, the harder you try, the less you accomplish. Brute force is fine when it suffices, but it does not always suffice, and when it does not, finesse and alternative approaches are called for. Such is the case with rapidly cycling through colors by repeatedly loading the VGA’s Digital to Analog Converter (DAC). No matter how much you optimize your code, you just can’t reliably load the whole DAC cleanly in a single frame, so you had best find other ways to use the DAC to cycle colors. What’s more, BIOS support for DAC loading is so inconsistent that it’s unusable for color cycling; direct loading through the I/O ports is the only way to go. We’ll see why next, as we explore color cycling, and then finish up this chapter and this section by cleaning up some odds and ends about VGA color.

    There’s a lot to be said about loading the DAC, so let’s dive right in and see where the complications lie.

    -

    Color Cycling

    +

    Color Cycling

    As we’ve learned in past chapters, the VGA’s DAC contains 256 storage locations, each holding one 18-bit value representing an RGB color triplet organized as 6 bits per primary color. Each and every pixel generated by the VGA is fed into the DAC as an 8-bit value (refer to Chapter 33 and to Chapter A on the companion CD-ROM to see how pixels become 8-bit values in non-256 color modes) and each 8-bit value is used to look up one of the 256 values stored in the DAC. The looked-up value is then converted to analog red, green, and blue signals and sent to the monitor to form one pixel.

    @@ -56,7 +49,7 @@

    Actually, so far as I know, you can’t. At least you can’t load the entire DAC—all 256 locations—frame after frame without producing distressing on-screen effects on at least some computers. In non-256 color modes, it is indeed possible to load the DAC quickly enough to cycle all displayed colors (of which there are 16 or fewer), so color cycling could be used successfully to cycle all colors in such modes. On the other hand, color paging (which flips among a number of color sets stored within the DAC in all modes other than 256 color mode, as discussed in Chapter A on the companion CD-ROM) can be used in non-256 color modes to produce many of the same effects as color cycling and is considerably simpler and more reliable then color cycling, so color paging is generally superior to color cycling whenever it’s available. In short, color cycling is really the method of choice for dynamic color effects only in 256-color mode—but, regrettably, color cycling is at its least reliable and capable in that mode, as we’ll see next.

    -

    The Heart of the Problem

    +

    The Heart of the Problem

    Here’s the problem with loading the entire DAC repeatedly: The DAC contains 256 color storage locations, each loaded via either 3 or 4 OUT instructions (more on that next), so at least 768 OUTs are needed to load the entire DAC. That many OUTs take a considerable amount of time, all the more so because OUTs are painfully slow on 486s and Pentiums, and because the DAC is frequently on the ISA bus (although VLB and PCI are increasingly common), where wait states are inserted in fast computers. In an 8 MHz AT, 768 OUTs alone would take 288 microseconds, and the data loading and looping that are also required would take in the ballpark of 1,800 microseconds more, for a minimum of 2 milliseconds total.

    @@ -83,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/34-02.html b/34-02.html index 875ebf1..dd87c94 100644 --- a/34-02.html +++ b/34-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - @@ -37,7 +30,7 @@


    -

    Loading the DAC via the BIOS

    +

    Loading the DAC via the BIOS

    The DAC can be loaded either directly or through subfunctions 10H (for a single DAC register) or 12H (for a block of DAC registers) of the BIOS video service interrupt 10H, function 10H, described in Chapter 33. For cycling the contents of the entire DAC, the block-load function (invoked by executing INT 10H with AH = 10H and AL = 12H to load a block of CX DAC locations, starting at location BX, from the block of RGB triplets—3 bytes per triplet—starting at ES:DX into the DAC) would be the better of the two, due to the considerably greater efficiency of calling the BIOS once rather than 256 times. At any rate, we’d like to use one or the other of the BIOS functions for color cycling, because we know that whenever possible, one should use a BIOS function in preference to accessing hardware directly, in the interests of avoiding compatibility problems. In the case of color cycling, however, it is emphatically not possible to use either of the BIOS functions, for they have problems. Serious problems.

    @@ -57,7 +50,7 @@

    Which is not to say that loading the DAC directly is a picnic either, as we’ll see next.

    -

    Loading the DAC Directly

    +

    Loading the DAC Directly

    So we must load the DAC directly in order to perform color cycling. The DAC is loaded directly by sending (with an OUT instruction) the number of the DAC location to be loaded to the DAC Write Index register at 3C8H and then performing three OUTs to write an RGB triplet to the DAC Data register at 3C9H. This approach must be repeated 256 times to load the entire DAC, requiring over a thousand OUTs in all.

    @@ -84,10 +77,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/34-03.html b/34-03.html index 92765dd..39d4587 100644 --- a/34-03.html +++ b/34-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - @@ -37,13 +30,13 @@


    -

    A Test Program for Color Cycling

    +

    A Test Program for Color Cycling

    Anyway, the choice of how to load the DAC is yours. Given that I’m not providing you with any hard-and-fast rules (mainly because there don’t seem to be any), what you need is a tool so that you can experiment with various DAC-loading approaches for yourself, and that’s exactly what you’ll find in Listing 34.1.

    Listing 34.1 draws a band of vertical lines, each one pixel wide, across the screen. The attribute of each vertical line is one greater than that of the preceding line, so there’s a smooth gradient of attributes from left to right. Once everything is set up, the program starts cycling the colors stored in however many DAC locations are specified by the CYCLE_SIZE equate; as many as all 256 DAC locations can be cycled. (Actually, CYCLE_SIZE-1 locations are cycled, because location 0 is kept constant in order to keep the background and border colors from changing, but CYCLE_SIZE locations are loaded, and it’s the number of locations we can load without problems that we’re interested in.)

    -

    LISTING 34.1 L34-1.ASM

    +

    LISTING 34.1 L34-1.ASM

     ; Fills a band across the screen with vertical bars in all 256
     ; attributes, then cycles a portion of the palette until a key is
    @@ -294,7 +287,7 @@ endif;USE_BIOS
             int     21h
     
             endstart
    -
    +


    @@ -313,10 +306,6 @@ endif;USE_BIOS
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/34-04.html b/34-04.html index 133deae..525afc6 100644 --- a/34-04.html +++ b/34-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - @@ -59,7 +52,7 @@

    One thing’s for sure, though—you’re not going to be able to cycle all 256 DAC locations cleanly once per frame on a reliable basis across the current generation of PCs. That’s why I said at the outset that brute force isn’t appropriate to the task of color cycling. That doesn’t mean that color cycling can’t be used, just that subtler approaches must be employed. Let’s look at some of those alternatives.

    -

    Color Cycling Approaches that Work

    +

    Color Cycling Approaches that Work

    First of all, I’d like to point out that when color cycling does work, it’s a thing of beauty. Assemble Listing 34.1 so that it doesn’t use the BIOS to load the DAC, doesn’t guard against interrupts, and uses 286-specific instructions if your computer supports them. Then tinker with CYCLE_SIZE until the color cycling is perfectly clean on your computer. Color cycling looks stunningly smooth, doesn’t it? And this is crude color cycling, working with the default color set; switch over to a color set that gradually works its way through various hues and saturations, and you could get something that looks for all the world like true-color animation (albeit working with a small subset of the full spectrum at any one time).

    @@ -92,10 +85,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/34-05.html b/34-05.html index b885bee..ccfeb31 100644 --- a/34-05.html +++ b/34-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Changing Colors without Writing Pixels - - @@ -47,15 +40,15 @@

    That’s what I’d do. Don’t let yourself be held back by my limited imagination, though! Color cycling may be the most complicated of all the color control techniques, but it’s also the most powerful.

    -

    Odds and Ends

    +

    Odds and Ends

    In my experience, when relying on the autoincrementing feature while loading the DAC, the Write Index register wraps back from 255 to 0, and likewise when you load a block of registers through the BIOS. So far as I know, this is a characteristic of the hardware, and should be consistent; also, Richard Wilton documents this behavior for the BIOS in the VGA bible, Programmer’s Guide to PC Video Systems, Second Edition (Microsoft Press), so you should be able to count on it. Not that I see that DAC index wrapping is especially useful, but it never hurts to understand exactly how your resources behave, and I never know when one of you might come up with a serviceable application for any particular quirk.

    -

    The DAC Mask

    +

    The DAC Mask

    There’s one register in the DAC that I haven’t mentioned yet, the DAC Mask register at 03C6H. The operation of this register is simple but powerful; it can mask off any or all of the 8 bits of pixel information coming into the DAC from the VGA. Whenever a bit of the DAC Mask register is 1, the corresponding bit of pixel information is passed along to the DAC to be used in looking up the RGB triplet to be sent to the screen. Whenever a bit of the DAC Mask register is 0, the corresponding pixel bit is ignored, and a 0 is used for that bit position in all look-ups of RGB triplets. At the extreme, a DAC Mask setting of 0 causes all 8 bits of pixel information to be ignored, so DAC location 0 is looked up for every pixel, and the entire screen displays the color stored in DAC location 0. This makes setting the DAC Mask register to 0 a quick and easy way to blank the screen.

    -

    Reading the DAC

    +

    Reading the DAC

    The DAC can be read directly, via the DAC Read Index register at 3C7H and the DAC Data register at 3C9H, in much the same way as it can be written directly by way of the DAC Write Index register—complete with autoincrementing the DAC Read Index register after every three reads. Everything I’ve said about writing to the DAC applies to reading from the DAC. In fact, reading from the DAC can even cause snow, just as loading the DAC does, so it should ideally be performed during vertical blanking.

    @@ -65,7 +58,7 @@

    Listing 34.1 illustrates reading the DAC both through the BIOS block-read function and directly, with the direct-read code capable of conditionally assembling to either guard against interrupts or not and to use REP INSB or not. As you can see, reading the DAC settings is very much symmetric with setting the DAC.

    -

    Cycling Down

    +

    Cycling Down

    And so, at long last, we come to the end of our discussion of color control on the VGA. If it has been more complex than anyone might have imagined, it has also been most rewarding. There’s as much obscure but very real potential in color control as there is anywhere on the VGA, which is to say that there’s a very great deal of potential indeed. Put color cycling or color paging together with the page flipping and image drawing techniques explored elsewhere in this book, and you’ll leave the audience gasping and wondering “How the heck did they do that?”

    @@ -86,10 +79,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-01.html b/35-01.html index 6ffb4c3..b315654 100644 --- a/35-01.html +++ b/35-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -37,10 +30,10 @@


    -

    Chapter 35
    +

    Chapter 35
    Bresenham Is Fast, and Fast Is Good

    -

    Implementing and Optimizing Bresenham’s Line-Drawing Algorithm

    +

    Implementing and Optimizing Bresenham’s Line-Drawing Algorithm

    For all the complexity of graphics design and programming, surprisingly few primitive functions lie at the heart of most graphics software. Heavily used primitives include routines that draw dots, circles, area fills, bit block logical transfers, and, of course, lines. For many years, computer graphics were created primarily with specialized line-drawing hardware, so lines are in a way the lingua franca of computer graphics. Lines are used in a wide variety of microcomputer graphics applications today, notably CAD/CAM and computer-aided engineering.

    @@ -58,7 +51,7 @@

    Notwithstanding, the line-drawing implementation in Listing 35.3 is plenty fast enough for most purposes, so let’s get the discussion underway.

    -

    The Task at Hand

    +

    The Task at Hand

    There are two important characteristics of any line-drawing function. First, it must draw a reasonable approximation of a line. A computer screen has limited resolution, and so a line-drawing function must actually approximate a straight line by drawing a series of pixels in what amounts to a jagged pattern that generally proceeds in the desired direction. That pattern of pixels must reliably suggest to the human eye the true line it represents. Second, to be usable, a line-drawing function must be fast. Minicomputers and mainframes generally have hardware that performs line drawing, but most microcomputers offer no such assistance. True, nowadays graphics accelerators such as the S3 and ATI chips have line drawing hardware, but some other accelerators don’t; when drawing lines on the latter sort of chip, when drawing on the CGA, EGA, and VGA, and when drawing sorts of lines not supported by line-drawing hardware as well, the PC’s CPU must draw lines on its own, and, as many users of graphics-oriented software know, that can be a slow process indeed.

    @@ -70,13 +63,12 @@

    Bresenham’s line-drawing algorithm, on the other hand, is uniquely suited to microcomputer implementation in that it requires no floating-point operations, no divides, and no multiplies inside the line-drawing loop. Moreover, it can be implemented with surprisingly little code.

    -

    Bresenham’s Line-Drawing Algorithm

    +

    Bresenham’s Line-Drawing Algorithm

    The key to grasping Bresenham’s algorithm is to understand that when drawing an approximation of a line on a finite-resolution display, each pixel drawn will lie either exactly on the true line or to one side or the other of the true line. The amount by which the pixel actually drawn deviates from the true line is the error of the line drawing at that point. As the drawing of the line progresses from one pixel to the next, the error can be used to tell when, given the resolution of the display, a more accurate approximation of the line can be drawn by placing a given pixel one unit of screen resolution away from its predecessor in either the horizontal or the vertical direction, or both.

    -


    - Figure 35.1
      Approximating a true line from a pixel array.

    +


    + Figure 35.1
      Approximating a true line from a pixel array.


    @@ -95,10 +87,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-02.html b/35-02.html index f1eabef..1ece3c7 100644 --- a/35-02.html +++ b/35-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -43,9 +36,8 @@

    In Figure 35.2, the X dimension is the major dimension. This means that 6 dots, one at each of X coordinates 0, 1, 2, 3, 4, and 5, will be drawn. The trick, then, is to decide on the correct Y coordinates to accompany those X coordinates.

    -


    - Figure 35.2
      Drawing between two pixel endpoints.

    +


    + Figure 35.2
      Drawing between two pixel endpoints.

    It’s easy enough to select the Y coordinates by eye in Figure 35.2. The appropriate Y coordinates are 0, 0, 1, 1, 2, 2, based on the Y coordinate closest to the line for each X coordinate. Bresenham’s algorithm makes the same selections, based on the same criterion. The manner in which it does this is by keeping a running record of the error of the line—that is, how far from the true line the current Y coordinate is—at each X coordinate, as shown in Figure 35.3. When the running error of the line indicates that the current Y coordinate deviates from the true line to the extent that the adjacent Y coordinate would be closer to the line, then the current Y coordinate is changed to that adjacent Y coordinate.

    @@ -55,9 +47,8 @@

    The third pixel has an X coordinate of 2. The running error at this point is C minus A, which is greater than 1/2 and therefore closer to the next than to the current Y coordinate. The third pixel is drawn at (2,1), and 1 is subtracted from the running error to compensate for the adjustment of one pixel in the current Y coordinate. The running error of the pixel actually drawn at this point is C minus D.

    -


    - Figure 35.3
      The error term in Bresenham’s algorithm.

    +


    + Figure 35.3
      The error term in Bresenham’s algorithm.

    The fourth pixel has an X coordinate of 3. The running error at this point is E minus D; since this is less than 1/2, the current Y coordinate doesn’t change. The fourth pixel is drawn at (3,1).

    @@ -69,7 +60,7 @@

    The above discussion summarizes the nature rather than the exact mechanism of Bresenham’s line-drawing algorithm. I’ll provide a brief seat-of-the-pants discussion of the algorithm in action when we get to the C implementation of the algorithm; for a full mathematical treatment, I refer you to pages 433-436 of Foley and Van Dam’s Fundamentals of Interactive Computer Graphics (Addison-Wesley, 1982), or pages 72-78 of the second edition of that book, which was published under the name Computer Graphics: Principles and Practice (Addison-Wesley, 1990). These sources provide the derivation of the integer-only, divide-free version of the algorithm, as well as Pascal code for drawing lines in one of the eight possible octants.

    -

    Strengths and Weaknesses

    +

    Strengths and Weaknesses

    The overwhelming strength of Bresenham’s line-drawing algorithm is speed. With no divides, no floating-point operations, and no need for variables that won’t fit in 16 bits, it is perfectly suited for PCs.

    @@ -94,10 +85,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-03.html b/35-03.html index 941f2dd..250d8f8 100644 --- a/35-03.html +++ b/35-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -37,11 +30,11 @@


    -

    An Implementation in C

    +

    An Implementation in C

    It’s time to get down and look at some actual working code. Listing 35.1 is a C implementation of Bresenham’s line-drawing algorithm for modes 0EH, 0FH, 10H, and 12H of the VGA, called as function EVGALine. Listing 35.2 is a sample program to demonstrate the use of EVGALine.

    -

    LISTING 35.1 L35-1.C

    +

    LISTING 35.1 L35-1.C

     /*
      * C implementation of Bresenham’s line drawing algorithm
    @@ -233,7 +226,7 @@ char Color;    /* color to draw line in */
        outportb(GC_INDEX, BIT_MASK_INDEX);
        outportb(GC_DATA, 0xFF);
     }
    -
    +


    @@ -252,10 +245,6 @@ char Color; /* color to draw line in */
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-04.html b/35-04.html index b6d7dfd..85c2e21 100644 --- a/35-04.html +++ b/35-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -37,7 +30,7 @@


    -

    LISTING 35.2 L35-2.C

    +

    LISTING 35.2 L35-2.C

     /*
      * Sample program to illustrate EGA/VGA line drawing routines.
    @@ -122,9 +115,9 @@ void main()
        _AX = TEXT_MODE;
        geninterrupt(BIOS_VIDEO_INT);
     }
    -
    + -

    Looking at EVGALine

    +

    Looking at EVGALine

    The EVGALine function itself performs four operations. EVGALine first sets up the VGA’s hardware so that all pixels drawn will be in the desired color. This is accomplished by setting two of the VGA’s registers, the Enable Set/Reset register and the Set/Reset register. Setting the Enable Set/Reset to the value 0FH, as is done in EVGALine, causes all drawing to produce pixels in the color contained in the Set/Reset register. Setting the Set/Reset register to the passed color, in conjunction with the Enable Set/Reset setting of 0FH, causes all drawing done by EVGALine and the functions it calls to generate the passed color. In summary, setting up the Enable Set/Reset and Set/Reset registers in this way causes the remainder of EVGALine to draw a line in the specified color.

    @@ -142,9 +135,8 @@ void main()

    Similarly, octants 0 (where X increases from start to finish) and 3 (where X decreases from start to finish) differ only in the direction in which the X coordinate moves when it changes. The difference between line drawing in octants 0 and 3 and line drawing in octants 1 and 2 is that in octants 0 and 3, since X is the major axis, the X coordinate changes on every pixel of the line and the Y coordinate changes only when the running error of the line dictates. In octants 1 and 2, the Y coordinate changes on every pixel and the X coordinate changes only when the running error dictates, since Y is the major axis.

    -


    - Figure 35.4
      Bresenham’s eight possible line orientations.

    +


    + Figure 35.4
      Bresenham’s eight possible line orientations.


    @@ -163,10 +155,6 @@ void main()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-05.html b/35-05.html index d7c1c2b..793920d 100644 --- a/35-05.html +++ b/35-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -43,11 +36,10 @@

    After calling the appropriate function to draw the line (more on those functions shortly), EVGALine restores the state of the Enable Set/Reset register to its default of zero. In this state, the Set/Reset register has no effect, so it is not necessary to restore the state of the Set/Reset register as well. EVGALine also restores the state of the Bit Mask register (which, as we will see, is modified by EVGADot, the pixel-drawing routine actually used to draw each pixel of the lines produced by EVGALine) to its default of 0FFH. While it would be more modular to have EVGADot restore the state of the Bit Mask register after drawing each pixel, it would also be considerably slower to do so. The same could be said of having EVGADot set the Enable Set/Reset and Set/Reset registers for each pixel: While modularity would improve, speed would suffer markedly.

    -


    - Figure 35.5
      EVGALine’s decision logic.

    +


    + Figure 35.5
      EVGALine’s decision logic.

    -

    Drawing Each Line

    +

    Drawing Each Line

    The Octant0 and Octant1 functions draw lines for which |DeltaX| is greater than DeltaY and lines for which |DeltaX| is less than or equal to DeltaY, respectively. The parameters to Octant0 and Octant1 are the starting point of the line, the length of the line in each dimension, and XDirection, the amount by which the X coordinate should be changed when it moves. XDirection must be either 1 (to draw toward the right edge of the screen) or -1 (to draw toward the left edge of the screen). No value is required for the amount by which the Y coordinate should be changed; since DeltaY is guaranteed to be positive, the Y coordinate always changes by 1 pixel.

    @@ -55,7 +47,7 @@

    Octant1 draws lines for which |DeltaX| is less than or equal to DeltaY. For these lines, the Y coordinate of each pixel drawn is 1 greater than the Y coordinate of the previous pixel. Whenever ErrorTerm becomes non-negative, indicating that the next X coordinate is a better approximation of the line being drawn, the X coordinate is advanced by either 1 or -1, depending on the value of XDirection. (This makes it possible for Octant1 to draw lines in both octant 1 and octant 2.)

    -

    Drawing Each Pixel

    +

    Drawing Each Pixel

    At the core of Octant0 and Octant1 is a pixel-drawing function, EVGADot. EVGADot draws a pixel at the specified coordinates in whatever color the hardware of the VGA happens to be set up for. As described earlier, since the entire line drawn by EVGALine is of the same color, line-drawing performance is improved by setting the VGA’s hardware up once in EVGALine before the line is drawn, and then drawing all the pixels in the line in the same color via EVGADot.

    @@ -84,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-06.html b/35-06.html index ab6feb0..ac5893f 100644 --- a/35-06.html +++ b/35-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -41,7 +34,7 @@

    Line drawing would be somewhat faster if the code of EVGADot were made an inline part of Octant0 and Octant1, thereby saving the overhead of preparing parameters and calling the function. Feel free to do this if you wish; I maintained EVGADot as a separate function for clarity and for ease of inserting a pixel-drawing function for a different graphics adapter, should that be desired. If you do install a pixel-drawing function for a different adapter, or a fundamentally different mode such as a 256-color SuperVGA mode, remember to remove the hardware-dependent outportb lines in EVGALine itself.

    -

    Comments on the C Implementation

    +

    Comments on the C Implementation

    EVGALine does no error checking whatsoever. My assumption in writing EVGALine was that it would be ultimately used as the lowest-level primitive of a graphics software package, with operations such as error checking and clipping performed at a higher level. Similarly, EVGALine is tied to the VGA’s screen coordinate system of (0,0) to (639,199) (in mode 0EH), (0,0) to (639,349) (in modes 0FH and 10H), or (0,0) to (639,479) (in mode 12H), with the upper left corner considered to be (0,0). Again, transformation from any coordinate system to the coordinate system used by EVGALine can be performed at a higher level. EVGALine is specifically designed to do one thing: draw lines into the display memory of the VGA. Additional functionality can be supplied by the code that calls EVGALine.

    @@ -51,7 +44,7 @@

    Given which, a high-speed assembly language version of EVGALine would seem to be a logical next step.

    -

    Bresenham’s Algorithm in Assembly

    +

    Bresenham’s Algorithm in Assembly

    Listing 35.3 is a high-performance implementation of Bresenham’s algorithm, written entirely in assembly language. The code is callable from C just as is Listing 35.1, with the same name, EVGALine, and with the same parameters. Either of the two can be linked to any program that calls EVGALine, since they appear to be identical to the calling program. The only difference between the two versions is that the sample program in Listing 35.2 runs over three times as fast on a 486 with an ISA-bus VGA when calling the assembly-language version of EVGALine as when calling the C version, and the difference would be considerably greater yet on a local bus, or with the use of write mode 3. Link each version with Listing 35.2 and compare performance—the difference is startling.

    @@ -72,10 +65,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/35-07.html b/35-07.html index 4632d71..ee480e5 100644 --- a/35-07.html +++ b/35-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bresenham Is Fast, and Fast Is Good - - @@ -37,7 +30,7 @@


    -

    LISTING 35.3 L35-3.ASM

    +

    LISTING 35.3 L35-3.ASM

     ; Fast assembler implementation of Bresenham’s line-drawing algorithm
     ; for the EGA and VGA. Works in modes 0Eh, 0Fh, 10h, and 12h.
    @@ -407,7 +400,7 @@ EVGALineDone:
     _EVGALine       endp
     
             end
    -
    +

    An explanation of the workings of the code in Listing 35.3 would be a lengthy one, and would be redundant since the basic operation of the code in Listing 35.3 is no different from that of the code in Listing 35.1, although the implementation is much changed due to the nature of assembly language and also due to designing for speed rather than for clarity. Given that you thoroughly understand the C implementation in Listing 35.1, the assembly language implementation in Listing 35.3, which is well-commented, should speak for itself.

    @@ -432,10 +425,6 @@ _EVGALine endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/36-01.html b/36-01.html index c5c710a..7c33818 100644 --- a/36-01.html +++ b/36-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - @@ -37,10 +30,10 @@


    -

    Chapter 36
    +

    Chapter 36
    The Good, the Bad, and the Run-Sliced

    -

    Faster Bresenham Lines with Run-Length Slice Line Drawing

    +

    Faster Bresenham Lines with Run-Length Slice Line Drawing

    Years ago, I worked at a company that asked me to write blazingly fast line-drawing code for an AutoCAD driver. I implemented the basic Bresenham’s line-drawing algorithm; streamlined it as much as possible; special-cased horizontal, diagonal, and vertical lines; broke out separate, optimized routines for lines in each octant; and massively unrolled the loops. When I was done, I had line drawing down to a mere five or six instructions per pixel, and I handed the code over to the AutoCAD driver person, content in the knowledge that I had pushed the theoretical limits of the Bresenham’s algorithm on the 80x86 architecture, and that this was as fast as line drawing could get on a PC. That feeling lasted for about a week, until Dave Miller, who these days is a Windows display-driver whiz at Engenious Solutions, casually mentioned Bresenham’s faster run-length slice line-drawing algorithm.

    @@ -54,7 +47,7 @@

    I learned two important safety tips from my line-drawing experience; neither involves the possible destruction of the universe, so far as I know, but they are nonetheless worth keeping in mind. First, never, never, never think you’ve written the fastest possible code. Odds are, you haven’t. Run your code past another good programmer, and he or she will probably say, “But why don’t you do this?” and you’ll realize that you could indeed do that, and your code would then be faster. Or relax and come back to your code later, and you may well see another, faster approach. There are a million ways to implement code for any task, and you can almost always find a faster way if you need to.

    -

    Second, when performance matters, never have your code perform the same calculation more than once. This sounds obvious, but it’s astonishing how often it’s ignored. For example, consider this snippet of code:

    +

    Second, when performance matters, never have your code perform the same calculation more than once. This sounds obvious, but it’s astonishing how often it’s ignored. For example, consider this snippet of code:

     for (i=0; i<RunLength; i++)
     {
    @@ -69,9 +62,9 @@ for (i=0; i<RunLength; i++)
        }
     }
     
    -
    + -

    Here, the programmer knows which way the line is going before the main loop begins—but nonetheless performs that test every time through the loop, when calculating the address of the next pixel. Far better to perform the test only once, outside the loop, as shown here:

    +

    Here, the programmer knows which way the line is going before the main loop begins—but nonetheless performs that test every time through the loop, when calculating the address of the next pixel. Far better to perform the test only once, outside the loop, as shown here:

     if (XDelta > 0)
     {
    @@ -88,21 +81,20 @@ else
        }
     }
     
    -
    +

    Think of it this way: A program is a state machine. It takes a set of inputs and produces a corresponding set of outputs by passing through a set of states. Your primary job as a programmer is to implement the desired state machine. Your additional job as a performance programmer is to minimize the lengths of the paths through the state machine. This means performing as many tests and calculations as possible outside the loops, so that the loops themselves can do as little work—that is, pass through as few states—as possible.

    Which brings us full circle to Bresenham’s run-length slice line-drawing algorithm, which just happens to be an excellent example of a minimized state machine. In case you’re fuzzy on the good/bad performance thing, that’s “good”—as in fast.

    -

    Run-Length Slice Fundamentals

    +

    Run-Length Slice Fundamentals

    First off, I have a confession to make: I’m not sure that the algorithm I’ll discuss is actually, precisely Bresenham’s run-length slice algorithm. It’s been a long time since I read about this algorithm; in the intervening years, I’ve misplaced Bresenham’s article, and have been unable to unearth it. As a result, I had to derive the algorithm from scratch, which was admittedly more fun than reading about it, and also ensured that I understood it inside and out. The upshot is that what I discuss may or may not be Bresenham’s run-length slice algorithm—but it surely is fast.

    The place to begin understanding the run-length slice algorithm is the standard Bresenham’s line-drawing algorithm. (I discussed the standard Bresenham’s line-drawing algorithm at length in the previous chapter.) The basis of the standard approach is stepping one pixel at a time along the major axis (the longer dimension of the line), while maintaining an integer error term that indicates at each major-axis step how close the line is to advancing halfway to the next pixel along the minor axis. Figure 36.1 illustrates standard Bresenham’s line drawing. The key point here is that a calculation and a test are performed once for each step along the major axis.

    -


    - Figure 36.1
      Standard Bresenham’s line drawing.

    +


    + Figure 36.1
      Standard Bresenham’s line drawing.


    @@ -121,10 +113,6 @@ else
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/36-02.html b/36-02.html index e0fb193..28adf94 100644 --- a/36-02.html +++ b/36-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - @@ -43,23 +36,20 @@

    Take a moment to let the idea behind run-length slice drawing soak in. Periodic decisions must be made to control pixel placement. The key to speed is to make those decisions as infrequently and as quickly as possible. Of course, it will work to make a decision at each pixel—that’s standard Bresenham’s. However, most of those per-pixel decisions are redundant, and in fact we have enough information before we begin drawing to know which are the redundant decisions. Run-length slice drawing is exactly equivalent to standard Bresenham’s, but it pares the decision-making process down to a minimum. It’s somewhat analogous to the difference between finding the greatest common divisor of two numbers using Euclid’s algorithm and finding it by trying every possible divisor. Both approaches produce the desired result, but that which takes maximum advantage of the available information and minimizes redundant work is preferable.

    -


    - Figure 36.2
      Run-length slice line drawing.

    +


    + Figure 36.2
      Run-length slice line drawing.

    -


    - Figure 36.3
      Runs in a slope 1/3.5 line.

    +


    + Figure 36.3
      Runs in a slope 1/3.5 line.

    -

    Run-Length Slice Implementation

    +

    Run-Length Slice Implementation

    We know that for any line, a given run will always be one of two possible lengths. How, though, do we know which length to select? Surprisingly, this is easy to determine. For the following discussion, assume that we have a slope of 1/3.5, so that X is the major axis; however, the discussion also applies to Y-major lines, with X and Y reversed.

    The minimum possible length for any run in an X-major line is int(XDelta/YDelta), where XDelta is the X-dimension of the line and YDelta is the Y-dimension. The maximum possible length is int(XDelta/YDelta)+ 1. The trick, then, is knowing which of these two lengths to select for each run. To see how we can make this selection, refer to Figure 36.4. For each one-pixel step along the minor axis (Y, in this case), we advance at least three pixels. The full advance distance along X (the major axis) is actually three-plus pixels, because there is also a fractional portion to the advance along X for a single-pixel Y step. This fractional advance is the key to deciding when to add an extra pixel to a run. The fraction indicates what portion of an extra pixel we advance along X (the major axis) during each run. If we keep a running sum of the fractional parts, we have a measure of how close we are to needing an extra pixel; when the fractional sum reaches 1, it’s time to add an extra pixel to the current run. Then, we can subtract 1 from the running sum (because we just advanced one pixel), and continue on.

    -


    - Figure 36.4
      How the error term determines run length.

    +


    + Figure 36.4
      How the error term determines run length.

    Practically speaking, however, we can’t work with fractions because floating-point arithmetic is slow and fixed-point arithmetic is imprecise. Therefore, we take a cue from standard Bresenham’s and scale all the error-term calculations up so that we can work with integers. The fractional X (major axis) advance per one-pixel Y (minor axis) advance is the fractional portion of XDelta/YDelta. This value is exactly equivalent to (XDelta % YDelta)/YDelta. We’ll scale this up by multiplying it by YDelta*2, so that the amount by which we adjust the error term up for each one-pixel minor-axis advance is (XDelta % YDelta)*2.

    @@ -86,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/36-03.html b/36-03.html index 2d0a73f..09c250c 100644 --- a/36-03.html +++ b/36-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - @@ -37,7 +30,7 @@


    -

    Run-Length Slice Details

    +

    Run-Length Slice Details

    A couple of run-length slice implementation details yet remain. First is the matter of how error-term turnover is detected. This is done in much the same way as it is with standard Bresenham’s: The error term is maintained as a negative valve and advances for each step; when the error term reaches 0, it’s time to add an extra pixel to the current run. This means that we only have to test for carry after advancing the error term to determine whether or not to add an extra pixel to each run. (Actually, the code in this chapter tests for the error term being greater than zero, but the assembly code in the next chapter will use the very efficient carry approach.)

    @@ -49,11 +42,10 @@

    That’s all there is to run-length slice line drawing; the partial first and last runs are the only tricky part. Listing 36.1 is a run-length slice implementation in C. This is not an optimized implementation, nor is it meant to be; this listing is provided so that you can see how the run-length slice algorithm works. In the next chapter, I’ll move on to an optimized version, but for now, Listing 36.1 will make it much easier to grasp the principles of run-length slice drawing, and to understand the optimized code I’ll present in the next chapter.

    -


    - Figure 36.5
      Balancing run-length slice lines: a) unbalanced; b) balanced.

    +


    + Figure 36.5
      Balancing run-length slice lines: a) unbalanced; b) balanced.

    -

    LISTING 36.1 L36-1.C

    +

    LISTING 36.1 L36-1.C

     /* Run-length slice line drawing implementation for mode 0x13, the VGA’s
     320x200 256-color mode. Not optimized! Tested with Borland C++ in
    @@ -293,7 +285,7 @@ void DrawVerticalRun(char far **ScreenPtr, int XAdvance,
        *ScreenPtr = WorkingScreenPtr;
     }
     
    -
    +


    @@ -312,10 +304,6 @@ void DrawVerticalRun(char far **ScreenPtr, int XAdvance,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/36-04.html b/36-04.html index c6e8b0b..1d9bce6 100644 --- a/36-04.html +++ b/36-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Good, the Bad, and the Run-Sliced - - @@ -39,7 +32,7 @@

    Notwithstanding that it’s not optimized, Listing 36.1 is reasonably fast. If you run Listing 36.2 (a sample line-drawing program that you can use to test-drive Listing 36.1), you may be as surprised as I was at how quickly the screen fills with vectors, considering that Listing 36.1 is entirely in C and has some redundant divides. Or perhaps you won’t be surprised—in which case I suggest you not miss the next chapter.

    -

    LISTING 36.2 L36-2.C

    +

    LISTING 36.2 L36-2.C

     /* Sample line-drawing program. Uses the optimized
     line-drawing functions coded in LListing L36.1.C.
    @@ -115,7 +108,7 @@ int main()
        regs.x.ax = TEXT_MODE;
        int86(BIOS_VIDEO_INT, &regs, &regs);
     }
    -
    +


    @@ -134,10 +127,6 @@ int main()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/37-01.html b/37-01.html index ceada18..6b61fd8 100644 --- a/37-01.html +++ b/37-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dead Cats and Lightning Lines - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dead Cats and Lightning Lines - - @@ -37,10 +30,10 @@


    -

    Chapter 37
    +

    Chapter 37
    Dead Cats and Lightning Lines

    -

    Optimizing Run-Length Slice Line Drawing in a Major Way

    +

    Optimizing Run-Length Slice Line Drawing in a Major Way

    As I write this, the wife, the kid, and I are in the throes of yet another lightning-quick transcontinental move, this time to Redmond, Washington, to work for You Know Who. Moving is never fun, but what makes it worse for us is the pets. Getting them into kennels and to the airport is hard; there’s always the possibility that they might not be allowed to fly because of the weather; and, worst of all, they might not make it. Animals don’t usually end up injured or dead, but it does happen.

    @@ -54,11 +47,11 @@

    Okay, but what’s the point? The point is, if it isn’t broken, don’t fix it. And if it is broken, maybe that’s all right, too. Which brings us, neat as a pin, to the topic of drawing lines in a serious hurry.

    -

    Fast Run-Length Slice Line Drawing

    +

    Fast Run-Length Slice Line Drawing

    In the last chapter, we examined the principles of run-length slice line drawing, which draws lines a run at a time rather than a pixel at a time, a run being a series of pixels along the major (longer) axis. It’s time to turn theory into useful practice by developing a fast assembly version. Listing 37.1 is the assembly version, in a form that’s plug-compatible with the C code from the previous chapter.

    -

    LISTING 37.1 L37-1.ASM

    +

    LISTING 37.1 L37-1.ASM

     ; Fast run-length slice line drawing implementation for mode 0x13, the VGA’s
     ; 320x200 256-color mode.
    @@ -408,7 +401,7 @@ Done:
           ret
     _LineDraw   endp
           end
    -
    +


    @@ -427,10 +420,6 @@ _LineDraw endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/37-02.html b/37-02.html index d86de02..948ba60 100644 --- a/37-02.html +++ b/37-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dead Cats and Lightning Lines - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dead Cats and Lightning Lines - - @@ -37,7 +30,7 @@


    -

    How Fast Is Fast?

    +

    How Fast Is Fast?

    Your first question is likely to be the following: Just how fast is Listing 37.1? Is it optimized to the hilt or just pretty fast? The quick answer is: It’s fast. Listing 37.1 draws lines at a rate of nearly 1 million pixels per second on my 486/33, and is capable of still faster drawing, as I’ll discuss shortly. (The heavily optimized AutoCAD line-drawing code that I mentioned in the last chapter drew 150,000 pixels per second on an EGA in a 386/16, and I thought I had died and gone to Heaven. Such is progress.) The full answer is a more complicated one, and ties in to the principle that if it is broken, maybe that’s okay—and to the principle of looking before you leap, also known as profiling before you optimize.

    @@ -59,7 +52,7 @@

    Profile before you optimize.

    -

    Further Optimizations

    +

    Further Optimizations

    Following is a quick tour of some of the many possible further optimizations to Listing 37.1.

    @@ -92,10 +85,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/38-01.html b/38-01.html index 93bae8d..69ad6b3 100644 --- a/38-01.html +++ b/38-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - @@ -37,10 +30,10 @@


    -

    Chapter 38
    +

    Chapter 38
    The Polygon Primeval

    -

    Drawing Polygons Efficiently and Quickly

    +

    Drawing Polygons Efficiently and Quickly

    “Give me but one firm spot on which to stand, and I will move the Earth.”

    @@ -52,7 +45,7 @@

    And slow computer graphics is scarcely worth the bother.

    -

    Filled Polygons

    +

    Filled Polygons

    A polygon is simply a shape formed by lines laid end to end to form a continuous, closed path. A polygon is filled by setting all pixels within the polygon’s boundaries to a color or pattern. For now, we’ll work only with polygons filled with solid colors.

    @@ -62,13 +55,12 @@

    Why bother to distinguish between convex, nonconvex, and complex polygons? Easy: performance, especially when it comes to filling convex polygons. We’re going to start with filled convex polygons; they’re widely useful and will serve well to introduce some of the subtler complexities of polygon drawing, not the least of which is the slippery concept of “inside.”

    -

    Which Side Is Inside?

    +

    Which Side Is Inside?

    The basic principle of polygon filling is decomposing each polygon into a series of horizontal lines, one for each horizontal row of pixels, or scan line, within the polygon (a process I’ll call scan conversion), and drawing the horizontal lines. I’ll refer to the entire process as rasterization. Rasterization of convex polygons is easily done by starting at the top of the polygon and tracing down the left and right sides, one scan line (one vertical pixel) at a time, filling the extent between the two edges on each scan line, until the bottom of the polygon is reached. At first glance, rasterization does not seem to be particularly complicated, although it should be apparent that this simple approach is inadequate for nonconvex polygons.

    -


    - Figure 38.1
      Convex, nonconvex, and complex polygons.

    +


    + Figure 38.1
      Convex, nonconvex, and complex polygons.

    There are a couple of complications, however. The lesser complication is how to rasterize the polygon efficiently, given that it’s difficult to write fast code that simultaneously traces two edges and fills the space between them. The solution is to decouple the process of scan-converting the polygon into a list of horizontal lines from that of drawing the horizontal lines. One device-independent routine can trace along the two edges and build a list of the beginning and end coordinates of the polygon on each raster line. Then a second, device-specific, routine can draw from the list after the entire polygon has been scanned. We’ll see this in action shortly.

    @@ -76,9 +68,8 @@

    It’s no crime to use standard lines to trace out a polygon, rather than drawing only interior pixels. In fact, there are certain advantages: For example, the edges of a filled polygon will match the edges of the same polygon drawn unfilled. Such polygons will look pretty much as they’re supposed to, and all drawing on raster displays is, after all, only an approximation of an ideal.

    -


    - Figure 38.2
      Drawing polygons with standard line-drawing algorithms.

    +


    + Figure 38.2
      Drawing polygons with standard line-drawing algorithms.

    There’s one great drawback to tracing polygons with standard lines, however: Adjacent polygons won’t fit together properly, as shown in Figure 38.3. If you use six equilateral triangles to make a hexagon, for example, the edges of the triangles will overlap when traced with standard lines, and more recently drawn triangles will wipe out portions of their predecessors. Worse still, odd color effects will show up along the polygon boundaries if XOR drawing is used. Consequently, filling out to the boundary lines just won’t do for drawing images composed of fitted-together polygons. And because fitting polygons together is exactly what I have in mind, we need a different approach.

    @@ -99,10 +90,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/38-02.html b/38-02.html index b68a529..85413cf 100644 --- a/38-02.html +++ b/38-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - @@ -37,7 +30,7 @@


    -

    How Do You Fit Polygons Together?

    +

    How Do You Fit Polygons Together?

    How, then, do you fit polygons together? Very carefully. First, the line-tracing algorithm must be adjusted so that it selects only those pixels that are truly inside the polygon. This basically requires shifting a standard line-drawing algorithm horizontally by one half-pixel toward the polygon’s interior. That leaves the issue of how to handle points that are exactly on the boundary, and points that lie at vertices, so that those points are drawn once and only once. To deal with that, we’re going to adopt the following rules:

    @@ -67,11 +60,11 @@

    For our purposes, nonoverlapping polygons are the way to go, so let’s have at them.

    -

    Filling Non-Overlapping Convex Polygons

    +

    Filling Non-Overlapping Convex Polygons

    Without further ado, Listing 38.1 contains a function, FillConvexPolygon, that accepts a list of points that describe a convex polygon, with the last point assumed to connect to the first, and scans it into a list of lines to fill, then passes that list to the function DrawHorizontalLineList in Listing 38.2. Listing 38.3 is a sample program that calls FillConvexPolygon to draw polygons of various sorts, and Listing 38.4 is a header file included by the other listings. Here are the listings; we’ll pick up discussion on the other side.

    -

    LISTING 38.1 L38-1.C

    +

    LISTING 38.1 L38-1.C

      /* Color-fills a convex polygon. All vertices are offset by (XOffset,
         YOffset). “Convex” means that every horizontal line drawn through
    @@ -280,9 +273,9 @@
         }
         *EdgePointPtr = WorkingEdgePointPtr;   /* advance caller’s ptr */
      }
    -
    + -

    LISTING 38.2 L38-2.C

    +

    LISTING 38.2 L38-2.C

      /* Draws all pixels in the list of horizontal lines passed in, in
         mode 13h, the VGA’s 320x200 256-color mode. Uses a slow pixel-by-
    @@ -329,7 +322,7 @@
      #endif
         *ScreenPtr = (unsigned char)Color;
      }
    -
    +


    @@ -348,10 +341,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/38-03.html b/38-03.html index 3cb50d2..e81982c 100644 --- a/38-03.html +++ b/38-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - @@ -37,7 +30,7 @@


    -

    LISTING 38.3 L38-3.C

    +

    LISTING 38.3 L38-3.C

      /* Sample program to exercise the polygon-filling routines. This code
         and all polygon-filling code has been tested with Borland and
    @@ -130,9 +123,9 @@
         regset.x.ax = 0x0003;   /* AL = 3 selects 80x25 text mode */
         int86(0x10, &regset, &regset);
      }
    -
    + -

    LISTING 38.4 POLYGON.H

    +

    LISTING 38.4 POLYGON.H

      /* POLYGON.H: Header file for polygon-filling code */
     
    @@ -168,7 +161,7 @@
         struct HLine * HLinePtr;   /* pointer to list of horz lines */
      };
     
    -
    +


    @@ -187,10 +180,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/38-04.html b/38-04.html index 02828f6..13d5d80 100644 --- a/38-04.html +++ b/38-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Polygon Primeval - - @@ -43,10 +36,10 @@

    Listing 38.1 first finds the top and bottom of the polygon, then works out from the top point to find the two ends of the top edge. If the ends are at different locations, the top is flat, which has two implications. First, it’s easy to find the starting vertices and directions through the vertex list for the left and right edges. (To scan-convert them properly, we must first determine which edge is which.) Second, the top scan line of the polygon should be drawn without the rightmost pixel, because only the rightmost pixel of the horizontal edge that makes up the top scan line is part of a right edge.

    -

    If, on the other hand, the ends of the top edge are at the same location, the top is pointed. In that case, the top scan line of the polygon isn’t drawn; it’s part of the right-edge line that starts at the top vertex. (It’s part of a left-edge line, too, but the right edge overrides.) When the top isn’t flat, it’s more difficult to tell in which direction through the vertex list the right and left edges go, because both edges start at the top vertex. The solution is to compare the slopes from the top vertex to the ends of the two lines coming out of it in order to see which is leftmost. The calculations in Listing 38.1 involving the various deltas do this, using a rearranged form of the slope-based equation:

    +

    If, on the other hand, the ends of the top edge are at the same location, the top is pointed. In that case, the top scan line of the polygon isn’t drawn; it’s part of the right-edge line that starts at the top vertex. (It’s part of a left-edge line, too, but the right edge overrides.) When the top isn’t flat, it’s more difficult to tell in which direction through the vertex list the right and left edges go, because both edges start at the top vertex. The solution is to compare the slopes from the top vertex to the ends of the two lines coming out of it in order to see which is leftmost. The calculations in Listing 38.1 involving the various deltas do this, using a rearranged form of the slope-based equation:

     (DeltaYN/DeltaXN)>(DeltaYP/DeltaXP)
    -
    +

    Once we know where the left edge starts in the vertex list, we can scan-convert it a line segment at a time until the bottom vertex is reached. Each point is stored as the starting X coordinate for the corresponding scan line in the list we’ll pass to DrawHorizontalLineList. The nearest X coordinate on each scan line that’s on or to the right of the left edge is selected. The last point of each line segment making up the left edge isn’t scan-converted, producing two desirable effects. First, it avoids drawing each vertex twice; two lines come into every vertex, but we want to scan-convert each vertex only once. Second, not scan-converting the last point of each line causes the bottom scan line of the polygon not to be drawn, as required by our rules. The first scan line of the polygon is also skipped if the top isn’t flat.

    @@ -56,7 +49,7 @@

    Finis.

    -

    Oddball Cases

    +

    Oddball Cases

    Listing 38.1 handles zero-length segments (multiple vertices at the same location) by ignoring them, which will be useful down the road because scaled-down polygons can end up with nearby vertices moved to the same location. Horizontal line segments are fine anywhere in a polygon, too. Basically, Listing 38.1 scan-converts between active edges (the edges that define the extent of the polygon on each scan line) and both horizontal and zero-length lines are non-active; neither advances to another scan line, so they don’t affect the edges being scanned.

    @@ -79,10 +72,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/39-01.html b/39-01.html index 372877d..89b18b0 100644 --- a/39-01.html +++ b/39-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - @@ -37,10 +30,10 @@


    -

    Chapter 39
    +

    Chapter 39
    Fast Convex Polygons

    -

    Filling Polygons in a Hurry

    +

    Filling Polygons in a Hurry

    In the previous chapter, we explored the surprisingly intricate process of filling convex polygons. Now we’re going to fill them an order of magnitude or so faster.

    @@ -60,7 +53,7 @@

    In short, even though doing so runs counter to current trends, it helps to understand how things work, especially when they’re very visible parts of the software you develop. That said, let’s learn more about filling convex polygons.

    -

    Fast Convex Polygon Filling

    +

    Fast Convex Polygon Filling

    In addressing the topic of filling convex polygons in the previous chapter, the implementation we came up with met all of our functional requirements. In particular, it met stringent rules that guaranteed that polygons would never overlap or have gaps at shared edges, an important consideration when building polygon-based images. Unfortunately, the implementation was also slow as molasses. In this chapter we’ll work up polygon-filling code that’s fast enough to be truly usable.

    @@ -76,7 +69,7 @@

    The amount of time that the previous chapter’s sample program spent in each of these areas is shown in Table 39.1. As you can see, half the time was spent drawing and the other half was spent tracing the polygon edges (the time spent in FillConvexPolygon was relatively minuscule), so we have our choice of where to begin optimizing.

    -

    Fast Drawing

    +

    Fast Drawing

    Let’s start with drawing, which is easily sped up. The previous chapter’s code used a double-nested loop that called a draw-pixel function to plot each pixel in the polygon individually. That’s a ridiculous approach in a graphics mode that offers linearly mapped memory, as does VGA mode 13H, the mode in which we’re working. At the very least, we could point a far pointer to the left edge of each polygon scan line, then draw each pixel in that scan line in quick succession, using something along the lines of *ScrPtr++ = FillColor; inside a loop.

    @@ -99,10 +92,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/39-02.html b/39-02.html index 7b5f3d2..bc407d9 100644 --- a/39-02.html +++ b/39-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - @@ -201,7 +194,7 @@ -

    LISTING 39.1 L39-1.C

    +

    LISTING 39.1 L39-1.C

     /* Draws all pixels in the list of horizontal lines passed in, in
        mode 13h, the VGA’s 320x200 256-color mode. Uses memset to fill
    @@ -242,13 +235,13 @@ void DrawHorizontalLineList(struct HLineList * HLineListPtr,
           ScreenPtr += SCREEN_WIDTH; /* point to next scan line start */
        }
     }
    -
    +

    At this point, I’d like to mention that benchmarks are notoriously unreliable; the results in Table 39.1 are accurate only for the test program, and only when running on a particular system. Results could be vastly different if smaller, larger, or more complex polygons were drawn, or if a faster or slower computer/VGA combination were used. These factors notwithstanding, the test program does fill a variety of polygons of varying complexity sized from large to small and in between, and certainly the order of magnitude difference between Listing 39.1 and the old version of DrawHorizontalLineList is a clear indication of which code is superior.

    Anyway, Listing 39.1 has the desired effect of vastly improving drawing time. There are cycles yet to be had in the drawing code, but as tracing polygon edges now takes 92 percent of the polygon filling time, it’s logical to optimize the tracing code next.

    -

    Fast Edge Tracing

    +

    Fast Edge Tracing

    There’s no secret as to why last chapter’s ScanEdge was so slow: It used floating point calculations. One secret of fast graphics is using integer or fixed-point calculations, instead. (Sure, the floating point code would run faster if a math coprocessor were installed, but it would still be slower than the alternatives; besides, why require a math coprocessor when you don’t have to?) Both integer and fixed-point calculations are fast. In many cases, fixed-point is faster, but integer calculations have one tremendous virtue: They’re completely accurate. The tiny imprecision inherent in either fixed or floating-point calculations can result in occasional pixels being one position off from their proper location. This is no great tragedy, but after going to so much trouble to ensure that polygons don’t overlap at common edges, why not get it exactly right?

    @@ -279,10 +272,6 @@ void DrawHorizontalLineList(struct HLineList * HLineListPtr,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/39-03.html b/39-03.html index 2fc0a6a..6aef6de 100644 --- a/39-03.html +++ b/39-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - @@ -41,7 +34,7 @@

    Ya gotta love that integer arithmetic.

    -

    LISTING 39.2 L39-2.C

    +

    LISTING 39.2 L39-2.C

     /* Scan converts an edge from (X1,Y1) to (X2,Y2), not including the
        point at (X2,Y2). If SkipFirst == 1, the point at (X1,Y1) isn’t
    @@ -158,9 +151,9 @@ void ScanEdge(int X1, int Y1, int X2, int Y2, int SetXStart,
     
        *EdgePointPtr = WorkingEdgePointPtr;   /* advance caller’s ptr */
     }
    -
    + -

    The Finishing Touch: Assembly Language

    +

    The Finishing Touch: Assembly Language

    The C implementation in Listing 39.2 is now nearly 20 times as fast as the original, which is good enough for most purposes. Still, it requires that one of the large data models be used (for memset ), and it’s certainly not the fastest possible code. The obvious next step is assembly language.

    @@ -183,10 +176,6 @@ void ScanEdge(int X1, int Y1, int X2, int Y2, int SetXStart,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/39-04.html b/39-04.html index 3da5d6f..66d17e5 100644 --- a/39-04.html +++ b/39-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - @@ -53,7 +46,7 @@ -

    LISTING 39.3 L39-3.ASM

    +

    LISTING 39.3 L39-3.ASM

     ; Draws all pixels in the list of horizontal lines passed in, in
     ; mode 13h, the VGA’s 320x200 256-color mode. Uses REP STOS to fill
    @@ -140,15 +133,15 @@ FillDone:
        ret
     _DrawHorizontalLineList   endp
        end
    -
    + -

    Maximizing REP STOS

    +

    Maximizing REP STOS

    Listing 39.3 doesn’t take the easy way out and use REP STOSB to fill each scan line; instead, it uses REP STOSW to fill as many pixel pairs as possible via word-sized accesses, using STOSB only to do odd bytes. Word accesses to odd addresses are always split by the processor into 2-byte accesses. Such word accesses take twice as long as word accesses to even addresses, so Listing 39.3 makes sure that all word accesses occur at even addresses, by performing a leading STOSB first if necessary.

    Listing 39.3 is another case in which it’s worth knowing the environment in which your code will run. Extra code is required to perform aligned word-at-a-time filling, resulting in extra overhead. For very small or narrow polygons, that overhead might overwhelm the advantage of drawing a word at a time, making plain old REP STOSB faster.

    -

    Faster Edge Tracing

    +

    Faster Edge Tracing

    Finally, Listing 39.4 is an assembly language version of ScanEdge. Listing 39.4 is a relatively straightforward translation from C to assembly, but is nonetheless about twice as fast as Listing 39.2.

    @@ -175,10 +168,6 @@ _DrawHorizontalLineList endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/39-05.html b/39-05.html index b600f02..94cbced 100644 --- a/39-05.html +++ b/39-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast Convex Polygons - - @@ -37,7 +30,7 @@


    -

    LISTING 39.4 L39-4.ASM

    +

    LISTING 39.4 L39-4.ASM

     ; Scan converts an edge from (X1,Y1) to (X2,Y2), not including the
     ; point at (X2,Y2). If SkipFirst == 1, the point at (X1,Y1) isn’t
    @@ -205,7 +198,7 @@ ScanEdgeExit:
             ret
     _ScanEdge   endp
             end
    -
    +


    @@ -224,10 +217,6 @@ _ScanEdge endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/40-01.html b/40-01.html index b7d42ad..c184de7 100644 --- a/40-01.html +++ b/40-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - @@ -37,10 +30,10 @@


    -

    Chapter 40
    +

    Chapter 40
    Of Songs, Taxes, and the Simplicity of Complex Polygons

    -

    Dealing with Irregular Polygonal Areas

    +

    Dealing with Irregular Polygonal Areas

    Every so often, my daughter asks me to sing her to sleep. (If you’ve ever heard me sing, this may cause you concern about either her hearing or her judgement, but love knows no bounds.) As any parent is well aware, singing a young child to sleep can easily take several hours, or until sunrise, whichever comes last. One night, running low on children’s songs, I switched to a Beatles medley, and at long last her breathing became slow and regular. At the end, I softly sang “A Hard Day’s Night,” then quietly stood up to leave. As I tiptoed out, she said, in a voice not even faintly tinged with sleep, “Dad, what do they mean, ‘working like a dog’? Chasing a stick? That doesn’t make sense; people don’t chase sticks.”

    @@ -50,7 +43,7 @@

    Filling arbitrary polygons is such a case.

    -

    Filling Arbitrary Polygons

    +

    Filling Arbitrary Polygons

    In Chapter 38, I described three types of polygons: convex, nonconvex, and complex. The RenderMan Companion, a terrific book by Steve Upstill (Addison-Wesley, 1990) has an intuitive definition of convex: If a rubber band stretched around a polygon touches all vertices in the order they’re defined, then the polygon is convex. If a polygon has intersecting edges, it’s complex. If a polygon doesn’t have intersecting edges but isn’t convex, it’s nonconvex. Nonconvex is a special case of complex, and convex is a special case of nonconvex. (Which, I’m well aware, makes nonconvex a lousy name—noncomplex would have been better—but I’m following X Window System nomenclature here.)

    @@ -58,21 +51,19 @@

    Before we dive into complex polygon filling, I’d like to point out that the code in this chapter, like all polygon filling code I’ve ever seen, requires that the caller describe the type of the polygon to be filled. Often, however, the caller doesn’t know what type of polygon it’s passing, or specifies complex for simplicity, because that will work for all polygons; in such a case, the polygon filler will use the slow complex-fill code even if the polygon is, in fact, a convex polygon. In Chapter 41, I’ll discuss one way to improve this situation.

    -

    Active Edges

    +

    Active Edges

    The basic premise of filling a complex polygon is that for a given scan line, we determine all intersections between the polygon’s edges and that scan line and then fill the spans between the intersections, as shown in Figure 40.1. (Section 3.6 of Foley and van Dam’s Computer Graphics, Second Edition provides an overview of this and other aspects of polygon filling.) There are several rules that might be used to determine which spans are drawn and which aren’t; we’ll use the odd/even rule, which specifies that drawing turns on after odd-numbered intersections (first, third, and so on) and off after even-numbered intersections.

    The question then becomes how can we most efficiently determine which edges cross each scan line and where? As it happens, there is a great deal of coherence from one scan line to the next in a polygon edge list, because each edge starts at a given Y coordinate and continues unbroken until it ends. In other words, edges don’t leap about and stop and start randomly; the X coordinate of an edge at one scan line is a consistent delta from that edge’s X coordinate at the last scan line, and that is consistent for the length of the line.

    -


    - Figure 40.1
      Filling one scan line by finding intersecting edges.

    +


    + Figure 40.1
      Filling one scan line by finding intersecting edges.

    This allows us to reduce the number of edges that must be checked for intersection; on any given scan line, we only need to check for intersections with the currently active edges—edges that start on that scan line, plus all edges that start on earlier (above) scan lines and haven’t ended yet—as shown in Figure 40.2. This suggests that we can proceed from the top scan line of the polygon to the bottom, keeping a running list of currently active edges—called the active edge table (AET)—with the edges sorted in order of ascending X coordinate of intersection with the current scan line. Then, we can simply fill each scan line in turn according to the list of active edges at that line.

    -


    - Figure 40.2
      Checking currently active edges (solid lines).

    +


    + Figure 40.2
      Checking currently active edges (solid lines).

    Maintaining the AET from one scan line to the next involves three steps: First, we must add to the AET any edges that start on the current scan line, making sure to keep the AET X-sorted for efficient odd/even scanning. Second, we must remove edges that end on the current scan line. Third, we must advance the X coordinates of active edges with the same sort of error term-based, Bresenham’s-like approach we used for convex polygons, again ensuring that the AET is X-sorted after advancing the edges.

    @@ -93,10 +84,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/40-02.html b/40-02.html index 22f8146..cad5986 100644 --- a/40-02.html +++ b/40-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - @@ -59,13 +52,12 @@
    7.  If either the AET or GET isn’t empty, go to step 2.
    -


    - Figure 40.3
      The global and active edge tables as linked lists.

    +


    + Figure 40.3
      The global and active edge tables as linked lists.

    That’s really all there is to it. Compare Listing 40.1 to the fast convex polygon filling code from Chapter 39, and you’ll see that, contrary to expectation, complex polygon filling is indeed one of the more sane and sensible corners of the universe.

    -

    LISTING 40.1 L40-1.C

    +

    LISTING 40.1 L40-1.C

      /* Color-fills an arbitrarily-shaped polygon described by VertexList.
         If the first and last points in VertexList are not the same, the path
    @@ -344,7 +336,7 @@
            CurrentEdge = CurrentEdge->NextEdge;
         }
      }
    -
    +


    @@ -363,10 +355,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/40-03.html b/40-03.html index 5ab819c..a6421c0 100644 --- a/40-03.html +++ b/40-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - @@ -37,13 +30,13 @@


    -

    Complex Polygon Filling: An Implementation

    +

    Complex Polygon Filling: An Implementation

    Listing 40.1 just shown presents a function, FillPolygon(), that fills polygons of all shapes. If CONVEX_FILL_LINKED is defined, the fast convex fill code from Chapter 39 is linked in and used to draw convex polygons. Otherwise, convex polygons are handled as if they were complex. Nonconvex polygons are also handled as complex, although this is not necessary, as discussed shortly.

    Listing 40.1 is a faithful implementation of the complex polygon filling approach just described, with separate functions corresponding to each of the tasks, such as building the GET and X-sorting the AET. Listing 40.2 provides the actual drawing code used to fill spans, built on a draw pixel routine that is the only hardware dependency anywhere in the C code. Listing 40.3 is the header file for the polygon filling code; note that it is an expanded version of the header file used by the fast convex polygon fill code from Chapter 39. (They may have the same name but are not the same file!) Listing 40.4 is a sample program that, when linked to Listings 40.1 and 40.2, demonstrates drawing polygons of various sorts.

    -

    LISTING 40.2 L40-2.C

    +

    LISTING 40.2 L40-2.C

      /* Draws all pixels in the horizontal line segment passed in, from
         (LeftX,Y) to (RightX,Y), in the specified color in mode 13h, the
    @@ -79,9 +72,9 @@
      #endif
         *ScreenPtr = (unsigned char) Color;
      }
    -
    + -

    LISTING 40.3 POLYGON.H

    +

    LISTING 40.3 POLYGON.H

      /* POLYGON.H: Header file for polygon-filling code */
     
    @@ -117,9 +110,9 @@
         int YStart;                /* Y coordinate of topmost line */
         struct HLine * HLinePtr;   /* pointer to list of horz lines */
      };
    -
    + -

    LISTING 40.4 L40-4.C

    +

    LISTING 40.4 L40-4.C

      /* Sample program to exercise the polygon-filling routines */
     
    @@ -200,7 +193,7 @@
         int86(0x10, &regset, &regset);
      }
     
    -
    +


    @@ -219,10 +212,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/40-04.html b/40-04.html index ad8a9ea..9f70b42 100644 --- a/40-04.html +++ b/40-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - @@ -43,15 +36,15 @@

    By the way, I have not seen polygon boundary filling handled precisely this way elsewhere. The boundary filling approach in Foley and van Dam is similar, but seems to me to not draw all boundary and vertex pixels once and only once.

    -

    More on Active Edges

    +

    More on Active Edges

    Edges of zero height—horizontal edges and edges defined by two vertices at the same location—never even make it into the GET in Listing 40.1. A polygon edge of zero height can never be an active edge, because it can never intersect a scan line; it can only run along the scan line, and the span it runs along is defined not by that edge but by the edges that connect to its endpoints.

    -

    Performance Considerations

    +

    Performance Considerations

    How fast is Listing 40.1? When drawing triangles on a 20-MHz 386, it’s less than one-fifth the speed of the fast convex polygon fill code. However, most of that time is spent drawing individual pixels; when Listing 40.2 is replaced with the fast assembly line segment drawing code in Listing 40.5, performance improves by two and one-half times, to about half as fast as the fast convex fill code. Even after conversion to assembly in Listing 40.5, DrawHorizontalLineSeg still takes more than half of the total execution time, and the remaining time is spread out fairly evenly over the various subroutines in Listing 40.1. Consequently, there’s no single place in which it’s possible to greatly improve performance, and the maximum additional improvement that’s possible looks to be a good deal less than two times; for that reason, and because of space limitations, I’m not going to convert the rest of the code to assembly. However, when filling a polygon with a great many edges, and especially one with a great many active edges at one time, relatively more time would be spent traversing the linked lists. In such a case, conversion to assembly (which does a very good job with linked list processing) could pay off reasonably well.

    -

    LISTING 40.5 L40-5.ASM

    +

    LISTING 40.5 L40-5.ASM

      ; Draws all pixels in the horizontal line segment passed in, from
      ;  (LeftX,Y) to (RightX,Y), in the specified color in mode 13h, the
    @@ -103,7 +96,7 @@
      _DrawHorizontalLineSeg  endp
              end
     
    -
    +

    The algorithm used to X-sort the AET is an interesting performance consideration. Listing 40.1 uses a bubble sort, usually a poor choice for performance. However, bubble sorts perform well when the data are already almost sorted, and because of the X coherence of edges from one scan line to the next, that’s generally the case with the AET. An insertion sort might be somewhat faster, depending on the state of the AET when any particular sort occurs, but a bubble sort will generally do just fine.

    @@ -124,10 +117,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/40-05.html b/40-05.html index 83a65e3..a0cdd7f 100644 --- a/40-05.html +++ b/40-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Of Songs, Taxes, and the Simplicity of Complex Polygons - - @@ -47,11 +40,11 @@ -

    Nonconvex Polygons

    +

    Nonconvex Polygons

    Nonconvex polygons can be filled somewhat faster than complex polygons. Because edges never cross or switch positions with other edges once they’re in the AET, the AET for a nonconvex polygon needs to be sorted only when new edges are added. In order for this to work, though, edges must be added to the AET in strict left-to-right order. Complications arise when dealing with two edges that start at the same point, because slopes must be compared to determine which edge is leftmost. This is certainly doable, but because of space limitations and limited performance returns, I haven’t implemented this in Listing 40.1.

    -

    Details, Details

    +

    Details, Details

    Every so often, a programming demon that I’d thought I’d forever laid to rest arises to haunt me once again. A minor example of this—an imp, if you will—is the use of “ = ” when I mean “ == ,” which I’ve done all too often in the past, and am sure I’ll do again. That’s minor deviltry, though, compared to the considerably greater evils of one of my personal scourges, of which I was recently reminded anew: too-close attention to detail. Not seeing the forest for the trees. Looking low when I should have looked high. Missing the big picture, if you catch my drift.

    @@ -84,10 +77,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/41-01.html b/41-01.html index 3a7d473..3c94c96 100644 --- a/41-01.html +++ b/41-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - @@ -37,10 +30,10 @@


    -

    Chapter 41
    +

    Chapter 41
    Those Way-Down Polygon Nomenclature Blues

    -

    Names Do Matter when You Conceptualize a Data Structure

    +

    Names Do Matter when You Conceptualize a Data Structure

    After I wrote the columns on polygons in Dr. Dobb’s Journal that became Chapters 38-40, long-time reader Bill Huber wrote to take me to task—and a well-deserved kick in the fanny it was, I might add—for my use of non-standard polygon terminology in those columns. Unix’s X-Window System (XWS) defines three categories of polygons: complex, nonconvex, and convex. These three categories, each a specialized subset of the preceding category, not-so-coincidentally map quite nicely to three increasingly fast polygon filling techniques. Therefore, I used the XWS names to describe the sorts of polygons that can be drawn with each of the polygon filling techniques.

    @@ -58,7 +51,7 @@

    Or, as Bill Huber put it, “You are free to adopt your own terminology when it suits your purposes well. But you risk losing or confusing those who could be among your most astute readers—those who already have been trained in the same or a related field.” Ditto. Likewise. D’accord. And mea culpa ; I shall endeavor to watch my language in the future.

    -

    Nomenclature in Action

    +

    Nomenclature in Action

    Just to show you how much difference proper description and interchange of ideas can make, consider the case of identifying convex polygons. When I was writing about polygons in my column in DDJ, a nonfunctional method for identifying such polygons—checking for exactly two X direction changes and two Y direction changes around the perimeter of the polygon—crept into the column by accident. That method, as I noted in a later column, does not work. (That’s why you won’t find it in this book.) Still, a fast method of checking for convex polygons would be highly desirable, because such polygons can be drawn with the fast code from Chapter 39, rather than the relatively slow, general-purpose code from Chapter 40.

    @@ -66,7 +59,7 @@

    What we have is an approach passed along by Jim Kent, of Autodesk Animator fame. If we modify the low-level code to check which edge is left-most on each scan line and start drawing there, as just described, then we can handle any polygon that’s monotone with respect to a vertical line regardless of whether the edges cross. (I’ll call this “monotone-vertical” from now on; if anyone wants to correct that terminology, jump right in.) In other words, we can then handle nonsimple polygons that are monotone-vertical; self-intersection is no longer a problem. We just scan around the polygon’s perimeter looking for exactly two direction reversals along the Y axis only, and if that proves to be the case, we can handle the polygon at high speed. Figure 41.1 shows polygons that can be drawn by a monotone-vertical capable filler; Figure 41.2 shows some that cannot. Listing 41.1 shows code to test whether a polygon is appropriately monotone.

    -

    LISTING 41.1 L41-1.C

    +

    LISTING 41.1 L41-1.C

     /* Returns 1 if polygon described by passed-in vertex list is monotone with
     respect to a vertical line, 0 otherwise. Doesn’t matter if polygon is simple 
    @@ -110,7 +103,7 @@ int PolygonIsMonotoneVertical(struct PointListHeader * VertexList)
        } while (i++ < (Length-1));
        return(1);  /* it’s a vertical-monotone polygon */
     }
    -
    +


    @@ -129,10 +122,6 @@ int PolygonIsMonotoneVertical(struct PointListHeader * VertexList)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/41-02.html b/41-02.html index f4a1e75..519d7f7 100644 --- a/41-02.html +++ b/41-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - @@ -39,15 +32,13 @@

    Listings 41.2 and 41.3 are variants of the fast convex polygon fill code from Chapter 39, modified to be able to handle all monotone-vertical polygons, including nonsimple ones; the edge-scanning code (Listing 39.4 from Chapter 39) remains the same, and so is not shown again here.

    -


    - Figure 41.1
      Monotone-vertical polygons.

    +


    + Figure 41.1
      Monotone-vertical polygons.

    -


    - Figure 41.2
      Non-monotone-vertical polygons.

    +


    + Figure 41.2
      Non-monotone-vertical polygons.

    -

    LISTING 41.2 L41-2.C

    +

    LISTING 41.2 L41-2.C

     /* Color-fills a convex polygon. All vertices are offset by (XOffset, YOffset).
     “Convex” means “monotone with respect to a vertical line”; that is, every 
    @@ -164,7 +155,7 @@ int FillMonotoneVerticalPolygon(struct PointListHeader * VertexList,
        free(WorkingHLineList.HLinePtr);
        return(1);
     }
    -
    +


    @@ -183,10 +174,6 @@ int FillMonotoneVerticalPolygon(struct PointListHeader * VertexList,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/41-03.html b/41-03.html index ebac9bc..ba6596a 100644 --- a/41-03.html +++ b/41-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - @@ -37,7 +30,7 @@


    -

    LISTING 41.3 L41-3.ASM

    +

    LISTING 41.3 L41-3.ASM

     ; Draws all pixels in list of horizontal lines passed in, in mode 13h, VGA’s 
     ; 320x200 256-color mode. Uses REP STOS to fill each line.
    @@ -129,7 +122,7 @@ FillDone:
             ret
     _DrawHorizontalLineList endp
             end
    -
    +

    Listing 41.4 is almost identical to Listing 40.1 from Chapter 40. I’ve modified Listing 40.1 to employ the vertical-monotone detection test we’ve been talking about and use the fast vertical-monotone drawing code whenever possible; that’s what Listing 41.4 is. Note well that Listing 40.5 from Chapter 40 is also required in order for this code to link. Listing 41.5 is an appropriately updated version of the POLYGON.H header file.

    @@ -150,10 +143,6 @@ _DrawHorizontalLineList endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/41-04.html b/41-04.html index 50ad270..6851fc7 100644 --- a/41-04.html +++ b/41-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Those Way-Down Polygon Nomenclature Blues - - @@ -37,7 +30,7 @@


    -

    LISTING 41.4 L41-4.C

    +

    LISTING 41.4 L41-4.C

     /* Color-fills an arbitrarily-shaped polygon described by VertexList.
     If the first and last points in VertexList are not the same, the path
    @@ -323,9 +316,9 @@ static void ScanOutAET(int YToScan, int Color) {
           CurrentEdge = CurrentEdge->NextEdge;
        }
     }
    -
    + -

    LISTING 41.5 POLYGON.H

    +

    LISTING 41.5 POLYGON.H

     /* Header file for polygon-filling code */
     
    @@ -364,7 +357,7 @@ struct HLineList {
     
     /* Describes a color as an RGB triple, plus one byte for other info */
     struct RGB { unsigned char Red, Green, Blue, Spare; };
    -
    +

    Is monotone-vertical polygon detection worth all this trouble? Under the right circumstances, you bet. In a situation where a great many polygons are being drawn, and the application either doesn’t know whether they’re monotone-vertical or has no way to tell the polygon filler that they are, performance can be increased considerably if most polygons are, in fact, monotone-vertical. This potential performance advantage is helped along by the surprising fact that Jim’s test for monotone-vertical status is simpler and faster than my original, nonfunctional test for convexity.

    @@ -387,10 +380,6 @@ struct RGB { unsigned char Red, Green, Blue, Spare; };
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/42-01.html b/42-01.html index 012fd4c..a4d73f3 100644 --- a/42-01.html +++ b/42-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - @@ -37,10 +30,10 @@


    -

    Chapter 42
    +

    Chapter 42
    Wu’ed in Haste; Fried, Stewed at Leisure

    -

    Fast Antialiased Lines Using Wu’s Algorithm

    +

    Fast Antialiased Lines Using Wu’s Algorithm

    The thought first popped into my head as I unenthusiastically picked through the salad bar at a local “family” restaurant, trying to decide whether the meatballs, the fried clams, or the lasagna was likely to shorten my life the least. I decided on the chicken in mystery sauce.

    @@ -58,7 +51,7 @@

    Literature that’s applicable to fast PC graphics is hard enough to find, but what we’d really like is above-average image quality combined with terrific speed, and there’s almost no literature of that sort around. There is some, however, and you folks are right on top of it. For example, alert reader Michael Chaplin, of San Diego, wrote to suggest that I might enjoy the line-antialiasing algorithm presented in Xiaolin Wu’s article, “An Efficient Antialiasing Technique,” in the July 1991 issue of Computer Graphics. Michael was dead-on right. This is a great algorithm, combining excellent antialiased line quality with speed that’s close to that of non-antialiased Bresenham’s line drawing. This is the sort of algorithm that makes you want to go out and write a wire-frame animation program, just so you can see how good those smooth lines look in motion. Wu antialiasing is a wonderful example of what can be accomplished on inexpensive, mass-market hardware with the proper programming perspective. In short, it’s a splendid example of appropriate technology for PCs.

    -

    Wu Antialiasing

    +

    Wu Antialiasing

    Antialiasing, as we’ve been discussing for the past few chapters, is the process of smoothing lines and edges so that they appear less jagged. Antialiasing is partly an aesthetic issue, because it makes images more attractive. It’s also partly an accuracy issue, because it makes it possible to position and draw images with effectively more precision than the resolution of the display. Finally, it’s partly a flat-out necessity, to avoid the horrible, crawling, jagged edges of temporal aliasing when performing animation.

    @@ -66,9 +59,8 @@

    The intensities of the two pixels that bracket the line are selected so that they always sum to exactly 1; that is, to the intensity of one fully illuminated pixel of the drawing color. The presence of aggregate full-pixel intensity means that at each step, the line has the same brightness it would have if a single pixel were drawn at precisely the correct location. Moreover, thanks to the distribution of the intensity weighting, that brightness is centered at the ideal line. Not coincidentally, a line drawn with pixel pairs of aggregate single-pixel intensity, centered on the ideal line, is perceived by the eye not as a jagged collection of pixel pairs, but as a smooth line centered on the ideal line. Thus, by weighting the bracketing pixels properly at each step, we can readily produce what looks like a smooth line at precisely the right location, rather than the jagged pattern of line segments that non-antialiased line-drawing algorithms such as Bresenham’s (see Chapters 35, 36, and 37) trace out.

    -


    - Figure 42.1
      The basic concept of Wu antialiasing.

    +


    + Figure 42.1
      The basic concept of Wu antialiasing.

    You might expect that the implementation of Wu antialiasing would fall into two distinct areas: tracing out the line (that is, finding the appropriate pixel pairs to draw) and calculating the appropriate weightings for each pixel pair. Not so, however. The weighting calculations involve only a few shifts, XORs, and adds; for all practical purposes, tracing and weighting are rolled into one step—and a very fast step it is. How fast is it? On a 33-MHz 486 with a fast VGA, a good but not maxed-out assembly implementation of Wu antialiasing draws a more than respectable 5,000 150-pixel-long vectors per second. That’s especially impressive considering that about 1,500,000 actual pixels are drawn per second, meaning that Wu antialiasing is drawing at around 50 percent of the maximum memory bandwidth—half the fastest theoretically possible drawing speed—of an AT-bus VGA. In short, Wu antialiasing is about as fast an antialiased line approach as you could ever hope to find for the VGA.

    @@ -89,10 +81,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/42-02.html b/42-02.html index 3bd187e..d2387e8 100644 --- a/42-02.html +++ b/42-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - @@ -37,7 +30,7 @@


    -

    Tracing and Intensity in One

    +

    Tracing and Intensity in One

    Horizontal, vertical, and diagonal lines do not require Wu antialiasing because they pass through the center of every pixel they meet; such lines can be drawn with fast, special-case code. For all other cases, Wu lines are traced out one step at a time along the major axis by means of a simple, fixed-point algorithm. The move along the minor axis with respect to a one-pixel move along the major axis (the line slope for lines with slopes less than 1, 1/slope for lines with slopes greater than 1) is calculated with a single integer divide. This value, called the “error adjust,” is stored as a fixed-point fraction, in 0.16 format (that is, all bits are fractional, and the decimal point is just to the left of bit 15). An error accumulator, also in 0.16 format, is initialized to 0. Then the first pixel is drawn; no weighting is needed, because the line intersects its endpoints exactly.

    @@ -45,13 +38,12 @@

    So far, nothing special; but now we come to the true wonder of Wu antialiasing. We know which pair of pixels to draw at each step along the line, but we also need to generate the two proper intensities, which must be inversely proportional to distance from the ideal line and sum to 1, and that’s a potentially time-consuming operation. Let’s assume, however, that the number of possible intensity levels to be used for weighting is the value NumLevels = 2n for some integer n, with the minimum weighting (0 percent intensity) being the value 2n -1, and the maximum weighting (100 percent intensity) being the value 0. Given that, lo and behold, the most significant n bits of the error accumulator select the proper intensity value for one element of the pixel pair, as shown in Figure 42.2. Better yet, 2n-1 minus the intensity of the first pixel selects the intensity of the other pixel in the pair, because the intensities of the two pixels must sum to 1; as it happens, this result can be obtained simply by flipping the n least-significant bits of the first pixel’s value. All this works because what the error accumulator accumulates is precisely the ideal line’s current distance between the two bracketing pixels.

    -


    - Figure 42.2
      Wu intensity calculations.

    +


    + Figure 42.2
      Wu intensity calculations.

    The intensity calculations take longer to describe than they do to perform. All that’s involved is a shift of the error accumulator to right-justify the desired intensity weighting bits, and then an XOR to flip the least-significant n bits of the first pixel’s value in order to generate the second pixel’s value. Listing 42.1 illustrates just how efficient Wu antialiasing is; the intensity calculations take only three statements, and the entire Wu line-drawing loop is only nine statements long. Of course, a single C statement can hide a great deal of complexity, but Listing 42.6, an assembly implementation, shows that only 15 instructions are required per step along the major axis—and the number of instructions could be reduced to ten by special-casing and loop unrolling. Make no mistake about it, Wu antialiasing is fast.

    -

    LISTING 42.1 L42-1.C

    +

    LISTING 42.1 L42-1.C

     /* Function to draw an antialiased line from (X0,Y0) to (X1,Y1), using an
      * antialiasing approach published by Xiaolin Wu in the July 1991 issue of
    @@ -187,7 +179,7 @@ void DrawWuLine(int X0, int Y0, int X1, int Y1, int BaseColor, int NumLevels,
           and so needs no weighting */
        DrawPixel(X1, Y1, BaseColor);
     }
    -
    +


    @@ -206,10 +198,6 @@ void DrawWuLine(int X0, int Y0, int X1, int Y1, int BaseColor, int NumLevels,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/42-03.html b/42-03.html index 8526481..a0bd485 100644 --- a/42-03.html +++ b/42-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - @@ -37,11 +30,11 @@


    -

    Sample Wu Antialiasing

    +

    Sample Wu Antialiasing

    The true test of any antialiasing technique is how good it looks, so let’s have a look at Wu antialiasing in action. Listing 42.1 is a C implementation of Wu antialiasing. Listing 42.2 is a sample program that draws a variety of Wu-antialiased lines, followed by non-antialiased lines, for comparison. Listing 42.3 contains DrawPixel() and SetMode() functions for mode 13H, the VGA’s 320x200 256-color mode. Finally, Listing 42.4 is a simple, non-antialiased line-drawing routine. Link these four listings together and run the resulting program to see both Wu-antialiased and non-antialiased lines.

    -

    LISTING 42.2 L42-2.C

    +

    LISTING 42.2 L42-2.C

     /* Sample line-drawing program to demonstrate Wu antialiasing. Also draws
      * non-antialiased lines for comparison.
    @@ -168,9 +161,9 @@ void SetPalette(struct WuColor * WColors)
           int86x(0x10, &regset, &regset, &sregset); /* load the palette block */
        }
     }
    -
    + -

    LISTING 42.3 L42-3.C

    +

    LISTING 42.3 L42-3.C

     /* VGA mode 13h pixel-drawing and mode set functions.
      * Tested with Borland C++ in C compilation mode and the small model.
    @@ -201,7 +194,7 @@ void SetMode()
        regset.x.ax = 0x0013;
        int86(0x10, &regset, &regset);
     }
    -
    +


    @@ -220,10 +213,6 @@ void SetMode()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/42-04.html b/42-04.html index 4d85582..e4a96cf 100644 --- a/42-04.html +++ b/42-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - @@ -37,7 +30,7 @@


    -

    LISTING 42.4 L42-4.C

    +

    LISTING 42.4 L42-4.C

     /* Function to draw a non-antialiased line from (X0,Y0) to (X1,Y1), using a
      * simple fixed-point error accumulation approach.
    @@ -105,11 +98,11 @@ void DrawLine(int X0, int Y0, int X1, int Y1, int Color)
           DrawPixel(X0, Y0, Color);
        } while (--DeltaX);
     }
    -
    +

    Listing 42.1 isn’t particularly fast, because it calls DrawPixel() for each pixel. On the other hand, DrawPixel() makes it easy to try out Wu antialiasing in a variety of modes; just adapt the code in Listing 42.3 for the 256-color mode you want to support. For example, Listing 42.5 shows code to draw Wu-antialiased lines in 640x480 256-color mode on SuperVGAs built around the Tseng Labs ET4000 chip with at least 512K of display memory installed. It’s well worth checking out Wu antialiasing at 640x480. Although antialiased lines look much smoother than normal lines at 320x200 resolution, they’re far from perfect, because the pixels are so big that the eye can’t blend them properly. At 640x480, however, Wu-antialiased lines look fabulous; from a couple of feet away, they look as straight and smooth as if they were drawn with a ruler.

    -

    LISTING 42.5 L42-5.C

    +

    LISTING 42.5 L42-5.C

     /* Mode set and pixel-drawing functions for the 640x480 256-color mode of
      * Tseng Labs ET4000-based SuperVGAs.
    @@ -151,7 +144,7 @@ void SetMode()
        regset.x.ax = 0x002E;
        int86(0x10, &regset, &regset);
     }
    -
    +

    Listing 42.1 requires that the DAC palette be set up so that a NumLevel-long block of palette entries contains linearly decreasing intensities of the drawing color. The size of the block is programmable, but must be a power of two. The more intensity levels, the better. Wu says that 32 intensities are enough; on my system, eight and even four levels looked pretty good. I found that gamma correction, which gives linearly spaced intensity steps, improved antialiasing quality significantly. Fortunately, we can program the palette with gamma-corrected values, so our drawing code doesn’t have to do any extra work.

    @@ -174,10 +167,6 @@ void SetMode()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/42-05.html b/42-05.html index 2cb8ebc..8635d1e 100644 --- a/42-05.html +++ b/42-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Wu'ed in Haste; Fried, Stewed at Leisure - - @@ -37,7 +30,7 @@


    -

    LISTING 42.6 L42-6.ASM

    +

    LISTING 42.6 L42-6.ASM

     ; C near-callable function to draw an antialiased line from
     ; (X0,Y0) to (X1,Y1), in mode 13h, the VGA's standard 320x200 256-color
    @@ -308,9 +301,9 @@ Done:                           ;we're done with this line
             ret                     ;done
     _DrawWuLine endp
             end
    -
    + -

    Notes on Wu Antialiasing

    +

    Notes on Wu Antialiasing

    Wu antialiasing can be applied to any curve for which it’s possible to calculate at each step the positions and intensities of two bracketing pixels, although the implementation will generally be nowhere near as efficient as it is for lines. However, Wu’s article in Computer Graphics does describe an efficient algorithm for drawing antialiased circles. Wu also describes a technique for antialiasing solids, such as filled circles and polygons. Wu’s approach biases the edges of filled objects outward. Although this is no good for adjacent polygons of the sort used in rendering, it’s certainly possible to design a more accurate polygon-antialiasing approach around Wu’s basic weighting technique. The results would not be quite so good as more sophisticated antialiasing techniques, but they would be much faster.

    @@ -347,10 +340,6 @@ _DrawWuLine endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/43-01.html b/43-01.html index 4cf7784..ad9e71a 100644 --- a/43-01.html +++ b/43-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - @@ -37,10 +30,10 @@


    -

    Chapter 43
    +

    Chapter 43
    Bit-Plane Animation

    -

    A Simple and Extremely Fast Animation Method for Limited Color

    +

    A Simple and Extremely Fast Animation Method for Limited Color

    When it comes to computers, my first love is animation. There’s nothing quite like the satisfaction of fooling the eye and creating a miniature reality simply by rearranging a few bytes of display memory. What makes animation particularly interesting is that it has to happen fast (as measured in human time), and without blinking and flickering, or else you risk destroying the illusion of motion and solidity. Those constraints make animation the toughest graphics challenge—and also the most rewarding.

    @@ -56,7 +49,7 @@

    It doesn’t much matter if bit-plane animation isn’t perfect for all applications, though. The real point of showing you bit-plane animation is to bring home the reality that the VGA is a complex adapter with many resources, and that you can do remarkable things if you understand those resources and come up with creative ways to put them to work at specific tasks.

    -

    Bit-Planes: The Basics

    +

    Bit-Planes: The Basics

    The underlying principle of bit-plane animation is extremely simple. The VGA has four separate bit planes in modes 0DH, 0EH, 10H, and 12H. Plane 0 normally contains data for the blue component of pixel color, plane 1 normally contains green pixel data, plane 2 red pixel data, and plane 3 intensity pixel data—but we’re going to mix that up a bit in a moment, so we’ll simply refer to them as planes 0, 1, 2, and 3 from now on.

    @@ -64,21 +57,18 @@

    Take a good look at Figure 43.1. Any light bulbs going on over your head yet? If not, consider this. The general problem with VGA animation is that it’s complex and time-consuming to manipulate images that span the four planes (as most do), and that it’s hard to avoid interference problems when images intersect, since those images share the same bits in display memory. Since the four bit planes can be written to and read from independently, it should be apparent that if we could come up with a way to display images from each plane independently of whatever images are stored in the other planes, we would have four sets of images that we could manipulate very easily. There would be no interference effects between images in different planes, because images in one plane wouldn’t share bits with images in another plane. What’s more, since all the bits for a given image would reside in a single plane, we could do away with the cumbersome programming of the VGA’s complex hardware that is needed to manipulate images that span multiple planes.

    -


    - Figure 43.1
      How 4 bits of video data become 6 bits of color.

    +


    + Figure 43.1
      How 4 bits of video data become 6 bits of color.

    All in all, it would be a good deal if we could store each image in a single plane, as shown in Figure 43.2. However, a problem arises when images in different planes overlap, as shown in Figure 43.3. The combined bits from overlapping images generate new colors, so the overlapping parts of the images don’t look like they belong to either of the two images. What we really want, of course, is for one of the images to appear to be in front of the other. It would be better yet if the rearward image showed through any transparent (that is, background-colored) parts of the forward image. Can we do that?

    You bet.

    -


    - Figure 43.2
      Storing images in separate planes.

    +


    + Figure 43.2
      Storing images in separate planes.

    -


    - Figure 43.3
      The problem of overlapping colors.

    +


    + Figure 43.3
      The problem of overlapping colors.


    @@ -97,10 +87,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/43-02.html b/43-02.html index 31e8bf7..b0ecb22 100644 --- a/43-02.html +++ b/43-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - @@ -37,7 +30,7 @@


    -

    Stacking the Palette Registers

    +

    Stacking the Palette Registers

    Suppose that instead of viewing the four bits per pixel coming out of display memory as selecting one of sixteen colors,we view those bits as selecting one of four colors. If the bit from plane 0 is 1, that would select color 0 (say, red). The bit from plane 1 would select color 1 (say, green), the bit from plane 2 would select color 2 (say, blue), and the bit from plane 3 would select color 3 (say, white). Whenever more than 1 bit is 1, the 1 bit from the lowest-numbered plane would determine the color, and 1 bits from all other planes would be ignored. Finally, the absence of any 1 bits at all would select the background color (say, black).

    @@ -214,13 +207,12 @@ -


    - Figure 43.4
      How pixel precedence works.

    +


    + Figure 43.4
      How pixel precedence works.

    Seems almost too easy, doesn’t it? Nonetheless, it works beautifully, as we’ll see very shortly. First, though, I’d like to point out that there’s nothing sacred about plane 0 having precedence. We could rearrange the palette register settings so that any plane had the highest precedence, followed by the other planes in any order. I’ve chosen to make plane 0 the highest precedence only because it seems simplest to think of plane 0 as appearing in front of plane 1, which is in front of plane 2, which is in front of plane 3.

    -

    Bit-Plane Animation in Action

    +

    Bit-Plane Animation in Action

    Without further ado, Listing 43.1 shows bit-plane animation in action. Listing 43.1 animates 13 rather large images (each 32 pixels on a side) over a complex background at a good clip even on a primordial 8088-based PC. Five of the images move very quickly, while the other 8 bounce back and forth at a steady pace.

    @@ -241,10 +233,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/43-03.html b/43-03.html index 56de638..c282a7f 100644 --- a/43-03.html +++ b/43-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - @@ -37,7 +30,7 @@


    -

    LISTING 43.1 L43-1.ASM

    +

    LISTING 43.1 L43-1.ASM

     ; Program to demonstrate bit-plane animation. Performs
     ; flicker-free animation with image transparency and
    @@ -534,7 +527,7 @@ DrawObjectendp
     ;
     Code    ends
             end     Start
    -
    +


    @@ -553,10 +546,6 @@ Code ends
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/43-04.html b/43-04.html index b9ebc85..f12d292 100644 --- a/43-04.html +++ b/43-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - @@ -51,7 +44,7 @@

    The addition of code to support rotated images would also open the door to support for internal animation, where the appearance of a given image changes over time to suggest that the image is an active entity. For example, propellers could whirl, jaws could snap, and jets could flare. Bit-plane animation with bit-aligned images and internal animation can look truly spectacular. It’s a sight worth seeing, particularly for those who doubt the PC’s worth when it comes to animation.

    -

    Limitations of Bit-Plane Animation

    +

    Limitations of Bit-Plane Animation

    As I’ve said, bit-plane animation is not perfect. For starters, bit-plane animation can only be used in the VGA’s planar modes, modes 0DH, 0EH, 10H, and 12H. Also, the reprogramming of the palette registers that provides image precedence also reduces the available color set from the normal 16 colors to just 5 (one color per plane plus the background color). Worse still, each image must consist entirely of only one of the four colors. Mixing colors within an image is not allowed, since the bits for each image are limited to a single plane and can therefore select only one color. Finally, all images of the same precedence must be the same color.

    @@ -76,10 +69,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/43-05.html b/43-05.html index f465587..bcce8b4 100644 --- a/43-05.html +++ b/43-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - @@ -175,13 +168,12 @@

    Not allowing images in the same plane to overlap is actually less of a limitation than it seems. Run Listing 43.1 again. Unless you were looking for it, you’d never notice that images of the same color almost never overlap—there’s plenty of action to distract the eye, and the trajectories of images of the same color are arranged so that they have a full range of motion without running into each other. The only exception is the chain of green images, which occasionally doubles back on itself when it bounces directly into a corner and reverses direction. Here, however, the images are moving so quickly that the brief moment during which one image’s fringe blanks a portion of another image is noticeable only upon close inspection, and not particularly unaesthetic even then.

    -


    - Figure 43.5
      Pixel precedence for plane 3 only.

    +


    + Figure 43.5
      Pixel precedence for plane 3 only.

    When a technique has such tremendous visual and performance advantages as does bit-plane animation, it behooves you to design your animation software so that the limitations of the animation technique don’t get in the way. For example, you might design a shooting gallery game with all the images in a given plane marching along in step in a continuous band. The images could never overlap, so bit-plane animation would produce very high image quality.

    -

    Shearing and Page Flipping

    +

    Shearing and Page Flipping

    As Listing 43.1 runs, you may occasionally see an image shear, with the top and bottom parts of the image briefly offset. This is a consequence of drawing an image directly into memory as that memory is being scanned for video data. Occasionally the CRT controller scans a given area of display memory for pixel data just as the program is changing that same memory. If the CRT controller scans memory faster than the CPU can modify that memory, then the CRT controller can scan out the bytes of display memory that have been already been changed, pass the point in the image that the CPU is currently drawing, and start scanning out bytes that haven’t yet been changed. The result: Mismatched upper and lower portions of the image.

    @@ -202,10 +194,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/43-06.html b/43-06.html index 901f7fc..dc912c0 100644 --- a/43-06.html +++ b/43-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Bit-Plane Animation - - @@ -51,7 +44,7 @@

    To sum up, bit-plane animation by itself is very fast and looks good. In conjunction with page flipping, bit-plane animation looks a little better but is slower, and the overall animation scheme is more difficult to implement and perhaps a bit less reliable on some computers.

    -

    Beating the Odds in the Jaw-Dropping Contest

    +

    Beating the Odds in the Jaw-Dropping Contest

    Bit-plane animation is neat stuff. Heck, good animation of any sort is fun, and the PC is as good a place as any (well, almost any) to make people’s jaws drop. (Certainly it’s the place to go if you want to make a lot of jaws drop.) Don’t let anyone tell you that you can’t do good animation on the PC. You can—if you stretch your mind to find ways to bring the full power of the VGA to bear on your applications. Bit-plane animation isn’t for every task; neither are page flipping, exclusive-ORing, pixel panning, or any of the many other animation techniques you have available. One or more tricks from that grab-bag should give you what you need, though, and the bigger your grab-bag, the better your programs.

    @@ -72,10 +65,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/44-01.html b/44-01.html index 984aa58..c06d018 100644 --- a/44-01.html +++ b/44-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - @@ -37,10 +30,10 @@


    -

    Chapter 44
    +

    Chapter 44
    Split Screens Save the Page Flipped Day

    -

    640x480 Page Flipped Animation in 64K...Almost

    +

    640x480 Page Flipped Animation in 64K...Almost

    Almost doesn’t count, they say—at least in horseshoes and maybe a few other things. This is especially true in digital circles, where if you need 12 MB of hard disk to install something and you only have 10 MB left (a situation that seems to be some sort of eternal law) you’re stuck.

    @@ -50,7 +43,7 @@

    No horseshoes here.

    -

    A Plethora of Challenges

    +

    A Plethora of Challenges

    In its simplest terms, computer animation consists of rapidly redrawing similar images at slightly differing locations, so that the eye interprets the successive images as a single object in motion over time. The fact that the world is an analog realm and the images displayed on a computer screen consist of discrete pixels updated at a maximum rate of about 70 Hz is irrelevant; your eye can interpret both real-world images and pixel patterns on the screen as objects in motion, and that’s that.

    @@ -58,7 +51,7 @@

    Another problem of animation is that the screen must update often enough so that motion appears continuous. A moving object that moves just once every second, shifting by hundreds of pixels each time it does move, will appear to jump, not to move smoothly. Therefore, there are two overriding requirements for smooth animation: 1) the bitmap must be updated quickly (once per frame—60 to 70 Hz—is ideal, although 30 Hz will do fine), and, 2) the process of redrawing the screen must be invisible to the user; only the end result should ever be seen. Both of these requirements are met by the program presented in Listings 44.1 and 44.2.

    -

    A Page Flipping Animation Demonstration

    +

    A Page Flipping Animation Demonstration

    The listings taken together form a sample animation program, in which a single object bounces endlessly off other objects, with instructions and a count of bounces displayed at the bottom of the screen. I’ll discuss various aspects of Listings 44.1 and 44.2 during the balance of this article. The listings are too complex and involve too much VGA and animation knowledge for for me to discuss it all in exhaustive detail (and I’ve covered a lot of this stuff earlier in the book); instead, I’ll cover the major elements, leaving it to you to explore the finer points—and, hope, to experiment with and expand on the code I’ll provide.

    @@ -79,10 +72,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/44-02.html b/44-02.html index 8222439..3463b35 100644 --- a/44-02.html +++ b/44-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - @@ -37,7 +30,7 @@


    -

    LISTING 44.1 L44-1.C

    +

    LISTING 44.1 L44-1.C

     /* Split screen VGA animation program. Performs page flipping in the
     top portion of the screen while displaying non-page flipped
    @@ -326,7 +319,7 @@ void MoveBouncer(bouncer *Bouncer, bumper *BumperPtr, int NumBumpers) {
        Bouncer->LeftX = NewLeftX; /* set the final new coordinates */
        Bouncer->TopY = NewTopY;
     }
    -
    +


    @@ -345,10 +338,6 @@ void MoveBouncer(bouncer *Bouncer, bumper *BumperPtr, int NumBumpers) {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/44-03.html b/44-03.html index bcd6c5a..b5419aa 100644 --- a/44-03.html +++ b/44-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - @@ -37,7 +30,7 @@


    -

    LISTING 44.2 L44-2.ASM

    +

    LISTING 44.2 L44-2.ASM

     ; Low-level animation routines.
     ; Tested with TASM
    @@ -361,7 +354,7 @@ CharUpLoop:
             ret
     -SetBIOS8x8Font endp
             end
    -
    +


    @@ -380,10 +373,6 @@ CharUpLoop:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/44-04.html b/44-04.html index 99b1729..57f355c 100644 --- a/44-04.html +++ b/44-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - @@ -39,7 +32,7 @@

    Listing 44.1 is written in C. It could equally well have been written in assembly language, and would then have been somewhat faster. However, wanted to make the point (as I’ve made again and again) that assembly language, and, indeed, optimization in general, is needed only in the most critical portions of any program, and then only when the program would otherwise be too slow. Only in a highly performance-sensitive situation would the performance boost resulting from converting Listing 44.1 to assembly justify the time spent in coding and the bugs that would likely creep in—and the sample program already updates the screen at the maximum possible rate of once per frame even on a 1985-vintage 8-MHz AT. In this case, faster performance would result only in a longer wait for the page to flip.

    -

    Write Mode 3

    +

    Write Mode 3

    It’s possible to update the bitmap very efficiently on the VGA, because the VGA can draw up to 8 pixels at once, and because the VGA provides a number of hardware features to speed up drawing. This article makes considerable use of one particularly unusual hardware feature, write mode 3. We discussed write mode 3 back in Chapter 26, but we’ve covered a lot of ground since then—so I’m going to run through a quick refresher on write mode 3.

    @@ -49,32 +42,32 @@

    Write mode 3 is useful when you want to set some but not all of the pixels in a single byte of display memory to the same color. That is, if you want to draw a number of pixels within a byte in a single color, write mode 3 is a good way to do it.

    -

    Write mode 3 works like this. First, set the Graphics Controller Mode register to write mode 3. (Look at Listing 44.2 for code that does everything described here.) Next, set the Set/Reset register to the color with which you wish to draw, in the range 0-15. (It is not necessary to explicitly enable set/reset via the Enable Set/Reset register; write mode 3 does that automatically.) Then, to draw individual pixels within a single byte, simply read display memory, and then write a byte to display memory with 1-bits where you want the color to be drawn and 0-bits where you want the current bitmap contents to be preserved. (Note well that the data actually read by the CPU doesn’t matter; the read operation latches all four planes’ data, as described way back in Chapter 24.) So, for example, if write mode 3 is enabled and the Set/Reset register is set to 1 (blue), then the following sequence of operations:

    +

    Write mode 3 works like this. First, set the Graphics Controller Mode register to write mode 3. (Look at Listing 44.2 for code that does everything described here.) Next, set the Set/Reset register to the color with which you wish to draw, in the range 0-15. (It is not necessary to explicitly enable set/reset via the Enable Set/Reset register; write mode 3 does that automatically.) Then, to draw individual pixels within a single byte, simply read display memory, and then write a byte to display memory with 1-bits where you want the color to be drawn and 0-bits where you want the current bitmap contents to be preserved. (Note well that the data actually read by the CPU doesn’t matter; the read operation latches all four planes’ data, as described way back in Chapter 24.) So, for example, if write mode 3 is enabled and the Set/Reset register is set to 1 (blue), then the following sequence of operations:

     mov   dx,0a000h
     mov   es,dx
     mov   al,es:[0]
     mov   byte ptr es:[0],0f0h
    -
    +

    will change the first 4 pixels on the screen (the left nibble of the byte at offset 0 in display memory) to blue, and will leave the next 4 pixels (the right nibble of the byte at offset 0) unchanged.

    -

    Using one MOV to read from display memory and another to write to display memory is not particularly efficient on some processors. In Listing 44.2, I instead use XCHG, which reads and then writes a memory location in a single operation, as in:

    +

    Using one MOV to read from display memory and another to write to display memory is not particularly efficient on some processors. In Listing 44.2, I instead use XCHG, which reads and then writes a memory location in a single operation, as in:

     mov    dx,0a000h
     mov    es,dx
     mov    al,0f0h
     xchg   es:[0],al
    -
    +

    Again, the actual value that’s read is irrelevant. In general, the XCHG approach is more compact than two MOVs, and is faster on 386 and earlier processors, but slower on 486s and Pentiums.

    -

    If all pixels in a byte of display memory are to be drawn in a single color, it’s not necessary to read before writing, because none of the information in display memory at that byte needs to be preserved; a simple write of 0FFH (to draw all bits) will set all 8 pixels to the set/reset color:

    +

    If all pixels in a byte of display memory are to be drawn in a single color, it’s not necessary to read before writing, because none of the information in display memory at that byte needs to be preserved; a simple write of 0FFH (to draw all bits) will set all 8 pixels to the set/reset color:

     mov   dx,0a000h
     mov   es,dx
     mov   byte ptr es:[di],0ffh
    -
    + @@ -86,7 +79,7 @@ mov byte ptr es:[di],0ffh

    In short, write mode 3 is a good choice for single-color drawing that modifies individual pixels within display memory bytes. Not coincidentally, the sample application draws only single-color objects within the animation area; this allows write mode 3 to be used for all drawing, in keeping with our desire for speedy screen updates.

    -

    Drawing Text

    +

    Drawing Text

    We’ll need text in the sample application; is that also a good use for write mode 3? Sometimes it is, but not in this particular case.

    @@ -109,10 +102,6 @@ mov byte ptr es:[di],0ffh
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/44-05.html b/44-05.html index 786bb4e..635dbce 100644 --- a/44-05.html +++ b/44-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - @@ -49,7 +42,7 @@

    I’m not going to delve any deeper into the considerable issues of drawing VGA text; I just want to sensitize you to the existence of approaches other than the ones used in Listings 44.1 and 44.2. On the VGA, the rule is: If there’s something you want to do, there probably are 10 ways to do it, each with unique strengths and weaknesses. Your mission, should you decide to accept it, is to figure out which one is best for your particular application.

    -

    Page Flipping

    +

    Page Flipping

    Now that we know how to update the screen reasonably quickly, it’s time to get on to the fun stuff. Page flipping answers the second requirement for animation, by keeping bitmap changes off the screen until they’re complete. In other words, page flipping guarantees that partially updated bitmaps are never seen.

    @@ -57,9 +50,8 @@

    The VGA bitmap is a linear 64 K block of memory. (True, most adapters nowadays are SuperVGAs with more than 256 K of display memory, but every make of SuperVGA has its own way of letting you access that extra memory, so going beyond standard VGA is a daunting and difficult task. Also, it’s hard to manipulate the large frame buffers of SuperVGA modes fast enough for real-time animation.) Normally, the VGA picks up the first byte of memory (the byte at offset 0) and displays the corresponding 8 pixels on the screen, then picks up the byte at offset 1 and displays the next 8 pixels, and so on to the end of the screen. However, the offset of the first byte of display memory picked up during each frame is not fixed at 0, but is rather programmable by way of the Start Address High and Low registers, which together store the 16-bit offset in display memory at which the bitmap to be displayed during the next frame starts. So, for example, in mode 10H (640x350, 16 colors), a large enough bitmap to store a complete screen of information can be stored at display memory offsets 0 through 27,999, and another full bitmap could be stored at offsets 28,000 through 55,999, as shown in Figure 44.1. (I’m discussing 640x350 mode at the moment for good reason; we’ll get to 640x480 shortly.) When the Start Address registers are set to 0, the first bitmap (or page) is displayed; when they are set to 28,000, the second bitmap is displayed. Page flipped animation can be performed by displaying page 0 and drawing to page 1, then setting the start address to page 1 to display that page and drawing to page 0, and so on ad infinitum.

    -


    - Figure 44.1
      Memory allocation for mode 10h page flipping.

    +


    + Figure 44.1
      Memory allocation for mode 10h page flipping.


    @@ -78,10 +70,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/44-06.html b/44-06.html index 17c4488..fd4b28e 100644 --- a/44-06.html +++ b/44-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Split Screens Save the Page Flipped Day - - @@ -37,7 +30,7 @@


    -

    Knowing When to Flip

    +

    Knowing When to Flip

    There’s a hitch, though, and that hitch is knowing exactly when it is that the page has flipped. The page doesn’t flip the instant that you set the Start Address registers. The VGA loads the starting offset from the Start Address registers once before starting each frame, then pays those registers no nevermind until the next frame comes around. This means that you can set the Start Address registers whenever you want—but the page actually being displayed doesn’t change until after the VGA loads that new offset in preparation for the next frame.

    @@ -49,7 +42,7 @@

    So, to flip pages, you must complete all drawing to the non-displayed page, wait for Display Enable to be active, set the new start address, and wait for Vertical Sync to be active. At that point, you can be fully confident that the page that you just flipped off the screen is not displayed and can safely (invisibly) be updated. A side benefit of page flipping is that your program will automatically have a constant time base, with the rate at which new screens are drawn synchronized to the frame rate of the display (typically 60 or 70 Hz). However, complex updates may take more than one frame to complete, especially on slower processors; this can be compensated for by maintaining a count of new screens drawn and cross-referencing that to the BIOS timer count periodically, accelerating the overall pace of the animation (moving farther each time and the like) if updates are happening too slowly.

    -

    Enter the Split Screen

    +

    Enter the Split Screen

    So far, I’ve discussed page flipping in 640x350 mode. There’s a reason for that: 640x350 is the highest-resolution standard mode in which there’s enough display memory for two full pages on a standard VGA. It’s possible to program the VGA to a non-standard 640x400 mode and still have two full pages, but that’s pretty much the limit. One 640x480 page takes 38,400 bytes of display memory, and clearly there isn’t enough room in 64 K of display memory for two of those monster pages.

    @@ -59,9 +52,8 @@

    That, in turn, allows us to divvy up display memory into three areas, as shown in Figure 44.2. The area from 0 to 11,279 is reserved for the split screen, the area from 11,280 to 38,399 is used for page 0, and the area from 38,400 to 65,519 is used for page 1. This allows page flipping to be performed in the top 339 scan lines (about 70 percent) of the screen, and leaves the bottom 141 scan lines for non-animation purposes, such as showing scores, instructions, statuses, and suchlike. (Note that the allocation of display memory and number of scan lines are dictated by the desire to have as many page-flipped scan lines as possible; you may, if you wish, have fewer page-flipped lines and reserve part of the bitmap for other uses, such as off-screen storage for images.)

    -


    - Figure 44.2
      Memory allocation for mode 12h page flipping.

    +


    + Figure 44.2
      Memory allocation for mode 12h page flipping.

    The sample program for this chapter uses the split screen and page flipping exactly as described above. The playfield through which the object bounces is the page-flipped portion of the screen, and the rectangle at the bottom containing the bounce count and the instructions is the split (that is, not animatable) portion of the screen. Of course, to the user it all looks like one screen. There are no visible boundaries between the two unless you choose to create them.

    @@ -86,10 +78,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/45-01.html b/45-01.html index 4c14440..7483c51 100644 --- a/45-01.html +++ b/45-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - @@ -37,10 +30,10 @@


    -

    Chapter 45
    +

    Chapter 45
    Dog Hair and Dirty Rectangles

    -

    Different Angles on Animation

    +

    Different Angles on Animation

    We brought our pets with us when we moved to Seattle. At about the same time, our Golden Retriever, Sam, observed his third birthday. Sam is relatively intelligent, in the sense that he is clearly smarter than a banana slug, although if he were in the same room with Jeff Duntemann’s dog Mr. Byte, there’s a reasonable chance that he would mistake Mr. Byte for something edible (a category that includes rocks, socks, and a surprising number of things too disgusting to mention), and Jeff would have to find a new source of things to write about.

    @@ -50,7 +43,7 @@

    And then we took Sam to the vet for his annual check-up and found that he had an ear infection. Thanks to the wonders of modern animal medicine, a $5 bottle of liquid restored his health in just two days. And with his health, we got, as a bonus, the old Sam. You see, Sam hadn’t changed. He was just tired from being sick. Now he once again joyously knocks down any stranger who makes the mistake of glancing in his direction, and will, quite possibly, be booked any day now on suspicion of homicide by licking.

    -

    Plus ça Change

    +

    Plus ça Change

    Okay, you give up. What exactly does this have to do with graphics? I’m glad you asked. The lesson to be learned from Sam, The Dog With A Brain The Size Of A Walnut, is that while things may look like they’ve changed, in fact they often haven’t. Take VGA performance. If you buy a 486 with a SuperVGA, you’ll get performance that knocks your socks off, especially if you run Windows. Things are liable to be so fast that you’ll figure the SuperVGA has to deserve some of the credit. Well, maybe it does if it’s a local-bus VGA. But maybe it doesn’t, even if it is local bus—and it certainly doesn’t if it’s an ISA bus VGA, because no ISA bus VGA can run faster than about 300 nanoseconds per access, and VGAs capable of that speed have been common for at least a couple of years now.

    @@ -58,7 +51,7 @@

    So, as I say, sometimes things don’t change. Of course, sometimes they do change. For example, in just 49 dog years, I fully expect to own at least one pair of underwear without a single hole in it. Which brings us, deus ex machina and the creek don’t rise, to yet another animation method: dirty-rectangle animation.

    -

    VGA Access Times

    +

    VGA Access Times

    Actually, before we get to dirty rectangles, I’d like to take you through a quick refresher on VGA memory and I/O access times. I want to do this partly because the slow access times of the VGA make dirty-rectangle animation particularly attractive, and partly as a public service, because even I was shocked by the results of some I/O performance tests I recently ran.

    @@ -83,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/45-02.html b/45-02.html index e2b4e9d..302a319 100644 --- a/45-02.html +++ b/45-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - @@ -165,21 +158,19 @@

    It is indeed a strange concept: The key to fast graphics is staying away from the graphics adapter as much as possible.

    -

    Dirty-Rectangle Animation

    +

    Dirty-Rectangle Animation

    The relative slowness of VGA hardware is part of the appeal of the technique that I call “dirty-rectangle” animation, in which a complete copy of the contents of display memory is maintained in offscreen system (nondisplay) memory. All drawing is done to this system buffer. As offscreen drawing is done, a list is maintained of the bounding rectangles for the drawn-to areas; these are the dirty rectangles, “dirty” in the sense that that have been altered and no longer match the contents of the screen. After all drawing for a frame is completed, all the dirty rectangles for that frame are copied to the screen in a burst, and then the cycle of off-screen drawing begins again.

    Why, exactly, would we want to go through all this complication, rather than simply drawing to the screen in the first place? The reason is visual quality. If we were to do all our drawing directly to the screen, there’d be a lot of flicker as objects were erased and then redrawn. Similarly, overlapped drawing done with the painter’s algorithm (in which farther objects are drawn first, so that nearer objects obscure them) would flicker as farther objects were visible for short periods. With dirty-rectangle animation, only the finished pixels for any given frame ever appear on the screen; intermediate results are never visible. Figure 45.1 illustrates the visual problems associated with drawing directly to the screen; Figure 45.2 shows how dirty-rectangle animation solves these problems.

    -


    - Figure 45.1
      Drawing directly to the screen.

    +


    + Figure 45.1
      Drawing directly to the screen.

    -


    - Figure 45.2
      Dirty rectangle animation.

    +


    + Figure 45.2
      Dirty rectangle animation.

    -

    So Why Not Use Page Flipping?

    +

    So Why Not Use Page Flipping?

    Well, then, if we want good visual quality, why not use page flipping? For one thing, not all adapters and all modes support page flipping. The CGA and MCGA don’t, and neither do the VGA’s 640x480 16-color or 320x200 256-color modes, or many SuperVGA modes. In contrast, all adapters support dirty-rectangle animation. Another advantage of dirty-rectangle animation is that it’s generally faster. While it may seem strange that it would be faster to draw off-screen and then copy the result to the screen, that is often the case, because dirty-rectangle animation usually reduces the number of times the VGA’s hardware needs to be touched, especially in 256-color modes.

    @@ -202,10 +193,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/45-03.html b/45-03.html index a57aeec..9558a51 100644 --- a/45-03.html +++ b/45-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - @@ -39,11 +32,11 @@

    Also, page flipping wastes a good deal of time waiting for the page to flip at the end of the frame. Dirty-rectangle animation never needs to wait for anything because partially drawn images are never present in display memory. Actually, in one sense, partially drawn images are sometimes present because it’s possible for a rectangle to be partially drawn when the scanning raster beam reaches that part of the screen. This causes the rectangle to appear partially drawn for one frame, producing a phenomenon I call “shearing.” Fortunately, shearing tends not to be particularly distracting, especially for fairly small images, but it can be a problem when copying large areas. This is one area in which dirty-rectangle animation falls short of page flipping, because page flipping has perfect display quality, never showing anything other than a completely finished frame. Similarly, dirty-rectangle copying may take two or more frame times to finish, so even if shearing doesn’t happen, it’s still possible to have the images in the various dirty rectangles show up non-simultaneously. In my experience, this latter phenomenon is not a serious problem, but do be aware of it.

    -

    Dirty Rectangles in Action

    +

    Dirty Rectangles in Action

    Listing 45.1 demonstrates dirty-rectangle animation. This is a very simple implementation, in several respects. For one thing, it’s written entirely in C, and animation fairly cries out for assembly language. For another thing, it uses far pointers, which C often handles with less than optimal efficiency, especially because I haven’t used library functions to copy and fill memory. (I did this so the code would work in any memory model.) Also, Listing 45.1 doesn’t attempt to coalesce rectangles so as to perform a minimum number of display-memory accesses; instead, it copies each dirty rectangle to the screen, even if it overlaps with another rectangle, so some pixels are copied multiple times. Listing 45.1 runs pretty well, considering all of its failings; on my 486/33, 10 11x11 images animate at a very respectable clip.

    -

    LISTING 45.1 L45-1.C

    +

    LISTING 45.1 L45-1.C

     /* Sample simple dirty-rectangle animation program. Doesn’t attempt to coalesce
        rectangles to minimize display memory accesses. Not even vaguely optimized!
    @@ -303,7 +296,7 @@ void EraseEntities()
           }
        }
     }
    -
    +


    @@ -322,10 +315,6 @@ void EraseEntities()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/45-04.html b/45-04.html index 6028d35..9bd677a 100644 --- a/45-04.html +++ b/45-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - @@ -49,7 +42,7 @@

    If you keep a close watch, you’ll notice that many high-performance animation games similarly restrict their full-featured animation area to a relatively small region. Often, it’s hard to tell that this is the case, because the animation region is surrounded by flashy digitized graphics and by items such as scoreboards and status screens, but look closely and see if the animation region in your favorite game isn’t smaller than you thought.

    -

    Hi-Res VGA Page Flipping

    +

    Hi-Res VGA Page Flipping

    On a standard VGA, hi-res mode is mode 12H, which offers 640x480 resolution with 16 colors. That’s a nice mode, with plenty of pixels, and square ones at that, but it lacks one thing—page flipping. The problem is that the mode 12H bitmap is 150 K in size, and the standard VGA has only 256 K total, too little memory for two of those monster mode 12H pages. With only one page, flipping is obviously out of the question, and without page flipping, top-flight, hi-res animation can’t be implemented. The standard fallback is to use the EGA’s hi-res mode, mode 10H (640x350, 16 colors) for page flipping, but this mode is less than ideal for a couple of reasons: It offers sharply lower vertical resolution, and it’s lousy for handling scaled-up CGA graphics, because the vertical resolution is a fractional multiple—1.75 times, to be exact—of that of the CGA. CGA resolution may not seem important these days, but many images were originally created for the CGA, as were many graphics packages and games, and it’s at least convenient to be able to handle CGA graphics easily. Then, too, 640x350 is also a poor multiple of the 200 scan lines of the popular 320x200 256-color mode 13H of the VGA.

    @@ -59,7 +52,7 @@

    The key to 640x400 mode is understanding that on a VGA, mode 10H (640x350) is, at heart, a 400-scan-line mode. What I mean by that is that in mode 10H, the Vertical Total register, which controls the total number of scan lines, both displayed and nondisplayed, is set to 447, exactly the same as in the VGA’s text modes, which do in fact support 400 scan lines. A properly sized and centered display is achieved in mode 10H by setting the polarity of the sync pulses to tell the monitor to scan vertically at a faster rate (to make fewer lines fill the screen), by starting the overscan after 350 lines, and by setting the vertical sync and blanking pulses appropriately for the faster vertical scanning rate. Changing those settings is all that’s required to turn mode 10H into a 640x400 mode, and that’s easy to do, as illustrated by Listing 45.2, which provides mode set code for 640x400 mode.

    -

    LISTING 45.2 L45-2.C

    +

    LISTING 45.2 L45-2.C

     /* Mode set routine for VGA 640x400 16-color mode. Tested with
        Borland C++ in C compilation mode. */
    @@ -92,7 +85,7 @@ void Set640x400()
        outpw(0x3D4, 0xB916);   /* adjust the Vertical Blank End register
                                   for 400 scan lines */
     }
    -
    +


    @@ -111,10 +104,6 @@ void Set640x400()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/45-05.html b/45-05.html index b0f036d..1d4db9e 100644 --- a/45-05.html +++ b/45-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - @@ -39,7 +32,7 @@

    In 640x400, 16-color mode, page 0 runs from offset 0 to offset 31,999 (7CFFH), and page 1 runs from offset 32,000 (7D00H) to 63,999 (0F9FFH). Page 1 is selected by programming the Start Address registers (CRTC registers 0CH, the high 8 bits, and 0DH, the low 8 bits) to 7D00H. Actually, because the low byte of the start address is 0 for both pages, you can page flip simply by writing 0 or 7DH to the Start Address High register (CRTC register 0CH); this has the benefit of eliminating a nasty class of potential synchronization bugs that can arise when both registers must be set. Listing 45.3 illustrates simple 640x400 page flipping.

    -

    LISTING 45.3 L45-3.C

    +

    LISTING 45.3 L45-3.C

     /* Sample program to exercise VGA 640x400 16-color mode page flipping, by
        drawing a horizontal line at the top of page 0 and another at bottom of page 1,
    @@ -112,11 +105,11 @@ void Wait30Frames()
           while ((inp(INPUT_STATUS_1) & 0x08) == 0) ;
        }
     }
    -
    +

    After I described 640x400 mode in a magazine article, Bill Lindley, of Mesa, Arizona, wrote me to suggest that when programming the VGA to a nonstandard mode of this sort, it’s a good idea to tell the BIOS about the new screen size, for a couple of reasons. For one thing, pop-up utilities often use the BIOS variables; Bill’s memory-resident screen printer, EGAD Screen Print, determines the number of scan lines to print by multiplying the BIOS “number of text rows” variable times the “character height” variable. For another, the BIOS itself may do a poor job of displaying text if not given proper information; the active text area may not match the screen dimensions, or an inappropriate graphics font may be used. (Of course, the BIOS isn’t going to be able to display text anyway in highly nonstandard modes such as Mode X, but it will do fine in slightly nonstandard modes such as 640x400 16-color mode.) In the case of the 640x400 16-color model described a little earlier, Bill suggests that the code in Listing 45.4 be called immediately after putting the VGA into that mode to tell the BIOS that we’re working with 25 rows of 16-pixel-high text. I think this is an excellent suggestion; it can’t hurt, and may save you from getting aggravating tech support calls down the road.

    -

    LISTING 45.4 L45-4.C

    +

    LISTING 45.4 L45-4.C

     /* Function to tell the BIOS to set up properly sized characters for 25 rows of
        16 pixel high text in 640x400 graphics mode. Call immediately after mode set.
    @@ -134,7 +127,7 @@ void Set640x400()
        int86(0x10, &regs, &regs);           /* invoke the BIOS video interrupt
                                                to set up the text */
     }
    -
    +


    @@ -153,10 +146,6 @@ void Set640x400()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/45-06.html b/45-06.html index d98ba53..39841bb 100644 --- a/45-06.html +++ b/45-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Dog Hair and Dirty Rectangles - - @@ -39,7 +32,7 @@

    The 640x400 mode I’ve described here isn’t exactly earthshaking, but it can come in handy for page flipping and CGA emulation, and I’m sure that some of you will find it useful at one time or another.

    -

    Another Interesting Twist on Page Flipping

    +

    Another Interesting Twist on Page Flipping

    I’ve spent a fair amount of time exploring various ways to do animation. I thought I had pegged all the possible ways to do animation: exclusive-ORing; simply drawing and erasing objects; drawing objects with a blank fringe to erase them at their old locations as they’re drawn; page flipping; and, finally, drawing to local memory and copying the dirty (modified) rectangles to the screen, as I’ve discussed in this chapter.

    @@ -102,10 +95,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/46-01.html b/46-01.html index 4bd0172..23c2d24 100644 --- a/46-01.html +++ b/46-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Who Was that Masked Image? - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Who Was that Masked Image? - - @@ -37,10 +30,10 @@


    -

    Chapter 46
    +

    Chapter 46
    Who Was that Masked Image?

    -

    Optimizing Dirty-Rectangle Animation

    +

    Optimizing Dirty-Rectangle Animation

    Programming is, by and large, a linear process. One statement or instruction follows another, in predictable sequences, with tiny building blocks strung together to make thinking, which is, of course, A Good Thing. Still, it’s important to keep in mind that there’s a large chunk of the human mind that doesn’t work in a linear fashion.

    @@ -52,7 +45,7 @@

    We’re strange thinking machines, but we’re the best ones yet invented, and it’s worth learning how to tap our full potential. And with that, it’s back to dirty-rectangle animation.

    -

    Dirty-Rectangle Animation, Continued

    +

    Dirty-Rectangle Animation, Continued

    In the last chapter, Introduced the idea of dirty-rectangle animation. This technique is an alternative to page flipping that’s capable of producing animation of very high visual quality, without any help at all from video hardware, and without the need for any extra, nondisplayed video memory. This makes dirty-rectangle animation more widely usable than page flipping, because many adapters don’t support page flipping. Dirty-rectangle animation also tends to be simpler to implement than page flipping, because there’s only one bitmap to keep track of. A final advantage of dirty-rectangle animation is that it’s potentially somewhat faster than page flipping, because display-memory accesses can theoretically be reduced to exactly one access for each pixel that changes from one frame to the next.

    @@ -77,10 +70,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/46-02.html b/46-02.html index 5838aa1..fab5007 100644 --- a/46-02.html +++ b/46-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Who Was that Masked Image? - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Who Was that Masked Image? - - @@ -37,7 +30,7 @@


    -

    LISTING 46.1 L46-1.C

    +

    LISTING 46.1 L46-1.C

     /* Sample simple dirty-rectangle animation program, partially optimized and
        featuring internal animation, masked images (sprites), and nonoverlapping dirty
    @@ -406,9 +399,9 @@ void EraseEntities()
        DirtyPtr->Next = TempPtr->Next;
        TempPtr->Next = DirtyPtr;
     }
    -
    + -

    LISTING 46.2 L46-2.ASM

    +

    LISTING 46.2 L46-2.ASM

     ; Assembly language helper routines for dirty rectangle animation. Tested with
     ; TASM. 
    @@ -556,7 +549,7 @@ RowLoop3:
             ret
     -CopyRect       endp
             end
    -
    +


    @@ -575,10 +568,6 @@ RowLoop3:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/46-03.html b/46-03.html index 49a8a0e..b4e92f7 100644 --- a/46-03.html +++ b/46-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Who Was that Masked Image? - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Who Was that Masked Image? - - @@ -37,7 +30,7 @@


    -

    Masked Images

    +

    Masked Images

    Masked images are rendered by drawing an object’s pixels through a mask; pixels are actually drawn only where the mask specifies that drawing is allowed. This makes it possible to draw nonrectangular objects that don’t improperly interfere with one another when they overlap. Masked images also make it possible to have transparent areas (windows) within objects. Masked images produce far more realistic animation than do rectangular images, and therefore are more desirable. Unfortunately, masked images are also considerably slower to draw—however, a good assembly language implementation can go a long way toward making masked images draw rapidly enough, as illustrated by this chapter’s code. (Masked images are also known as sprites; some video hardware supports sprites directly, but on the PC it’s necessary to handle sprites in software.)

    @@ -45,11 +38,11 @@

    In this chapter, I’ve used the approach of having separate, paired masks and images. Another, quite different approach to masking is to specify a transparent color for copying, and copy only those pixels that are not the transparent color. This has the advantage of not requiring separate mask data, so it’s more compact, and the code to implement this is a little less complex than the full masking I’ve implemented. On the other hand, the transparent color approach is less flexible because it makes one color undrawable. Also, with a transparent color, it’s not possible to keep the same base image but use different masks, because the mask information is embedded in the image data.

    -

    Internal Animation

    +

    Internal Animation

    I’ve added another feature essential to producing convincing animation: internal animation, which is the process of changing the appearance of a given object over time, as distinguished from changing only the location of a given object. Internal animation makes images look active and alive. I’ve implemented the simplest possible form of internal animation in Listing 46.1—alternation between two images—but even this level of internal animation greatly improves the feel of the overall animation. You could easily increase the number of images cycled through, simply by increasing the value of InternalAnimateMax for a given entity. You could also implement more complex image-selection logic to produce more interesting and less predictable internal-animation effects, such as jumping, ducking, running, and the like.

    -

    Dirty-Rectangle Management

    +

    Dirty-Rectangle Management

    As mentioned above, dirty-rectangle animation makes it possible to access display memory a minimum number of times. The previous chapter’s code didn’t do any of that; instead, it copied all portions of every dirty rectangle to the screen, regardless of overlap between rectangles. The code I’ve presented in this chapter goes to the other extreme, taking great pains never to draw overlapped portions of rectangles more than once. This is accomplished by checking for overlap whenever a rectangle is to be added to the dirty list. When overlap with an existing rectangle is detected, the new rectangle is reduced to between zero and four nonoverlapping rectangles. Those rectangles are then again considered for addition to the dirty list, and may again be reduced, if additional overlap is detected.

    @@ -59,7 +52,7 @@

    You might also try taking advantage of the natural coherence of animated graphics screens. In particular, because the rectangle used to erase an image at its old location often overlaps the rectangle within which the image resides at its new location, you could just directly generate the two or three nonoverlapped rectangles required to copy both the erase rectangle and the new-image rectangle for any single moving image. The calculation of these rectangles could be very efficient, given that you know in advance the direction of motion of your images. Handling this particular overlap case would eliminate most overlapped drawing, at a minimal cost. You might then decide to ignore overlapped drawing between different images, which tends to be both less common and more expensive to identify and handle.

    -

    Drawing Order and Visual Quality

    +

    Drawing Order and Visual Quality

    A final note on dirty-rectangle animation concerns the quality of the displayed screen image. In the last chapter, we simply stuffed dirty rectangles into a list in the order they became dirty, and then copied all of the rectangles in that same order. Unfortunately, this caused all of the erase rectangles to be copied first, followed by all of the rectangles of the images at their new locations. Consequently, there was a significant delay between the appearance of the erase rectangle for a given image and the appearance of the new rectangle. A byproduct was the fact that a partially complete—part old, part new—image was visible long enough to be noticed. In short, although the pixels ended up correct, they were in an intermediate, incorrect state for a sufficient period of time to make the animation look wrong.

    @@ -84,10 +77,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-01.html b/47-01.html index e49c418..f6321b3 100644 --- a/47-01.html +++ b/47-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -37,10 +30,10 @@


    -

    Chapter 47
    +

    Chapter 47
    Mode X: 256-Color VGA Magic

    -

    Introducing the VGA’s Undocumented “Animation-Optimal” Mode

    +

    Introducing the VGA’s Undocumented “Animation-Optimal” Mode

    At a book signing for my book Zen of Code Optimization, an attractive young woman came up to me, holding my book, and said, “You’re Michael Abrash, aren’t you?” I confessed that I was, prepared to respond in an appropriately modest yet proud way to the compliments I was sure would follow. (It was my own book signing, after all.) It didn’t work out quite that way, though. The first thing out of her mouth was:

    @@ -54,7 +47,7 @@

    So, in the end, I’m thoroughly pleased with Mode X; the world is a better place for it, even if it did cost me my one potential female fan. (Contrary to popular belief, the lives of computer columnists and rock stars are not, repeat, not, all that similar.) This and the following two chapters are based on the DDJ columns that started it all back in 1991, three columns that generated a tremendous amount of interest and spawned a ton of games, and about which I still regularly get letters and e-mail. Ladies and gentlemen, I give you...Mode X.

    -

    What Makes Mode X Special?

    +

    What Makes Mode X Special?

    Consider the strange case of the VGA’s 320x256-color mode—Mode X—which is undeniably complex to program and isn’t even documented by IBM—but which is, nonetheless, perhaps the single best mode the VGA has to offer, especially for animation.

    @@ -89,10 +82,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-02.html b/47-02.html index 3914a5a..459cc1d 100644 --- a/47-02.html +++ b/47-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -41,7 +34,7 @@

    The mode set code is the logical place to begin.

    -

    Selecting 320x240 256-Color Mode

    +

    Selecting 320x240 256-Color Mode

    We could, if we wished, write our own mode set code for Mode X from scratch—but why bother? Instead, we’ll let the BIOS do most of the work by having it set up mode 13H, which we’ll then turn into Mode X by changing a few registers. Listing 47.1 does exactly that.

    @@ -57,7 +50,7 @@

    When people ask why software isn’t bulletproof; why it crashes or doesn’t coexist with certain programs; why PC clones aren’t always compatible; why, in short, the myriad irritations of using a PC exist—this is a big part of the reason. I guess that’s just the price we pay for the unfettered creativity and vast choice of the PC market.

    -

    LISTING 47.1 L47-1.ASM

    +

    LISTING 47.1 L47-1.ASM

     ; Mode X (320x240, 256 colors) mode set routine. Works on all VGAs.
     ; ****************************************************************
    @@ -147,7 +140,7 @@ SetCRTParmsLoop:
             ret
     _Set320x240Mode endp
             end
    -
    +


    @@ -166,10 +159,6 @@ _Set320x240Mode endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-03.html b/47-03.html index f26695e..4f166af 100644 --- a/47-03.html +++ b/47-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -45,11 +38,10 @@

    That’s all right, though, because most graphics software spends little time drawing individual pixels. I’ve provided the write and read pixel routines as basic primitives, and so you’ll understand how the bitmap is organized, but the building blocks of high-performance graphics software are fills, copies, and bitblts, and it’s there that Mode X shines.

    -


    - Figure 47.1
      Mode X display memory organization.

    +


    + Figure 47.1
      Mode X display memory organization.

    -

    LISTING 47.2 L47-2.ASM

    +

    LISTING 47.2 L47-2.ASM

     ; Mode X (320x240, 256 colors) write pixel routine. Works on all VGAs.
     ; No clipping is performed.
    @@ -103,9 +95,9 @@ _WritePixelX    proc    near
             ret
     _WritePixelX    endp
             end
    -
    + -

    LISTING 47.3 L47-3.ASM

    +

    LISTING 47.3 L47-3.ASM

     ; Mode X (320x240, 256 colors) read pixel routine. Works on all VGAs.
     ; No clipping is performed.
    @@ -156,7 +148,7 @@ _ReadPixelX     proc    near
             ret
     _ReadPixelX     endp
             end
    -
    +


    @@ -175,10 +167,6 @@ _ReadPixelX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-04.html b/47-04.html index 507bc66..284aaec 100644 --- a/47-04.html +++ b/47-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -37,11 +30,11 @@


    -

    Designing from a Mode X Perspective

    +

    Designing from a Mode X Perspective

    Listing 47.4 shows Mode X rectangle fill code. The plane is selected for each pixel in turn, with drawing cycling from plane 0 to plane 3, then wrapping back to plane 0. This is the sort of code that stems from a write-pixel line of thinking; it reflects not a whit of the unique perspective that Mode X demands, and although it looks reasonably efficient, it is in fact some of the slowest graphics code you will ever see. I’ve provided Listing 47.4 partly for illustrative purposes, but mostly so we’ll have a point of reference for the substantial speed-up that’s possible with code that’s designed from a Mode X perspective.

    -

    LISTING 47.4 L47-4.ASM

    +

    LISTING 47.4 L47-4.ASM

     ; Mode X (320x240, 256 colors) rectangle fill routine. Works on all
     ; VGAs. Uses slow approach that selects the plane explicitly for each
    @@ -133,7 +126,7 @@ FillDone:
             ret
     _FillRectangleX endp
             end
    -
    +

    The two major weaknesses of Listing 47.4 both result from selecting the plane on a pixel by pixel basis. First, endless OUTs (which are particularly slow on 386s, 486s, and Pentiums, much slower than accesses to display memory) must be performed, and, second, REP STOS can’t be used. Listing 47.5 overcomes both these problems by tailoring the fill technique to the organization of display memory. Each plane is filled in its entirety in one burst before the next plane is processed, so only five OUTs are required in all, and REP STOS can indeed be used; I’ve used REP STOSB in Listings 47.5 and 47.6. REP STOSW could be used and would improve performance on most VGAs; however, REP STOSW requires extra overhead to set up, so it can be slower for small rectangles, especially on 8-bit VGAs. Note that doing an entire plane at a time can produce a “fading-in” effect for large images, because all columns for one plane are drawn before any columns for the next. If this is a problem, the four planes can be cycled through once for each scan line, rather than once for the entire rectangle.

    @@ -156,10 +149,6 @@ _FillRectangleX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-05.html b/47-05.html index 11dbc10..2f91c65 100644 --- a/47-05.html +++ b/47-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -45,7 +38,7 @@
    -

    LISTING 47.5 L47-5.ASM

    +

    LISTING 47.5 L47-5.ASM

     ; Mode X (320x240, 256 colors) rectangle fill routine. Works on all
     ; VGAs. Uses medium-speed approach that selects each plane only once
    @@ -174,7 +167,7 @@ FillDone:
             ret
     _FillRectangleX endp
             end
    -
    +


    @@ -193,10 +186,6 @@ _FillRectangleX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-06.html b/47-06.html index cda7b6b..9ff8605 100644 --- a/47-06.html +++ b/47-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -37,7 +30,7 @@


    -

    Hardware Assist from an Unexpected Quarter

    +

    Hardware Assist from an Unexpected Quarter

    Listing 47.5 illustrates the benefits of designing code from a Mode X perspective; this is the software aspect of Mode X optimization, which suffices to make Mode X about as fast as mode 13H. That alone makes Mode X an attractive mode, given its square pixels, page flipping, and offscreen memory, but superior performance would nonetheless be a pleasant addition to that list. Superior performance is indeed possible in Mode X, although, oddly enough, it comes courtesy of the VGA’s hardware, which was never designed to be used in 256-color modes.

    @@ -47,9 +40,8 @@

    In 16-color modes, each plane contains one-quarter of each of eight pixels, with the 4 bits of each pixel spanning all four planes. Not so in Mode X. Look at Figure 47.1 again; each plane contains one pixel in its entirety, with four pixels at any given address, one per plane. Still, the Map Mask register does the same job in Mode X as in 16-color modes; set it to 0FH (all 1-bits), and all four planes will be written to by each CPU access. Thus, it would seem that up to four pixels could be set by a single Mode X byte-sized write to display memory, potentially speeding up operations like rectangle fills by four times.

    -


    - Figure 47.2
      Selecting planes with the Map Mask register.

    +


    + Figure 47.2
      Selecting planes with the Map Mask register.

    And, as it turns out, four-plane parallelism works quite nicely indeed. Listing 47.6 is yet another rectangle-fill routine, this time using the Map Mask to set up to four pixels per STOS. The only trick to Listing 47.6 is that any left or right edge that isn’t aligned to a multiple-of-four pixel column (that is, a column at which one four-pixel set ends and the next begins) must be clipped via the Map Mask register, because not all pixels at the address containing the edge are modified. Performance is as expected; Listing 47.6 is nearly ten times faster at clearing the screen than Listing 47.4 and just about four times faster than Listing 47.5—and also about four times faster than the same rectangle fill in mode 13H. Understanding the bitmap organization and display hardware of Mode X does indeed pay.

    @@ -72,10 +64,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/47-07.html b/47-07.html index 2b2f597..9ac7cc0 100644 --- a/47-07.html +++ b/47-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X: 256-color VGA Magic - - @@ -37,7 +30,7 @@


    -

    LISTING 47.6 L47-6.ASM

    +

    LISTING 47.6 L47-6.ASM

     ; Mode X (320x240, 256 colors) rectangle fill routine. Works on all
     ; VGAs. Uses fast approach that fans data out to up to four planes at
    @@ -152,13 +145,13 @@ FillDone:
             ret
     _FillRectangleX endp
             end
    -
    +

    Just so you can see Mode X in action, Listing 47.7 is a sample program that selects Mode X and draws a number of rectangles. Listing 47.7 links to any of the rectangle fill routines I’ve presented.

    And now, I hope, you’re beginning to see why I’m so fond of Mode X. In the next chapter, we’ll continue with Mode X by exploring the wonders that the latches and parallel plane hardware can work on scrolls, copies, blits, and pattern fills.

    -

    LISTING 47.7 L47-7.C

    +

    LISTING 47.7 L47-7.C

     /* Program to demonstrate mode X (320x240, 256-colors) rectangle
        fill by drawing adjacent 20x20 rectangles in successive colors from
    @@ -184,7 +177,7 @@ void main() {
        regset.x.ax = 0x0003;   /* switch back to text mode and done */
        int86(0x10, &regset, &regset);
     }
    -
    +


    @@ -203,10 +196,6 @@ void main() {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/48-01.html b/48-01.html index 2ca460a..351ef1f 100644 --- a/48-01.html +++ b/48-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - @@ -37,22 +30,20 @@


    -

    Chapter 48
    +

    Chapter 48
    Mode X Marks the Latch

    -

    The Internals of Animation’s Best Video Display Mode

    +

    The Internals of Animation’s Best Video Display Mode

    In the previous chapter, I introduced you to what I call Mode X, an undocumented 320x240 256-color mode of the VGA. Mode X is distinguished from mode 13H, the documented 320x200 256-color VGA mode, in that it supports page flipping, makes off-screen memory available, has square pixels, and, above all, lets you use the VGA’s hardware to increase performance by as much as four times. (Of course, those four times come at the cost of more complex and demanding programming, to be sure—but end users care about results, not how hard the code was to write, and Mode X delivers results in a big way.) In the previous chapter we saw how the VGA’s plane-oriented hardware can be used to speed solid fills. That’s a nice technique, but now we’re going to move up to the big guns—the VGA latches.

    The VGA has four latches, one for each plane of display memory. Each latch stores exactly one byte, and that byte is always the last byte read from the corresponding plane of display memory, as shown in Figure 48.1. Furthermore, whenever a given address in display memory is read, all four planes’ bytes at that address are read and stored in the corresponding latches, regardless of which plane supplied the byte returned to the CPU (as determined by the Read Map register). As with so much else about the VGA, the above will make little sense to VGA neophytes, but the important point is this: By reading one display memory byte, 4 bytes—one from each plane—can be loaded into the latches at once. Any or all of those 4 bytes can then be written anywhere in display memory with a single byte-sized write, as shown in Figure 48.2.

    -


    - Figure 48.1
      How the VGA latches are loaded.

    +


    + Figure 48.1
      How the VGA latches are loaded.

    -


    - Figure 48.2
      Writing 4 bytes to display memory in a single operation.

    +


    + Figure 48.2
      Writing 4 bytes to display memory in a single operation.

    The upshot is that the latches make it possible to copy data around from one part of display memory to another, 32 bits (four pixels) at a time—four times as fast as normal. (Recall from the previous chapter that in Mode X, pixels are stored one per byte, with four pixels in a row stored in successive planes at the same address, one pixel per plane.) However, any one latch can only be loaded from and written to the corresponding plane, so an individual latch can only work with every fourth pixel on the screen; the latch for plane 0 can work with pixels 0, 4, 8..., the latch for plane 1 with pixels 1, 5, 9..., and so on.

    @@ -60,7 +51,7 @@

    Fast Mode X fills using patterns that are four pixels in width can be performed by drawing the pattern once to the four pixels at any one address in display memory, reading that address to load the pattern into the latches, setting the Bit Mask register to 0 to specify that all bits drawn to display memory should come from the latches, and then performing the fill pretty much as we did in the previous chapter—except that each line of the pattern must be loaded into the latches before the corresponding scan line on the screen is filled. Listings 48.1 and 48.2 together demonstrate a variety of fast Mode X four-by-four pattern fills. (The mode set function called by Listing 48.1 is from the previous chapter’s listings.)

    -

    LISTING 48.1 L48-1.C

    +

    LISTING 48.1 L48-1.C

     /* Program to demonstrate Mode X (320x240, 256 colors) patterned
        rectangle fills by filling the screen with adjacent 80x60
    @@ -106,7 +97,7 @@ void main() {
        regset.x.ax = 0x0003;   /* switch back to text mode and done */
        int86(0x10, &regset, &regset);
     }
    -
    +


    @@ -125,10 +116,6 @@ void main() {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/48-02.html b/48-02.html index b10673b..c33e4ab 100644 --- a/48-02.html +++ b/48-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - @@ -37,7 +30,7 @@


    -

    LISTING 48.2 L48-2.ASM

    +

    LISTING 48.2 L48-2.ASM

     ; Mode X (320x240, 256 colors) rectangle 4x4 pattern fill routine.
     ; Upper-left corner of pattern is always aligned to a multiple-of-4
    @@ -214,7 +207,7 @@ FillDone:
             ret
     _FillPatternX endp
             end
    -
    +


    @@ -233,10 +226,6 @@ _FillPatternX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/48-03.html b/48-03.html index 1369dea..2fbf4b0 100644 --- a/48-03.html +++ b/48-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - @@ -41,7 +34,7 @@

    Furthermore, eight-wide patterns, which are widely used, can be drawn with two passes, one for each half of the pattern. This principle can in fact be extended to patterns of arbitrary multiple-of-four widths. (Widths that aren’t multiples of four are considerably more difficult to handle, because the latches are four pixels wide; one possible solution is expanding such patterns via repetition until they are multiple-of-four widths.)

    -

    Allocating Memory in Mode X

    +

    Allocating Memory in Mode X

    Listing 48.2 raises some interesting questions about the allocation of display memory in Mode X. In Listing 48.2, whenever a pattern is to be drawn, that pattern is first drawn in its entirety at the very end of display memory; the latches are then loaded from that copy of the pattern before each scan line of the actual fill is drawn. Why this double copying process, and why is the pattern stored in that particular area of display memory?

    @@ -49,13 +42,12 @@

    As for why the pattern is stored exactly where it is, that’s part of a master memory allocation plan that will come to fruition in the next chapter, when I implement a Mode X animation program. Figure 48.3 shows this master plan; the first two pages of memory (each 76,800 pixels long, spanning 19,200 addresses—that is, 19,200 pixel quadruplets—in display memory) are reserved for page flipping, the next page of memory (also 76,800 pixels long) is reserved for storing the background (which is used to restore the holes left after images move), the last 16 pixels (four addresses) of display memory are reserved for the pattern buffer, and the remaining 31,728 pixels (7,932 addresses) of display memory are free for storage of icons, images, temporary buffers, or whatever.

    -


    - Figure 48.3
      A useful Mode X display memory layout.

    +


    + Figure 48.3
      A useful Mode X display memory layout.

    This is an efficient organization for animation, but there are certainly many other possible setups. For example, you might choose to have a solid-colored background, in which case you could dispense with the background page (instead using the solid rectangle fill routine to replace the background after images move), freeing up another 76,800 pixels of off-screen storage for images and buffers. You could even eliminate page-flipping altogether if you needed to free up a great deal of display memory. For example, with enough free display memory it is possible in Mode X to create a virtual bitmap three times larger than the screen, with the screen becoming a scrolling window onto that larger bitmap. This technique has been used to good effect in a number of animated games, with and without the use of Mode X.

    -

    Copying Pixel Blocks within Display Memory

    +

    Copying Pixel Blocks within Display Memory

    Another fine use for the latches is copying pixels from one place in display memory to another. Whenever both the source and the destination share the same nibble alignment (that is, their start addresses modulo four are the same), it is not only possible but quite easy to use the latches to copy four pixels at a time. Listing 48.3 shows a routine that copies via the latches. (When the source and destination do not share the same nibble alignment, the latches cannot be used because the source and destination planes for any given pixel differ. In that case, you can set the Read Map register to select a source plane and the Map Mask register to select the corresponding destination plane. Then, copy all pixels in that plane, repeating for all four planes.)

    @@ -84,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/48-04.html b/48-04.html index 18277b9..046561e 100644 --- a/48-04.html +++ b/48-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - @@ -37,7 +30,7 @@


    -

    LISTING 48.3 L48-3.ASM

    +

    LISTING 48.3 L48-3.ASM

     ; Mode X (320x240, 256 colors) display memory to display memory copy
     ; routine. Left edge of source rectangle modulo 4 must equal left edge
    @@ -212,7 +205,7 @@ CopyDone:
             ret
     _CopyScreenToScreenX endp
             end
    -
    +


    @@ -231,10 +224,6 @@ _CopyScreenToScreenX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/48-05.html b/48-05.html index 09866ed..662345e 100644 --- a/48-05.html +++ b/48-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X Marks the Latch - - @@ -41,13 +34,13 @@

    Now that we have a fast way to copy images around in display memory, we can draw icons and other images as much as four times faster than in mode 13H, depending on the speed of the VGA’s display memory. (In case you’re worried about the nibble-alignment limitation on fast copies, don’t be; I’ll address that fully in due time, but the secret is to store all four possible rotations in off-screen memory, then select the correct one for each copy.) However, before our fast display memory-to-display memory copy routine can do us any good, we must have a way to get pixel patterns from system memory into display memory, so that they can then be copied with the fast copy routine.

    -

    Copying to Display Memory

    +

    Copying to Display Memory

    The final piece of the puzzle is the system memory to display-memory-copy-routine shown in Listing 48.4. This routine assumes that pixels are stored in system memory in exactly the order in which they will ultimately appear on the screen; that is, in the same linear order that mode 13H uses. It would be more efficient to store all the pixels for one plane first, then all the pixels for the next plane, and so on for all four planes, because many OUTs could be avoided, but that would make images rather hard to create. And, while it is true that the speed of drawing images is, in general, often a critical performance factor, the speed of copying images from system memory to display memory is not particularly critical in Mode X. Important images can be stored in off-screen memory and copied to the screen via the latches much faster than even the speediest system memory-to-display memory copy routine could manage.

    I’m not going to present a routine to perform Mode X copies from display memory to system memory, but such a routine would be a straightforward inverse of Listing 48.4.

    -

    LISTING 48.4 L48-4.ASM

    +

    LISTING 48.4 L48-4.ASM

     ; Mode X (320x240, 256 colors) system memory to display memory copy
     ; routine. Uses approach of changing the plane for each pixel copied;
    @@ -167,9 +160,9 @@ CopyDone:
             ret
     _CopySystemToScreenX endp
             end
    -
    + -

    Who Was that Masked Image Copier?

    +

    Who Was that Masked Image Copier?

    At this point, it’s getting to be time for us to take all the Mode X tools we’ve developed, together with one more tool—masked image copying—and the remaining unexplored feature of Mode X, page flipping, and build an animation application. I hope that when we’re done, you’ll agree with me that Mode X is the way to animate on the PC.

    @@ -194,10 +187,6 @@ _CopySystemToScreenX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/49-01.html b/49-01.html index 4b4ac72..0546842 100644 --- a/49-01.html +++ b/49-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - @@ -37,14 +30,14 @@


    -

    Chapter 49
    +

    Chapter 49
    Mode X 256-Color Animation

    -

    How to Make the VGA Really Get up and Dance

    +

    How to Make the VGA Really Get up and Dance

    Okay—no amusing stories or informative anecdotes to kick off this chapter; lotta ground to cover, gotta hurry—you’re impatient, I can smell it. I won’t talk about the time a friend made the mistake of loudly saying “$100 bill” during an animated discussion while walking among the bums on Market Street in San Francisco one night, thereby graphically illustrating that context is everything. I can’t spare a word about how my daughter thinks my 11-year-old floppy-disk-based CP/M machine is more powerful than my 386 with its 100-MB hard disk because the CP/M machine’s word processor loads and runs twice as fast as the 386’s Windows-based word processor, demonstrating that progress is not the neat exponential curve we’d like to think it is, and that features and performance are often conflicting notions. And, lord knows, I can’t take the time to discuss the habits of small white dogs, notwithstanding that such dogs seem to be relevant to just about every aspect of computing, as Jeff Duntemann’s writings make manifest. No lighthearted fluff for us; we have real work to do, for today we animate with 256 colors in Mode X.

    -

    Masked Copying

    +

    Masked Copying

    Over the past two chapters, we’ve put together most of the tools needed to implement animation in the VGA’s undocumented 320x240 256-color Mode X. We now have mode set code, solid and 4x4 pattern fills, system memory-to-display memory block copies, and display memory-to-display memory block copies. The final piece of the puzzle is the ability to copy a nonrectangular image to display memory. I call this masked copying.

    @@ -54,7 +47,7 @@

    The system memory to display memory masked copy routine in Listing 49.1 implements masked copying in a straightforward fashion. In the main drawing loop, the corresponding mask byte is consulted as each image pixel is encountered, and the image pixel is copied only if the mask byte is nonzero. As with most of the system-to-display code I’ve presented, Listing 49.1 is not heavily optimized, because it’s inherently slow; there’s a better way to go when performance matters, and that’s to use the VGA’s hardware.

    -

    LISTING 49.1 L49-1.ASM

    +

    LISTING 49.1 L49-1.ASM

     ; Mode X (320x240, 256 colors) system memory-to-display memory masked copy
     ; routine. Not particularly fast; images for which performance is critical
    @@ -190,7 +183,7 @@ CopyDone:
             ret
     _CopySystemToScreenMaskedX endp
             end
    -
    +


    @@ -209,10 +202,6 @@ _CopySystemToScreenMaskedX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/49-02.html b/49-02.html index 98234b6..6db6dbf 100644 --- a/49-02.html +++ b/49-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - @@ -37,7 +30,7 @@


    -

    Faster Masked Copying

    +

    Faster Masked Copying

    In the previous chapter we saw how the VGA’s latches can be used to copy four pixels at a time from one area of display memory to another in Mode X. We’ve further seen that in Mode X the Map Mask register can be used to select which planes are copied. That’s all we need to know to be able to perform fast masked copies; we can store an image in off-screen display memory, and set the Map Mask to the appropriate mask value as up to four pixels at a time are copied.

    @@ -45,7 +38,7 @@

    Listing 49.2 performs fast masked copying. This code expects to receive a pointer to a MaskedImage structure, which in turn points to four AlignedMaskedImage structures that describe the four possible image and mask alignments. The aligned images are already stored in display memory, and the aligned masks are already stored in system memory; further, the masks are predigested into Map Mask register-compatible form. Given all that ready-to-use data, Listing 49.2 selects and works with the appropriate image-mask pair for the destination’s left edge alignment.

    -

    LISTING 49.2 L49-2.ASM

    +

    LISTING 49.2 L49-2.ASM

     ; Mode X (320x240, 256 colors) display memory to display memory masked copy
     ; routine. Works on all VGAs. Uses approach of reading 4 pixels at a time from
    @@ -208,7 +201,7 @@ CopyDone:
     _CopyScreenToScreenMaskedX endp
             end
     
    -
    +


    @@ -227,10 +220,6 @@ _CopyScreenToScreenMaskedX endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/49-03.html b/49-03.html index 031c531..d758f0b 100644 --- a/49-03.html +++ b/49-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - @@ -39,7 +32,7 @@

    It would be handy to have a function that, given a base image and mask, generates the four image and mask alignments and fills in the MaskedImage structure. Listing 49.3, together with the include file in Listing 49.4 and the system memory-to-display memory block-copy routine in Listing 48.4 (in the previous chapter) does just that. It would be faster if Listing 49.3 were in assembly language, but there’s no reason to think that generating aligned images needs to be particularly fast; in such cases, I prefer to use C, for reasons of coding speed, fewer bugs, and maintainability.

    -

    LISTING 49.3 L49-3.C

    +

    LISTING 49.3 L49-3.C

     /* Generates all four possible mode X image/mask alignments, stores image
     alignments in display memory, allocates memory for and generates mask
    @@ -107,9 +100,9 @@ unsigned int CreateAlignedMaskedImage(MaskedImage * ImageToSet,
        }
        return DispMemOffset - DispMemStart;
     }
    -
    + -

    LISTING 49.4 MASKIM.H

    +

    LISTING 49.4 MASKIM.H

     /* MASKIM.H: structures used for storing and manipulating masked
        images */
    @@ -128,9 +121,9 @@ typedef struct {
                                              structs for four possible destination
                                              image alignments */
     } MaskedImage;
    -
    + -

    Notes on Masked Copying

    +

    Notes on Masked Copying

    Listings 49.1 and 49.2, like all Mode X code I’ve presented, perform no clipping, because clipping code would complicate the listings too much. While clipping can be implemented directly in the low-level Mode X routines (at the beginning of Listing 49.1, for instance), another, potentially simpler approach would be to perform clipping at a higher level, modifying the coordinates and dimensions passed to low-level routines such as Listings 49.1 and 49.2 as necessary to accomplish the desired clipping. It is for precisely this reason that the low-level Mode X routines support programmable start coordinates in the source images, rather than assuming (0,0); likewise for the distinction between the width of the image and the width of the area of the image to draw.

    @@ -161,10 +154,6 @@ typedef struct {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/49-04.html b/49-04.html index 16c545d..e9ef7be 100644 --- a/49-04.html +++ b/49-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - @@ -37,7 +30,7 @@


    -

    Animation

    +

    Animation

    Gosh. There’s just no way I can discuss high-level animation fundamentals in any detail here; I could spend an entire (and entirely separate) book on animation techniques alone. You might want to have a look at Chapters 43 through 46 before attacking the code in this chapter; that will have to do us for the present volume. (I will return to 3-D animation in the next chapter.)

    @@ -45,15 +38,14 @@

    Some of the code in this chapter was adapted for Mode X from the code in Chapter 44—yet another reason to read that chapter before finishing this one.

    -

    Mode X Animation in Action

    +

    Mode X Animation in Action

    Listing 49.5 ties together everything I’ve discussed about Mode X so far in a compact but surprisingly powerful animation package. Listing 49.5 first uses solid and patterned fills and system-memory-to-screen-memory masked copying to draw a static background containing a mountain, a sun, a plain, water, and a house with puffs of smoke coming out of the chimney, and sets up the four alignments of a masked kite image. The background is transferred to both display pages, and drawing of 20 kite images in the nondisplayed page using fast masked copying begins. After all images have been drawn, the page is flipped to show the newly updated screen, and the kites are moved and drawn in the other page, which is no longer displayed. Kites are erased at their old positions in the nondisplayed page by block copying from the background page. (See the discussion in the previous chapter for the display memory organization used by Listing 49.5.) So far as the displayed image is concerned, there is never any hint of flicker or disturbance of the background. This continues at a rate of up to 60 times a second until Esc is pressed to exit the program. See Figure 49.1 for a screen shot of the resulting image—add the animation in your imagination.

    -


    - Figure 49.1
      An animated Mode X screen.

    +


    + Figure 49.1
      An animated Mode X screen.

    -

    LISTING 49.5 L49-5.C

    +

    LISTING 49.5 L49-5.C

     /* Sample mode X VGA animation program. Portions of this code first appeared
        in PC Techniques. Compiled with Borland C++ 2.0 in C compilation mode. */
    @@ -289,7 +281,7 @@ void MoveObject(AnimatedObject * ObjectToMove) {
        ObjectToMove->Y = Y;
     }
     
    -
    +


    @@ -308,10 +300,6 @@ void MoveObject(AnimatedObject * ObjectToMove) {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/49-05.html b/49-05.html index 1247b2e..32ca7fd 100644 --- a/49-05.html +++ b/49-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Mode X 256-Color Animation - - @@ -43,7 +36,7 @@

    The external functions called by Listing 49.5 can be found in Listings 49.1, 49.2, 49.3, and 49.6, and in the listings for the previous two chapters.

    -

    LISTING 49.6 L49-6.ASM

    +

    LISTING 49.6 L49-6.ASM

     ; Shows the page at the specified offset in the bitmap. Page is displayed when
     ; this routine returns.
    @@ -91,9 +84,9 @@ WaitVS:
             ret
     _ShowPage       endp
             end
    -
    + -

    Works Fast, Looks Great

    +

    Works Fast, Looks Great

    We now end our exploration of Mode X, although we’ll use it again shortly for 3-D animation. Mode X admittedly has its complexities; that’s why I’ve provided a broad and flexible primitive set. Still, so what if it is complex? Take a look at Listing 49.5 in action. That sort of colorful, high-performance animation is worth jumping through a few hoops for; drawing 20, or even 10, fair-sized objects at a rate of 60 Hz, with no flicker, interference, or fringe, is no mean accomplishment, even on a 386.

    @@ -116,10 +109,6 @@ _ShowPage endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-01.html b/50-01.html index 9fac23a..f2d1b81 100644 --- a/50-01.html +++ b/50-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -37,10 +30,10 @@


    -

    Chapter 50
    +

    Chapter 50
    Adding a Dimension

    -

    3-D Animation Using Mode X

    +

    3-D Animation Using Mode X

    When I first started programming micros, more than 11 years ago now, there wasn’t much money in it, or visibility, or anything you could call a promising career. Sometimes, it was a way to accomplish things that would never have gotten done otherwise because minicomputer time cost too much; other times, it paid the rent; mostly, though, it was just for fun. Given free computer time for the first time in my life, I went wild, writing versions of all sorts of software I had seen on mainframes, in arcades, wherever. It was a wonderful way to learn how computers work: Trial and error in an environment where nobody minded the errors, with no meter ticking.

    @@ -56,13 +49,13 @@

    In a sense, I’ve saved the best for last, because, to my mind, real-time 3-D animation is one of the most exciting things of any stripe that can be done with a computer—and because, with today’s hardware, it can in fact be done. Nay, it can be done amazingly well.

    -

    References on 3-D Drawing

    +

    References on 3-D Drawing

    There are several good sources for information about 3-D graphics. Foley and van Dam’s Computer Graphics: Principles and Practice (Second Edition, Addison-Wesley, 1990) provides a lengthy discussion of the topic and a great many references for further study. Unfortunately, this book is heavy going at times; a more approachable discussion is provided in Principles of Interactive Computer Graphics, by Newman and Sproull (McGraw-Hill, 1979). Although the latter book lacks the last decade’s worth of graphics developments, it nonetheless provides a good overview of basic 3-D techniques, including many of the approaches likely to work well in realtime on a PC.

    A source that you may or may not find useful is the series of six books on C graphics by Lee Adams, as exemplified by High-Performance CAD Graphics in C (Windcrest/Tab, 1986). (I don’t know if all six books discuss 3-D graphics, but the four I’ve seen do.) To be honest, this book has a number of problems, including: Relatively little theory and explanation; incomplete and sometimes erroneous discussions of graphics hardware; use of nothing but global variables, with cryptic names like “array3” and “B21;” and—well, you get the idea. On the other hand, the book at least touches on a great many aspects of 3-D drawing, and there’s a lot of C code to back that up. A number of people have spoken warmly to me of Adams’ books as their introduction to 3-D graphics. I wouldn’t recommend these books as your only 3-D references, but if you’re just starting out, you might want to look at one and see if it helps you bridge the gap between the theory and implementation of 3-D graphics.

    -

    The 3-D Drawing Pipeline

    +

    The 3-D Drawing Pipeline

    Each 3-D object that we’ll handle will be built out of polygons that represent the surface of the object. Figure 50.1 shows the stages a polygon goes through enroute to being drawn on the screen. (For the present, we’ll avoid complications such as clipping, lighting, and shading.) First, the polygon is transformed from object space, the coordinate system the object is defined in, to world space, the coordinate system of the 3-D universe. Transformation may involve rotating, scaling, and moving the polygon. Fortunately, applying the desired transformation to each of the polygon vertices in an object is equivalent to transforming the polygon; in other words, transformation of a polygon is fully defined by transformation of its vertices, so it is not necessary to transform every point in a polygon, just the vertices. Likewise, transformation of all the polygon vertices in an object fully transforms the object.

    @@ -87,10 +80,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-02.html b/50-02.html index 760be0a..62a436c 100644 --- a/50-02.html +++ b/50-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -39,47 +32,42 @@

    One note: I’ll use a purely right-handed convention for coordinate systems. Right-handed means that if you hold your right hand with your fingers curled and the thumb sticking out, the thumb points along the Z axis and the fingers point in the direction of rotation from the X axis to the Y axis, as shown in Figure 50.2. Rotations about an axis are counter-clockwise, as viewed looking down an axis toward the origin. The handedness of a coordinate system is just a convention, and left-handed would do equally well; however, right-handed is generally used for object and world space. Sometimes, the handedness is flipped for view space, so that increasing Z equals increasing distance from the viewer along the line of sight, but I have chosen not to do that here, to avoid confusion. Therefore, Z decreases as distance along the line of sight increases; a view space coordinate of (0,0,-1000) is directly ahead, twice as far away as a coordinate of (0,0,-500).

    -


    - Figure 50.1
      The 3-D drawing pipeline.

    +


    + Figure 50.1
      The 3-D drawing pipeline.

    -


    - Figure 50.2
      A right-handed coordinate system.

    +


    + Figure 50.2
      A right-handed coordinate system.

    -

    Projection

    +

    Projection

    Working backward from the final image, we want to take the vertices of a polygon, as transformed into view space, and project them to 2-D coordinates on the screen, which, for projection purposes, is assumed to be centered on and perpendicular to the Z axis in view space, at some distance from the screen. We’re after visual realism, so we’ll want to do a perspective projection, in order that farther objects look smaller than nearer objects, and so that the field of view will widen with distance. This is done by scaling the X and Y coordinates of each point proportionately to the Z distance of the point from the viewer, a simple matter of similar triangles, as shown in Figure 50.3. It doesn’t really matter how far down the Z axis the screen is assumed to be; what matters is the ratio of the distance of the screen from the viewpoint to the width of the screen. This ratio defines the rate of divergence of the viewing pyramid—the full field of view—and is used for performing all perspective projections. Once perspective projection has been performed, all that remains before calling the polygon filler is to convert the projected X and Y coordinates to integers, appropriately clipped and adjusted as necessary to center the origin on the screen or otherwise map the image into a window, if desired.

    -

    Translation

    +

    Translation

    Translation means adding X, Y, and Z offsets to a coordinate to move it linearly through space. Translation is as simple as it seems; it requires nothing more than an addition for each axis. Translation is, for example, used to move objects from object space, in which the center of the object is typically the origin (0,0,0), into world space, where the object may be located anywhere.

    -


    - Figure 50.3
      Perspective projection.

    +


    + Figure 50.3
      Perspective projection.

    -

    Rotation

    +

    Rotation

    Rotation is the process of circularly moving coordinates around the origin. For our present purposes, it’s necessary only to rotate objects about their centers in object space, so as to turn them to the desired attitude before translating them into world space.

    Rotation of a point about an axis is accomplished by transforming it according to the formulas shown in Figure 50.4. These formulas map into the more generally useful matrix-multiplication forms also shown in Figure 50.4. Matrix representation is more useful for two reasons: First, it is possible to concatenate multiple rotations into a single matrix by multiplying them together in the desired order; that single matrix can then be used to perform the rotations more efficiently.

    -


    - Figure 50.4
      3-D rotation formulas.

    +


    + Figure 50.4
      3-D rotation formulas.

    Second, 3x3 rotation matrices can become the upper-left-hand portions of 4x4 matrices that also perform translation (and scaling as well, but we won’t need scaling in the near future), as shown in Figure 50.5. A 4x4 matrix of this sort utilizes homogeneous coordinates; that’s a topic way beyond this book, but, basically, homogeneous coordinates allow you to handle both rotations and translations with 4x4 matrices, thereby allowing the same code to work with either, and making it possible to concatenate a long series of rotations and translations into a single matrix that performs the same transformation as the sequence of rotations and transformations.

    There’s much more to be said about transformations and the supporting matrix math, but, in the interests of getting to working code in this chapter, I’ll leave that to be discussed as the need arises.

    -

    A Simple 3-D Example

    +

    A Simple 3-D Example

    At this point, we know enough to be able to put together a simple working 3-D animation example. The example will do nothing more complicated than display a single polygon as it sits in 3-D space, rotating around the Y axis. To make things a little more interesting, we’ll let the user move the polygon around in space with the arrow keys, and with the “A” (away), and “T” (toward) keys. The sample program requires two sorts of functionality: The ability to transform and project the polygon from object space onto the screen (3-D functionality), and the ability to draw the projected polygon (complete with clipping) and handle the other details of animation (2-D functionality).

    -


    - Figure 50.5
      A 4x4 Transformation Matrix.

    +


    + Figure 50.5
      A 4x4 Transformation Matrix.

    Happily (and not coincidentally), we put together a nice 2-D animation framework back in Chapters 47, 48, and 49, during our exploratory discussion of Mode X, so we don’t have much to worry about in terms of non-3-D details. Basically, we’ll use Mode X (320x240, 256 colors), and we’ll flip between two display pages, drawing to one while the other is displayed. One new 2-D element that we need is the ability to clip polygons; while we could avoid this for the moment by restricting the range of motion of the polygon so that it stays fully on the screen, certainly in the long run we’ll want to be able to handle partially or fully clipped polygons. Listing 50.1 is the low-level code for a Mode X polygon filler that supports clipping. (The high-level polygon fill code is mode independent, and is the same as that presented in Chapters 38, 39, and 40, as noted further on.) The clipping is implemented at the low level, by trimming the Y extent of the scan line list up front, then clipping the X coordinates of each scan line in turn. This is not a particularly fast approach to clipping—ideally, the polygon would be clipped before it was scanned into a line list, avoiding potentially wasted scanning and eliminating the line-by-line X clipping—but it’s much simpler, and, as we shall see, polygon filling performance is the least of our worries at the moment.

    @@ -100,10 +88,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-03.html b/50-03.html index 6287ef6..e875c15 100644 --- a/50-03.html +++ b/50-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -37,7 +30,7 @@


    -

    LISTING 50.1 L50-1.ASM

    +

    LISTING 50.1 L50-1.ASM

     ; Draws all pixels in the list of horizontal lines passed in, in
     ; Mode X, the VGA’s undocumented 320x240 256-color mode. Clips to
    @@ -195,7 +188,7 @@ FillDone:
             ret
     _DrawHorizontalLineList endp
             end
    -
    +


    @@ -214,10 +207,6 @@ _DrawHorizontalLineList endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-04.html b/50-04.html index c288c19..2a805fc 100644 --- a/50-04.html +++ b/50-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -43,7 +36,7 @@

    Other modules required are: Listings 47.1 and 47.6 from Chapter 47 (Mode X mode set, rectangle fill); Listing 49.6 from Chapter 49; Listing 39.4 from Chapter 39 (polygon edge scan); and the FillConvexPolygon() function from Listing 38.1 in Chapter 38. All necessary code modules, along with a project file, are present in the subdirectory for this chapter on the listings disk, whether they were presented in this chapter or some earlier chapter. This will be the case for the next several chapters as well, where listings from previous chapters are referenced. This scheme may crowd the listings diskette a little bit, but it will certainly reduce confusion!

    -

    LISTING 50.2 L50-2.C

    +

    LISTING 50.2 L50-2.C

     /* Matrix arithmetic functions.
        Tested with Borland C++ in the small model. */
    @@ -89,9 +82,9 @@ void ConcatXforms(double SourceXform1[4][4], double SourceXform2[4][4],
           }
        }
     }
    -
    + -

    LISTING 50.3 L50-3.C

    +

    LISTING 50.3 L50-3.C

     /* Transforms convex polygon Poly (which has PolyLength vertices),
        performing the transformation according to Xform (which generally
    @@ -146,7 +139,7 @@ void XformAndProjectPoly(double Xform[4][4], struct Point3 * Poly,
        /* Draw the polygon */
        DRAW_POLYGON(ProjectedPoly, PolyLength, Color, 0, 0);
     }
    -
    +


    @@ -165,10 +158,6 @@ void XformAndProjectPoly(double Xform[4][4], struct Point3 * Poly,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-05.html b/50-05.html index 9467b94..b00c9df 100644 --- a/50-05.html +++ b/50-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -37,7 +30,7 @@


    -

    LISTING 50.4 POLYGON.H

    +

    LISTING 50.4 POLYGON.H

     /* POLYGON.H: Header file for polygon-filling code, also includes
        a number of useful items for 3-D animation. */
    @@ -111,7 +104,7 @@ extern void FillRectangleX(int StartX, int StartY, int EndX,
        int EndY, unsigned int PageBase, int Color);
     extern int DisplayedPage, NonDisplayedPage;
     extern struct Rect EraseRect[];
    -
    +


    @@ -130,10 +123,6 @@ extern struct Rect EraseRect[];
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-06.html b/50-06.html index 54e65eb..f90f53c 100644 --- a/50-06.html +++ b/50-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -37,7 +30,7 @@


    -

    LISTING 50.5 L50-5.C

    +

    LISTING 50.5 L50-5.C

     /* Simple 3-D drawing program to view a polygon as it rotates in
        Mode X. View space is congruent with world space, with the
    @@ -159,7 +152,7 @@ void main() {
        regset.x.ax = 0x0003;   /* AL = 3 selects 80x25 text mode */
        int86(0x10, &regset, &regset);
     }
    -
    +


    @@ -178,10 +171,6 @@ void main() {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/50-07.html b/50-07.html index 6a88150..49a6601 100644 --- a/50-07.html +++ b/50-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Adding a Dimension - - @@ -37,7 +30,7 @@


    -

    Notes on the 3-D Animation Example

    +

    Notes on the 3-D Animation Example

    The sample program transforms the polygon’s vertices from object space to world space to view space to the screen, as described earlier. In this case, world space and view space are congruent—we’re looking right down the negative Z axis of world space—so the transformation matrix from world to view is the identity matrix; you might want to experiment with changing this matrix to change the viewpoint. The sample program uses 4x4 homogeneous coordinate matrices to perform transformations, as described above. Floating-point arithmetic is used for all 3-D calculations. Setting the translation from object space to world space is a simple matter of changing the appropriate entry in the fourth column of the object-to-world transformation matrix. Setting the rotation around the Y axis is almost as simple, requiring only the setting of the four matrix entries that control the Y rotation to the sines and cosines of the desired rotation. However, rotations involving more than one axis require multiple rotation matrices, one for each axis rotated around; those matrices are then concatenated together to produce the object-to-world transformation. This area is trickier than it might initially appear to be; more in the near future.

    @@ -49,7 +42,7 @@

    Finally, observe the jaggies crawling along the edges of the polygon as it rotates. This is temporal aliasing at its finest! We won’t address antialiasing further, realtime antialiasing being decidedly nontrivial, but this should give you an idea of why antialiasing is so desirable.

    -

    An Ongoing Journey

    +

    An Ongoing Journey

    In the next chapter, we’ll assign fronts and backs to polygons, and start drawing only those that are facing the viewer. That will enable us to handle convex polyhedrons, such as tetrahedrons and cubes. We’ll also look at interactively controllable rotation, and at more complex rotations than the simple rotation around the Y axis that we did this time. In time, we’ll use fixed-point arithmetic to speed things up, and do some shading and texture mapping. The journey has only begun; we’ll get to all that and more soon.

    @@ -70,10 +63,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/51-01.html b/51-01.html index 7c24ee5..118597d 100644 --- a/51-01.html +++ b/51-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - @@ -37,10 +30,10 @@


    -

    Chapter 51
    +

    Chapter 51
    Sneakers in Space

    -

    Using Backface Removal to Eliminate Hidden Surfaces

    +

    Using Backface Removal to Eliminate Hidden Surfaces

    As I’m fond of pointing out, computer animation isn’t a matter of mathematically exact modeling or raw technical prowess, but rather of fooling the eye and the mind. That’s especially true for 3-D animation, where we’re not only trying to convince viewers that they’re seeing objects on a screen—when in truth that screen contains no objects at all, only gaggles of pixels—but we’re also trying to create the illusion that the objects exist in three-space, possessing four dimensions (counting movement over time as a fourth dimension) of their own. To make this magic happen, we must provide cues for the eye not only to pick out boundaries, but also to detect depth, orientation, and motion. This involves perspective, shading, proper handling of hidden surfaces, and rapid and smooth screen updates; the whole deal is considerably more difficult to pull off on a PC than 2-D animation.

    @@ -58,7 +51,7 @@

    If it’s good enough for George Lucas, it’s good enough for us. And with that, let’s resume our quest for realtime 3-D animation on the PC.

    -

    One-sided Polygons: Backface Removal

    +

    One-sided Polygons: Backface Removal

    In the previous chapter, we implemented the basic polygon drawing pipeline, transforming a polygon all the way from its basic definition in object space, through the shared 3-D world space, and into the 3-D space as seen from the viewpoint, called view space. From view space, we performed a perspective projection to convert the polygon into screen space, then mapped the transformed and projected vertices to the nearest screen coordinates and filled the polygon. Armed with code that implemented this pipeline, we were able to watch as a polygon rotated about its Y axis, and were able to move the polygon around in space freely.

    @@ -81,10 +74,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/51-02.html b/51-02.html index 7b77e70..a443298 100644 --- a/51-02.html +++ b/51-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - @@ -45,23 +38,21 @@

    The cross-product of two vectors is defined as the vector shown in Figure 51.1. One interesting property of the cross-product vector is that it is perpendicular to the plane in which the two original vectors lie. If we take the cross-product of the vectors that form two edges of a polygon, the result will be a vector perpendicular to the polygon; then, we’ll know that the polygon is visible if and only if the cross-product vector points toward the viewer. We need one more thing to make the cross-product approach work, though. The cross-product can actually point either way, depending on which edges of the polygon we choose to work with and the order in which we evaluate them, so we must establish some conventions for defining polygons and evaluating the cross-product.

    -


    - Figure 51.1
      The cross-product of two vectors.

    +


    + Figure 51.1
      The cross-product of two vectors.

    We’ll define only convex polygons, with the vertices defined in clockwise order, as viewed from the outside; that is, if you’re looking at the visible side of the polygon, the vertices will appear in the polygon definition in clockwise order. With those assumptions, the cross-product becomes a quick and easy indicator of polygon orientation with respect to the viewer; we’ll calculate it as the cross-product of the first and last vectors in a polygon, as shown in Figure 51.2, and if it’s pointing toward the viewer, we’ll know that the polygon is visible. Actually, we don’t even have to calculate the entire cross-product vector, because the Z component alone suffices to tell us which way the polygon is facing: positive Z means visible, negative Z means not. The Z component can be calculated very efficiently, with only two multiplies and a subtraction.

    The question remains of the proper space in which to perform backface removal. There’s a temptation to perform it in view space, which is, after all, the space defined with respect to the viewer, but view space is not a good choice. Screen space—the space in which perspective projection has been performed—is the best choice. The purpose of backface removal is to determine whether each polygon is visible to the viewer, and, despite its name, view space does not provide that information; unlike screen space, it does not reflect perspective effects.

    -


    - Figure 51.2
      Using the cross product to generate a polygon normal.

    +


    + Figure 51.2
      Using the cross product to generate a polygon normal.

    Backface removal may also be performed using the polygon vertices in screen coordinates, which are integers. This is less accurate than using the screen space coordinates, which are floating point, but is, by the same token, faster. In Listing 51.3, which we’ll discuss shortly, backface removal is performed in screen coordinates in the interests of speed.

    Backface removal, as implemented in Listing 51.3, will not work reliably if the polygon is not convex, if the vertices don’t appear in clockwise order, if either the first or last edge in a polygon has zero length, or if the first and last edges are collinear. These latter two points are the reason it’s preferable to work in screen space rather than screen coordinates (which suffer from rounding problems), speed considerations aside.

    -

    Backface Removal in Action

    +

    Backface Removal in Action

    Listings 51.1 through 51.5 together form a program that rotates a solid cube in real-time under user control. Listing 51.1 is the main program; Listing 51.2 performs transformation and projection; Listing 51.3 performs backface removal and draws visible faces; Listing 51.4 concatenates incremental rotations to the object-to-world transformation matrix; Listing 51.5 is the general header file. Also required from previous chapters are: Listings 50.1 and 50.2 from Chapter 50 (draw clipped line list, matrix math functions); Listings 47.1 and 47.6 from Chapter 47, (Mode X mode set, rectangle fill); Listing 49.6 from Chapter 49; Listing 39.4 from Chapter 39 (polygon edge scan); and the FillConvexPolygon() function from Listing 38.1 from Chapter 38. All necessary modules, along with a project file, will be present in the subdirectory for this chapter on the listings diskette, whether they were presented in this chapter or some earlier chapter. This may crowd the listings diskette a little bit, but it will certainly reduce confusion!

    @@ -82,10 +73,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/51-03.html b/51-03.html index f2c5491..b109f72 100644 --- a/51-03.html +++ b/51-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - @@ -37,7 +30,7 @@


    -

    LISTING 51.1 L51-1.C

    +

    LISTING 51.1 L51-1.C

     /* 3D animation program to view a cube as it rotates in Mode X. The viewpoint
        is fixed at the origin (0,0,0) of world space, looking in the direction of
    @@ -199,7 +192,7 @@ void main() {
        regset.x.ax = 0x0003;   /* AL = 3 selects 80x25 text mode */
        int86(0x10, &regset, &regset);
     }
    -
    +


    @@ -218,10 +211,6 @@ void main() {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/51-04.html b/51-04.html index c593253..113185f 100644 --- a/51-04.html +++ b/51-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - @@ -37,7 +30,7 @@


    -

    LISTING 51.2 L51-2.C

    +

    LISTING 51.2 L51-2.C

     /* Transforms all vertices in the specified object into view spa ce, then
        perspective projects them to screen space and maps them to screen coordinates,
    @@ -73,9 +66,9 @@ void XformAndProjectPoints(double Xform[4][4],
                    SCREEN_HEIGHT/2;
        }
     }
    -
    + -

    LISTING 51.3 L51-3.C

    +

    LISTING 51.3 L51-3.C

     /* Draws all visible faces (faces pointing toward the viewer) in the specified
        object. The object must have previously been transformed and projected, so
    @@ -133,19 +126,17 @@ void DrawVisibleFaces(struct Object * ObjectToXform)
        }
     }
     
    -
    +

    The sample program, as shown in Figure 51.3, places a cube, floating in three-space, under the complete control of the user. The arrow keys may be used to move the cube left, right, up, and down, and the A and T keys may be used to move the cube away from or toward the viewer. The F1 and F2 keys perform rotation around the Z axis, the axis running from the viewer straight into the screen. The 4 and 6 keys perform rotation around the Y (vertical) axis, and the 2 and 8 keys perform rotation around the X axis, which runs horizontally across the screen; the latter four keys are most conveniently used by flipping the keypad to the numeric state.

    -


    - Figure 51.3
      Sample screens from the 3-D cube program.

    +


    + Figure 51.3
      Sample screens from the 3-D cube program.

    The demo involves six polygons, one for each side of the cube. Each of the polygons must be transformed and projected, so it would seem that 24 vertices (four for each polygon) must be handled, but some steps have been taken to improve performance. All vertices for the object have been stored in a single list; the definition of each face contains not the vertices for that face themselves, but rather indexes into the object’s vertex list, as shown in Figure 51.4. This reduces the number of vertices to be manipulated from 24 to 8, for there are, after all, only eight vertices in a cube, with three faces sharing each vertex. In this way, the transformation burden is lightened by two-thirds. Also, as mentioned earlier, backface removal is performed with integers, in screen coordinates, rather than with floating-point values in screen space. Finally, the RecalcXForm flag is set whenever the user changes the object-to-world transformation. Only when this flag is set is the full object-to-view transformation recalculated and the object’s vertices transformed and projected again; otherwise, the values already stored within the object are reused. In the sample application, this brings no visual improvement, because there’s only the one object, but the underlying mechanism is sound: In a full-blown 3-D animation application, with multiple objects moving about the screen, it would help a great deal to flag which of the objects had moved with respect to the viewer, performing a new transformation and projection only for those that had.

    -


    - Figure 51.4
      The object data structure

    +


    + Figure 51.4
      The object data structure


    @@ -164,10 +155,6 @@ void DrawVisibleFaces(struct Object * ObjectToXform)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/51-05.html b/51-05.html index 0c9cdd8..5c7c7f1 100644 --- a/51-05.html +++ b/51-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - @@ -39,13 +32,13 @@

    With the above optimizations, the sample program is certainly adequately responsive on a 20 MHz 386 (sans 387; I’m sure it’s wonderfully responsive with a math coprocessor). Still, it couldn’t quite keep up with the keyboard when I modified it to read only one key each time through the loop—and we’re talking about only eight vertices here. This indicates that we’re already near the limit of animation complexity possible with our current approach. It’s time to start rethinking that approach; over two-thirds of the overall time is spent in floating-point calculations, and it’s there that we’ll begin to attack the performance bottleneck we find ourselves up against.

    -

    Incremental Transformation

    +

    Incremental Transformation

    Listing 51.4 contains three functions; each concatenates an additional rotation around one of the three axes to an existing rotation. To improve performance, only the matrix entries that are affected in a rotation around each particular axis are recalculated (all but four of the entries in a single-axis rotation matrix are either 0 or 1, as shown in Chapter 50). This cuts the number of floating-point multiplies from the 64 required for the multiplication of two 4x4 matrices to just 12, and floating point adds from 48 to 6.

    Be aware that Listing 51.4 performs an incremental rotation on top of whatever rotation is already in the matrix. The cube may already have been turned left, right, up, down, and sideways; regardless, Listing 51.4 just tacks the specified rotation onto whatever already exists. In this way, the object-to-world transformation matrix contains a history of all the rotations ever specified by the user, concatenated one after another onto the original matrix. Potential loss of precision is a problem associated with using such an approach to represent a very long concatenation of transformations, especially with fixed-point arithmetic; that’s not a problem for us yet, but we’ll run into it eventually.

    -

    LISTING 51.4 L51-4.C

    +

    LISTING 51.4 L51-4.C

     /* Routines to perform incremental rotations around the three axes */
     #include <math.h>
    @@ -108,7 +101,7 @@
        XformToChange[0][2] = Temp02; XformToChange[1][0] = Temp10;
        XformToChange[1][1] = Temp11; XformToChange[1][2] = Temp12;
     }
    -
    +


    @@ -127,10 +120,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/51-06.html b/51-06.html index 7aa35e9..486f04b 100644 --- a/51-06.html +++ b/51-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sneakers in Space - - @@ -37,7 +30,7 @@


    -

    LISTING 51.5 POLYGON.H

    +

    LISTING 51.5 POLYGON.H

     /* POLYGON.H: Header file for polygon-filling code, also includes a number of
        useful items for 3D animation. */
    @@ -121,13 +114,13 @@ extern void AppendRotationY(double XformToChange[4][4], double Angle);
     extern void AppendRotationZ(double XformToChange[4][4], double Angle);
     extern int DisplayedPage, NonDisplayedPage;
     extern struct Rect EraseRect[];
    -
    + -

    A Note on Rounding Negative Numbers

    +

    A Note on Rounding Negative Numbers

    In the previous chapter, I added 0.5 and truncated in order to round values from floating-point to integer format. Here, in Listing 51.2, I’ve switched to adding 0.5 and using the floor() function. For positive values, the two approaches are equivalent; for negative values, only the floor() approach works properly.

    -

    Object Representation

    +

    Object Representation

    Each object consists of a list of vertices and a list of faces, with the vertices of each face defined by pointers into the vertex list; this allows each vertex to be transformed exactly once, even though several faces may share a single vertex. Each object contains the vertices not only in their original, untransformed state, but in three other forms as well: transformed to view space, transformed and projected to screen space, and converted to screen coordinates. Earlier, we saw that it can be convenient to store the screen coordinates within the object, so that if the object hasn’t moved with respect to the viewer, it can be redrawn without the need for recalculation, but why bother storing the view and screen space forms of the vertices as well?

    @@ -150,10 +143,6 @@ extern struct Rect EraseRect[];
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-01.html b/52-01.html index d3e191b..4fcf3a5 100644 --- a/52-01.html +++ b/52-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,10 +30,10 @@


    -

    Chapter 52
    +

    Chapter 52
    Fast 3-D Animation: Meet X-Sharp

    -

    The First Iteration of a Generalized 3-D Animation Package

    +

    The First Iteration of a Generalized 3-D Animation Package

    Across the lake from Vermont, a few miles into upstate New York, the Ausable River has carved out a fairly impressive gorge known as “Ausable Chasm.” Impressive for the East, anyway; you might think of it as the poor man’s Grand Canyon. Some time back, I did the tour with my wife and five-year-old, and it was fun, although I confess that I didn’t loosen my grip on my daughter’s hand until we were on the bus and headed for home; that gorge is deep, and the railings tend to be of the single-bar, rusted-out variety.

    @@ -52,7 +45,7 @@

    In our 3-D animation work so far, we’ve used floating-point arithmetic. Floating-point arithmetic—even with a floating-point processor but especially without one—is the microcomputer animation equivalent of working in a school bus: It takes forever to do anything, and you just know you’re never going to accomplish as much as you want to. In this chapter, we’ll address fixed-point arithmetic, which will give us an instant order-of-magnitude performance boost. We’ll also give our 3-D animation code a much more powerful and extensible framework, making it easy to add new and different sorts of objects. Taken together, these alterations will let us start to do some really interesting real-time animation.

    -

    This Chapter’s Demo Program

    +

    This Chapter’s Demo Program

    Three-dimensional animation is a complicated business, and it takes an astonishing amount of functionality just to get off the launching pad: page flipping, polygon filling, clipping, transformations, list management, and so forth. I’ve been building toward a critical mass of animation functionality over the course of this book, and this chapter’s code builds on the code from no fewer than five previous chapters. The code that’s required in order to link this chapter’s animation demo program is the following:

    @@ -87,10 +80,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-02.html b/52-02.html index 7c37051..ab01bcf 100644 --- a/52-02.html +++ b/52-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,7 +30,7 @@


    -

    LISTING 52.1 L52-1.C

    +

    LISTING 52.1 L52-1.C

     /* 3-D animation program to rotate 12 cubes. Uses fixed point. All C code
        tested with Borland C++ in C compilation mode and the small model. */
    @@ -113,9 +106,9 @@ void main() {
        int86(0x10, &regset, &regset);
        exit(1);
     }
    -
    + -

    LISTING 52.2 L52-2.C

    +

    LISTING 52.2 L52-2.C

     /* Transforms all vertices in the specified polygon-based object into view
        space, then perspective projects them to screen space and maps them to screen
    @@ -160,7 +153,7 @@ void XformAndProjectPObject(PObject * ObjectToXform)
                 DOUBLE_TO_FIXED(0.5)) >> 16))) + SCREEN_HEIGHT/2;
        }
     }
    -
    +


    @@ -179,10 +172,6 @@ void XformAndProjectPObject(PObject * ObjectToXform)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-03.html b/52-03.html index 743df15..51df732 100644 --- a/52-03.html +++ b/52-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,7 +30,7 @@


    -

    LISTING 52.3 L52-3.C

    +

    LISTING 52.3 L52-3.C

     /* Routines to perform incremental rotations around the three axes. */
     
    @@ -123,9 +116,9 @@ void AppendRotationZ(Xform XformToChange, double Angle)
        XformToChange[0][2] = Temp02; XformToChange[1][0] = Temp10;
        XformToChange[1][1] = Temp11; XformToChange[1][2] = Temp12;
     }
    -
    + -

    LISTING 52.4 L52-4.C

    +

    LISTING 52.4 L52-4.C

     /* Fixed point matrix arithmetic functions. */
     
    @@ -165,9 +158,9 @@ void ConcatXforms(Xform SourceXform1, Xform SourceXform2,
                    SourceXform1[i][3];
        }
     }
    -
    + -

    LISTING 52.5 L52-5.C

    +

    LISTING 52.5 L52-5.C

     /* Set up basic data that needs to be in fixed point, to avoid data
        definition hassles. */
    @@ -196,7 +189,7 @@ void InitializeFixedPoint()
           CubeVerts[i].Z = INT_TO_FIXED(IntCubeVerts[i].Z);
        }
     }
    -
    +


    @@ -215,10 +208,6 @@ void InitializeFixedPoint()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-04.html b/52-04.html index 2fa1f63..e026161 100644 --- a/52-04.html +++ b/52-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,7 +30,7 @@


    -

    LISTING 52.6 L52-6.C

    +

    LISTING 52.6 L52-6.C

     /* Rotates and moves a polygon-based object around the three axes.
        Movement is implemented only along the Z axis currently. */
    @@ -68,9 +61,9 @@ void RotateAndMovePObject(PObject * ObjectToMove)
           ObjectToMove->RecalcXform = 1;
        }
     }
    -
    + -

    LISTING 52.7 L52-7.C

    +

    LISTING 52.7 L52-7.C

     /* Draws all visible faces in specified polygon-based object. Object must have
        previously been transformed and projected, so that ScreenVertexList array is
    @@ -137,7 +130,7 @@ void DrawPObject(PObject * ObjectToXform)
           }
        }
     }
    -
    +


    @@ -156,10 +149,6 @@ void DrawPObject(PObject * ObjectToXform)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-05.html b/52-05.html index 69e8736..c8355da 100644 --- a/52-05.html +++ b/52-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,7 +30,7 @@


    -

    LISTING 52.8 L52-8.C

    +

    LISTING 52.8 L52-8.C

     /* Initializes the cubes and adds them to the object list. */
     
    @@ -161,7 +154,7 @@ void InitializeCubes()
           ObjectList[NumObjects++] = (Object *)WorkingCube;
        }
     }
    -
    +


    @@ -180,10 +173,6 @@ void InitializeCubes()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-06.html b/52-06.html index 9106e05..91fd706 100644 --- a/52-06.html +++ b/52-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,7 +30,7 @@


    -

    LISTING 52.9 L52-9.ASM

    +

    LISTING 52.9 L52-9.ASM

     ; 386-specific fixed point multiply and divide.
     ;
    @@ -112,9 +105,9 @@ FDP3:   mov     edx,eax         ;return result in DX:AX; fractional
             ret
     _FixedDiv       endp
             end
    -
    + -

    LISTING 52.10 POLYGON.H

    +

    LISTING 52.10 POLYGON.H

     /* POLYGON.H: Header file for polygon-filling code, also includes
        a number of useful items for 3-D animation. */
    @@ -214,7 +207,7 @@ extern int NumObjects;
     extern Xform WorldViewXform;
     extern Object *ObjectList[];
     extern Point3 CubeVerts[];
    -
    +


    @@ -233,10 +226,6 @@ extern Point3 CubeVerts[];
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-07.html b/52-07.html index b8eb3c3..2fa652f 100644 --- a/52-07.html +++ b/52-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -37,7 +30,7 @@


    -

    A New Animation Framework: X-Sharp

    +

    A New Animation Framework: X-Sharp

    Listings 52.1 through 52.10 shown earlier represent not merely faster animation in library form, but also a nearly complete, extensible, data-driven animation framework. Whereas much of the earlier animation code I’ve presented in this book was hardwired to demonstrate certain concepts, this chapter’s code is intended to serve as the basis for a solid animation package. Objects are stored, in their entirety, in customizable structures; new structures can be devised for new sorts of objects. Drawing, preparing for drawing, and moving are all vectored functions, so that variations such as shading or texturing, or even radically different sorts of graphics objects, such as scaled bitmaps, could be supported. The cube initialization is entirely data driven; more or different cubes, or other sorts of convex polyhedrons, could be added by simply changing the initialization data in Listing 52.8.

    @@ -47,7 +40,7 @@

    I’m working toward a goal in this last section of the book, and there are many lessons to be learned and stories to be told along the way. So as X-Sharp grows, you’ll find its evolving implementations in the chapter subdirectories on the listings diskette. This chapter’s subdirectory, for example, contains the self-extracting archive file XSHARP14.EXE, (to extract its contents you simply run it as though it were a program) and the code in that archive is the code I’m speaking of specifically in this chapter, with all the limitations mentioned above. Chapter 53’s subdirectory, however, contains the file XSHARP15.EXE, which is the next step in the evolution of X-Sharp, and it is the version that I’ll be specifically talking about in that chapter. Later chapters will have their own implementations in their respective chapter subdirectories, in files of the form XSHARPxx.EXE, where xx is an ascending number indicating the version. The final and most recent X-Sharp version will be present in its own subdirectory called XSHARP22. If you’re intending to use X-Sharp in a real project, use the most recent version to be sure that you avail yourself of all new features and bug fixes.

    -

    Three Keys to Realtime Animation Performance

    +

    Three Keys to Realtime Animation Performance

    As of the previous chapter, we were at the point where we could rotate, move, and draw a solid cube in real time. Not too shabby...but the code I’m presenting in this chapter goes a bit further, rotating 12 solid cubes at an update rate of about 15 frames per second (fps) on a 20 MHz 386 with a slow VGA. That’s 12 transformation matrices, 72 polygons, and 96 vertices being handled in real time; not Star Wars, granted, but a giant step beyond a single cube. Run the program if you get a chance; you may be surprised at just how effective this level of animation is. I’d like to point out, in case anyone missed it, that this is fully general 3-D. I’m not using any shortcuts or tricks, like prestoring coordinates or pregenerating bitmaps; if you were to feed in different rotations or vertices, the animation would change accordingly.

    @@ -68,10 +61,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/52-08.html b/52-08.html index 47f6e5c..a89e43f 100644 --- a/52-08.html +++ b/52-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Fast 3-D Animation: Meet X-Sharp - - @@ -45,13 +38,13 @@

    Just for fun, I reimplemented the animation of Listings 52.1 through 52.10 with floating-point instructions. Together, the preceeding optimizations improve the performance of the entire animation—including drawing time and overhead, and not just math—by more than ten times over the code that uses the floating-point emulator. Amazing what one can accomplish with a few dozen lines of assembly and a switch in number format, isn’t it? Note that no assembly code other than the native 386 multiply and divide is used in Listings 52.1 through 52.10, although the polygon fill code is of course mostly in assembly; we’ve achieved 12 cubes animated at 15 fps while doing the 3-D work almost entirely in Borland C++, and we’re still doing sine and cosine via the floating-point emulator. Happily, we’re still nowhere near the upper limit on the animation potential of the PC.

    -

    Drawbacks

    +

    Drawbacks

    The techniques we’ve used to turbocharge 3-D animation are very powerful, but there’s a dark side to them as well. Obviously, native 386 instructions won’t work on 8088 and 286 machines. That’s rectifiable; equivalent multiplication and division routines could be implemented for real mode and performance would still be reasonable. It sure is nice to be able to plug in a 32-bit IMUL or DIV and be done with it, though. More importantly, 32-bit fixed-point arithmetic has limitations in range and accuracy. Points outside a 64Kx64Kx64K space can’t be handled, imprecision tends to creep in over the course of multiple matrix concatenations, and it’s quite possible to generate the dreaded divide by 0 interrupt if Z coordinates with absolute values less than one are used.

    I don’t have space to discuss these issues in detail, but here are some brief thoughts: The working 64Kx64Kx64K fixed-point space can be paged into a larger virtual space. Imprecision of a pixel or two rarely matters in terms of display quality, and deterioration of concatenated rotations can be corrected by restoring orthogonality, for example by periodically calculating one row of the matrix as the cross-product of the other two (forcing it to be perpendicular to both). Alternatively, transformations can be calculated from scratch each time an object or the viewer moves, so there’s no chance for cumulative error. 3-D clipping with a front clip plane of -1 or less can prevent divide overflow.

    -

    Where the Time Goes

    +

    Where the Time Goes

    The distribution of execution time in the animation code is no longer wildly biased toward transformation, but sine and cosine are certainly still sucking up cycles. Likewise, the overhead in the calls to FixedMul() and FixedDiv() is costly. Much of this is correctable with a little carefully crafted assembly language and a lookup table; I’ll provide that shortly.

    @@ -74,10 +67,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/53-01.html b/53-01.html index 3a9c0d8..ce52981 100644 --- a/53-01.html +++ b/53-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - @@ -37,10 +30,10 @@


    -

    Chapter 53
    +

    Chapter 53
    Raw Speed and More

    -

    The Naked Truth About Speed in 3-D Animation

    +

    The Naked Truth About Speed in 3-D Animation

    Years ago, this friend of mine—let’s call him Bert—went to Hawaii with three other fellows to celebrate their graduation from high school. This was an unchaperoned trip, and they behaved pretty much as responsibly as you’d expect four teenagers to behave, which is to say, not; there’s a story about a rental car that, to this day, Bert can’t bring himself to tell. They had a good time, though, save for one thing: no girls.

    @@ -52,7 +45,7 @@

    And with that, we come to this chapter’s topics: raw speed and hidden surfaces.

    -

    Raw Speed, Part 1: Assembly Language

    +

    Raw Speed, Part 1: Assembly Language

    I would like to state, here and for the record, that I am not an assembly language fanatic. Frankly, I prefer programming in C; assembly language is hard work, and I can get a whole lot more done with fewer hassles in C. However, I am a performance fanatic, performance being defined as having programs be as nimble as possible in those areas where the user wants fast response. And, in the course of pursuing performance, there are times when a little assembly language goes a long way.

    @@ -77,10 +70,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/53-02.html b/53-02.html index 223a28d..a14b743 100644 --- a/53-02.html +++ b/53-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - @@ -37,7 +30,7 @@


    -

    LISTING 53.1 FIXED.ASM

    +

    LISTING 53.1 FIXED.ASM

     ; 386-specific fixed point routines.
     ; Tested with TASM
    @@ -426,7 +419,7 @@ popbp;restore stack frame
     ret
     -ConcatXformsendp
     end
    -
    +


    @@ -445,10 +438,6 @@ end
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/53-03.html b/53-03.html index a7cc146..b57b498 100644 --- a/53-03.html +++ b/53-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - @@ -37,7 +30,7 @@


    -

    Raw Speed, Part II: Look it Up

    +

    Raw Speed, Part II: Look it Up

    It’s a funny thing about Turbo Profiler: Time spent in the Borland C++ 80x87 emulator doesn’t show up directly anywhere that I can see in the timing results. The only way to detect it is by way of the line that reports what percent of total time is represented by all the areas that were profiled; if you’re profiling all areas, whatever’s not explicitly accounted for seems to be the floating-point emulator time. This quirk fooled me for a while, leading me to think sine and cosine weren’t major drags on performance, because the sin() and cos() functions spend most of their time in the emulator, and that time doesn’t show up in Turbo Profiler’s statistics on those functions. Once I figured out what was going on, it turned out that not only were sin() and cos() major drags, they were taking up over half the total execution time by themselves.

    @@ -45,7 +38,7 @@

    FIXED.ASM (Listing 53.1) speeds X-Sharp up quite a bit, and it changes the performance balance a great deal. When we started out with 3-D animation, calculation time was the dragon we faced; more than 90 percent of the total time was spent doing matrix and projection math. Additional optimizations in the area of math could still be made (using 32-bit multiplies in the backface-removal code, for example), but fixed-point math, the sine and cosine lookup, and selective assembly optimizations have done a pretty good job already. The bulk of the time taken by X-Sharp is now spent drawing polygons, drawing rectangles (to erase objects), and waiting for the page to flip. In other words, we’ve slain the dragon of 3-D math, or at least wounded it grievously; now we’re back to the dragon of polygon filling. We’ll address faster polygon filling soon, but for the moment, we have more than enough horsepower to have some fun with. First, though, we need one more feature: hidden surfaces.

    -

    Hidden Surfaces

    +

    Hidden Surfaces

    So far, we’ve made a number of simplifying assumptions in order to get the animation to look good; for example, all objects must currently be convex polyhedrons. What’s more, right now, objects can never pass behind or in front of each other. What that means is that it’s time to have a look at hidden surfaces.

    @@ -57,9 +50,8 @@

    Listing 53.2 shows X-Sharp file OLIST.C, which includes the key routines for depth sorting. Objects are now stored in a linked list. The initial, empty list, created by InitializeObjectList(), consists of a sentinel entry at either end, one at the farthest possible z coordinate, and one at the nearest. New entries are inserted by AddObject() in z-sorted order. Each time the objects are moved, before they’re drawn at their new locations, SortObjects() is called to Z-sort the object list, so that drawing will proceed from back to front. The Z-sorting is done on the basis of the objects’ center points; a center-point field has been added to the object structure to support this, and the center point for each object is now transformed along with the vertices. That’s really all there is to depth sorting—and now we can have objects that overlap in X and Y.

    -


    - Figure 53.1
      Why back-to-front sorting doesn’t always work properly.

    +


    + Figure 53.1
      Why back-to-front sorting doesn’t always work properly.


    @@ -78,10 +70,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/53-04.html b/53-04.html index 9afd594..5f3d5ec 100644 --- a/53-04.html +++ b/53-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Raw Speed and More - - @@ -37,7 +30,7 @@


    -

    LISTING 53.2 OLIST.C

    +

    LISTING 53.2 OLIST.C

     /* Object list-related functions. */
     #include <stdio.h>
    @@ -120,15 +113,15 @@ void SortObjects()
           }
        }
     }
    -
    + -

    Rounding

    +

    Rounding

    FIXED.ASM contains the equate ROUNDING-ON. When this equate is 1, the results of multiplications and divisions are rounded to the nearest fixed-point values; when it’s 0, the results are truncated. The difference between the results produced by the two approaches is, at most, 2-16; you wouldn’t think that would make much difference, now, would you? But it does. When the animation is run with rounding disabled, the cubes start to distort visibly after a few minutes, and after a few minutes more they look like they’ve been run over. In contrast, I’ve never seen any significant distortion with rounding on, even after a half-hour or so. I think the difference with rounding is not that it’s so much more accurate, but rather that the errors are evenly distributed; with truncation, the errors are biased, and biased errors become very visible when they’re applied to right-angle objects. Even with rounding, though, the errors will eventually creep in, and reorthogonalization will become necessary at some point.

    The performance cost of rounding is small, and the benefits are highly visible. Still, truncation errors become significant only when they accumulate over time, as, for example, when rotation matrices are repeatedly concatenated over the course of many transformations. Some time could be saved by rounding only in such cases. For example, division is performed only in the course of projection, and the results do not accumulate over time, so it would be reasonable to disable rounding for division.

    -

    Having a Ball

    +

    Having a Ball

    So far in our exploration of 3-D animation, we’ve had nothing to look at but triangles and cubes. It’s time for something a little more visually appealing, so the demonstration program now features a 72-sided ball. What’s particularly interesting about this ball is that it’s created by the GENBALL.C program in the BALL subdirectory of X-Sharp, and both the size of the ball and the number of bands of faces are programmable. GENBALL.C spits out to a file all the arrays of vertices and faces needed to create the ball, ready for inclusion in INITBALL.C. True, if you change the number of bands, you must change the Colors array in INITBALL.C to match, but that’s a tiny detail; by and large, the process of generating a ball-shaped object is now automated. In fact, we’re not limited to ball-shaped objects; substitute a different vertex and face generation program for GENBALL.C, and you can make whatever convex polyhedron you want; again, all you have to do is change the Colors array correspondingly. You can easily create multiple versions of the base object, too; INITCUBE.C is an example of this, creating 11 different cubes.

    @@ -151,10 +144,6 @@ void SortObjects()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/54-01.html b/54-01.html index 942d02f..1b73f52 100644 --- a/54-01.html +++ b/54-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - @@ -37,14 +30,14 @@


    -

    Chapter 54
    +

    Chapter 54
    3-D Shading

    -

    Putting Realistic Surfaces on Animated 3-D Objects

    +

    Putting Realistic Surfaces on Animated 3-D Objects

    At the end of the previous chapter, X-Sharp had just acquired basic hidden-surface capability, and performance had been vastly improved through the use of fixed-point arithmetic. In this chapter, we’re going to add quite a bit more: support for 8088 and 80286 PCs, a general color model, and shading. That’s an awful lot to cover in one chapter (actually, it’ll spill over into the next chapter), so let’s get to it!

    -

    Support for Older Processors

    +

    Support for Older Processors

    To date, X-Sharp has run on only the 386 and 486, because it uses 32-bit multiply and divide instructions that sub-386 processors don’t support. I chose 32-bit instructions for two reasons: They’re much faster for 16.16 fixed-point arithmetic than any approach that works on the 8088 and 286; and they’re much easier to implement than any other approach. In short, I was after maximum performance, and I was perhaps just a little lazy.

    @@ -73,10 +66,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/54-02.html b/54-02.html index fd06a0d..fbcec4e 100644 --- a/54-02.html +++ b/54-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - @@ -37,7 +30,7 @@


    -

    LISTING 54.1 FIXED.ASM

    +

    LISTING 54.1 FIXED.ASM

     ; Fixed point routines.
     ; Tested with TASM
    @@ -898,7 +891,7 @@ endif ;USE386
     ret
     _ConcatXforms    endp
     end
    -
    +


    @@ -917,10 +910,6 @@ end
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/54-03.html b/54-03.html index 1434525..b1bf7d7 100644 --- a/54-03.html +++ b/54-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - @@ -37,31 +30,29 @@


    -

    Shading

    +

    Shading

    So far, the polygons out of which our animated objects have been built have had colors of fixed intensities. For example, a face of a cube might be blue, or green, or white, but whatever color it is, that color never brightens or dims. Fixed colors are easy to implement, but they don’t make for very realistic animation. In the real world, the intensity of the color of a surface varies depending on how brightly it is illuminated. The ability to simulate the illumination of a surface, or shading, is the next feature we’ll add to X-Sharp.

    The overall shading of an object is the sum of several types of shading components. Ambient shading is illumination by what you might think of as background light, light that’s coming from all directions; all surfaces are equally illuminated by ambient light, regardless of their orientation. Directed lighting, producing diffuse shading, is illumination from one or more specific light sources. Directed light has a specific direction, and the angle at which it strikes a surface determines how brightly it lights that surface. Specular reflection is the tendency of a surface to reflect light in a mirrorlike fashion. There are other sorts of shading components, including transparency and atmospheric effects, but the ambient and diffuse-shading components are all we’re going to deal with in X-Sharp.

    -

    Ambient Shading

    +

    Ambient Shading

    The basic model for both ambient and diffuse shading is a simple one. Each surface has a reflectivity between 0 and 1, where 0 means all light is absorbed and 1 means all light is reflected. A certain amount of light energy strikes each surface. The energy (intensity) of the light is expressed such that if light of intensity 1 strikes a surface with reflectivity 1, then the brightest possible shading is displayed for that surface. Complicating this somewhat is the need to support color; we do this by separating reflectance and shading into three components each—red, green, and blue—and calculating the shading for each color component separately for each surface.

    Given an ambient-light red intensity of IAred and a surface red reflectance Rred, the displayed red ambient shading for that surface, as a fraction of the maximum red intensity, is simply min(IAredx Rred, 1). The green and blue color components are handled similarly. That’s really all there is to ambient shading, although of course we must design some way to map displayed color components into the available palette of colors; I’ll do that in the next chapter. Ambient shading isn’t the whole shading picture, though. In fact, scenes tend to look pretty bland without diffuse shading.

    -

    Diffuse Shading

    +

    Diffuse Shading

    Diffuse shading is more complicated than ambient shading, because the effective intensity of directed light falling on a surface depends on the angle at which it strikes the surface. According to Lambert’s law, the light energy from a directed light source striking a surface is proportional to the cosine of the angle at which it strikes the surface, with the angle measured relative to a vector perpendicular to the polygon (a polygon normal), as shown in Figure 54.1. If the red intensity of directed light is IDred, the red reflectance of the surface is Rred, and the angle between the incoming directed light and the surface’s normal is theta, then the displayed red diffuse shading for that surface, as a fraction of the largest possible red intensity, is min (IDredxRredxcos(θ), 1).

    That’s easy enough to calculate—but seemingly slow. Determining the cosine of an angle can be sped up with a table lookup, but there’s also the task of figuring out the angle, and, all in all, it doesn’t seem that diffuse shading is going to be speedy enough for our purposes. Consider this, however: According to the properties of the dot product (denoted by the operator “•”, as shown in Figure 54.2), cos(q)=(v•w)/ |v| x |w| ), where v and w are vectors, q is the angle between v and w, and |v| is the length of v. Suppose, now, that v and w are unit vectors; that is, vectors exactly one unit long. Then the above equation reduces to cos(q)=v•w. In other words, we can calculate the cosine between N, the unit-normal vector (one-unit-long perpendicular vector) of a polygon, and L', the reverse of a unit vector describing the direction of a light source, with just three multiplies and two adds. (I’ll explain why the light-direction vector must be reversed later.) Once we have that, we can easily calculate the red diffuse shading from a directed light source as min(IDredxRredx(L'• N), 1) and likewise for the green and blue color components.

    -


    - Figure 54.1
      Illumination by a directed light source

    +


    + Figure 54.1
      Illumination by a directed light source

    -


    - Figure 54.2
      The dot product of two vectors.

    +


    + Figure 54.2
      The dot product of two vectors.

    The overall red shading for each polygon can be calculated by summing the ambient-shading red component with the diffuse-shading component from each light source, as in min((IAredxRred) + (IDred0xRredx(L0' • N)) + (IDred1xRredx(L1' • N)) +..., 1) where IDred0 and L0' are the red intensity and the reversed unit-direction vector, respectively, for spotlight 0. Listing 54.2 shows the X-Sharp module DRAWPOBJ.C, which performs ambient and diffuse shading. Toward the end, you will find the code that performs shading exactly as described by the above equation, first calculating the ambient red, green, and blue shadings, then summing that with the diffuse red, green, and blue shadings generated by each directed light source.

    @@ -82,10 +73,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/54-04.html b/54-04.html index ccb4324..ce074ed 100644 --- a/54-04.html +++ b/54-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - @@ -37,7 +30,7 @@


    -

    LISTING 54.2 DRAWPOBJ.C

    +

    LISTING 54.2 DRAWPOBJ.C

     /* Draws all visible faces in the specified polygon-based object. The object
        must have previously been transformed and projected, so that all vertex
    @@ -165,7 +158,7 @@ void DrawPObject(PObject * ObjectToXform)
           }
        }
     }
    -
    +


    @@ -184,10 +177,6 @@ void DrawPObject(PObject * ObjectToXform)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/54-05.html b/54-05.html index 225ec0c..6cc4683 100644 --- a/54-05.html +++ b/54-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Shading - - @@ -37,17 +30,15 @@


    -

    Shading: Implementation Details

    +

    Shading: Implementation Details

    In order to calculate the cosine of the angle between an incoming light source and a polygon’s unit normal, we must first have the polygon’s unit normal. This could be calculated by generating a cross-product on two polygon edges to generate a normal, then calculating the normal’s length and scaling to produce a unit normal. Unfortunately, that would require taking a square root, so it’s not a desirable course of action. Instead, I’ve made a change to X-Sharp’s polygon format. Now, the first vertex in a shaded polygon’s vertex list is the end-point of a unit normal that starts at the second point in the polygon’s vertex list, as shown in Figure 54.3. The first point isn’t one of the polygon’s vertices, but is used only to generate a unit normal. The second point, however, is a polygon vertex. Calculating the difference vector between the first and second points yields the polygon’s unit normal. Adding a unit-normal endpoint to each polygon isn’t free; each of those end-points has to be transformed, along with the rest of the vertices, and that takes time. Still, it’s faster than calculating a unit normal for each polygon from scratch.

    -


    - Figure 54.3
      The unit normal in the polygon data structure.

    +


    + Figure 54.3
      The unit normal in the polygon data structure.

    -


    - Figure 54.4
      The reversed light source vector.

    +


    + Figure 54.4
      The reversed light source vector.

    We also need a unit vector for each directed light source. The directed light sources I’ve implemented in X-Sharp are spotlights; that is, they’re considered to be point light sources that are infinitely far away. This allows the simplifying assumption that all light rays from a spotlight are parallel and of equal intensity throughout the displayed universe, so each spotlight can be represented with a single unit vector and a single intensity. The only trick is that in order to calculate the desired cos(theta) between the polygon unit normal and a spotlight’s unit vector, the direction of the spotlight’s unit vector must be reversed, as shown in Figure 54.4. This is necessary because the dot product implicitly places vectors with their start points at the same location when it’s used to calculate the cosine of the angle between two vectors. The light vector is incoming to the polygon surface, and the unit normal is outbound, so only by reversing one vector or the other will we get the cosine of the desired angle.

    @@ -70,10 +61,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/55-01.html b/55-01.html index abc4452..fca42fe 100644 --- a/55-01.html +++ b/55-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - @@ -37,10 +30,10 @@


    -

    Chapter 55
    +

    Chapter 55
    Color Modeling in 256-Color Mode

    -

    Pondering X-Sharp’s Color Model in an RGB State of Mind

    +

    Pondering X-Sharp’s Color Model in an RGB State of Mind

    Once she turned six, my daughter wanted some fairly sophisticated books read to her. Wind in the Willows. Little House on the Prairie. Pretty heady stuff for one so young, and sometimes I wondered how much of it she really understood. As an experiment, during one reading I stopped whenever I came to a word I thought she might not know, and asked her what it meant. One such word was “mulling.”

    @@ -58,7 +51,7 @@

    What does this anecdote tell us about the universe in which we live? Well, it certainly indicates that this universe is inhabited by at least one comedian and one good straight man. Beyond that, though, it can be construed as a parable about the difficulty of defining things properly; for example, consider the complications inherent in the definition of color on a 256-color display adapter such as the VGA. Coincidentally, VGA color modeling just happens to be this chapter’s topic, and the place to start is with color modeling in general.

    -

    A Color Model

    +

    A Color Model

    We’ve been developing X-Sharp for several chapters now. In the previous chapter, we added illumination sources and shading; that addition makes it necessary for us to have a general-purpose color model, so that we can display the gradations of color intensity necessary to render illuminated surfaces properly. In other words, when a bright light is shining straight at a green surface, we need to be able to display bright green, and as that light dims or tilts to strike the surface at a shallower angle, we need to be able to display progressively dimmer shades of green.

    @@ -66,15 +59,14 @@

    In the RGB model, a given color is modeled as the mix of specific fractions of full intensities of each of the three color primaries. For example, the brightest possible pure blue is 0.0*R, 0.0*G, 1.0*B. Half-bright cyan is 0.0*R, 0.5*G, 0.5*B. Quarter-bright gray is 0.25*R, 0.25*G, 0.25*B. You can think of RGB color space as being a cube, as shown in Figure 55.1, with any particular color lying somewhere inside or on the cube.

    -


    - Figure 55.1
      The RGB color cube.

    +


    + Figure 55.1
      The RGB color cube.

    RGB is good for modeling colors generated by light sources, because red, green, and blue are the additive primaries; that is, all other colors can be generated by mixing red, green, and blue light sources. They’re also the primaries for color computer displays, and the RGB model maps beautifully onto the display capabilities of 15- and 24-bpp display adapters, which tend to represent pixels as RGB combinations in display memory.

    How, then, are RGB colors represented in X-Sharp? Each color is represented as an RGB triplet, with eight bits each of red, green, and blue resolution, using the structure shown in Listing 55.1.

    -

    LISTING 55.1 L55-1.C

    +

    LISTING 55.1 L55-1.C

      typedef struct -ModelColor {
         unsigned char Red;   /* 255 = max red, 0 = no red */
    @@ -82,7 +74,7 @@
         unsigned char Blue;  /* 255 = max blue, 0 = no blue */
      } ModelColor;
     
    -
    +

    Here, each color is described by three color components—one each for red, green, and blue—and each primary color component is represented by eight bits. Zero intensity of a color component is represented by the value 0, and full intensity is represented by the value 255. This gives us 256 levels of each primary color component, and a total of 16,772,216 possible colors.

    @@ -109,10 +101,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/55-02.html b/55-02.html index 71e7896..3cd254c 100644 --- a/55-02.html +++ b/55-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - @@ -43,7 +36,7 @@

    One way to deal with the limited simultaneous color capabilities of the VGA is to build an application that uses only a subset of RGB space, then bias the VGA’s palette toward that subspace. This is the approach used in the DEMO1 sample program in X-Sharp; Listings 55.2 and 55.3 show the versions of InitializePalette() and ModelColorToColorIndex() that set up and perform the color mapping for DEMO1.

    -

    LISTING 55.2 L55-2.C

    +

    LISTING 55.2 L55-2.C

      /* Sets up the palette in mode X, to a 2-2-2 general R-G-B organization, with
         64 separate levels each of pure red, green, and blue. This is very good
    @@ -130,7 +123,7 @@
         sregset.es = DS; /* segment of array from which to load settings */
         int86x(0x10, &regset, &regset, &sregset); /* load the palette block */
      }
    -
    +


    @@ -149,10 +142,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/55-03.html b/55-03.html index c3d6105..055544e 100644 --- a/55-03.html +++ b/55-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - @@ -37,7 +30,7 @@


    -

    LISTING 55.3 L55-3.C

    +

    LISTING 55.3 L55-3.C

      /* Converts a model color (a color in the RGB color cube, in the current
         color model) to a color index for mode X. Pure primary colors are
    @@ -62,7 +55,7 @@
               ((Color->Blue & 0xC0) >> 6));
      }
     
    -
    +

    In DEMO1, three-quarters of the palette is set up with 64 intensity levels of each of the three pure primary colors (red, green, and blue), and then most drawing is done with only pure primary colors. The resulting rendering quality is very good because there are so many levels of each primary.

    @@ -74,7 +67,7 @@

    The sad truth is that the VGA’s 256-color palette is an inadequate resource for general RGB shading. The good news is that clever workarounds can make VGA graphics look nearly as good as 24-bpp graphics; but the burden falls on you, the programmer, to design your applications and color mapping to compensate for the VGA’s limitations. To experiment with a different 256-color model in X-Sharp, just change InitializePalette() to set up the desired palette and ModelColorToColorIndex() to map 24-bit RGB triplets into the palette you’ve set up. It’s that simple, and the results can be striking indeed.

    -

    A Bonus from the BitMan

    +

    A Bonus from the BitMan

    Finally, a note on fast VGA text, which came in from a correspondent who asked to be referred to simply as the BitMan. The BitMan passed along a nifty application of the VGA’s under-appreciated write mode 3 that is, under the proper circumstances, the fastest possible way to draw text in any 16-color VGA mode.

    @@ -82,9 +75,8 @@

    Solid text is useful for drawing menus, text areas, and the like; basically, it can be used whenever you want to display text on a solid-color background. The obvious way to implement solid text is to fill the rectangle representing the background box, then draw transparent text on top of the background box. However, there are two problems with doing solid text this way. First, there’s some flicker, because for a little while the box is there but the text hasn’t yet arrived. More important is that the background-followed-by-foreground approach accesses display memory three times for each byte of font data: once to draw the background box, once to read display memory to load the latches, and once to actually draw the font pattern. Display memory is incredibly slow, so we’d like to reduce the number of accesses as much as possible. With the BitMan’s approach, we can reduce the number of accesses to just one per font byte, and eliminate flicker, too.

    -


    - Figure 55.2
      Drawing solid text.

    +


    + Figure 55.2
      Drawing solid text.

    The keys to fast solid text are the latches and write mode 3. The latches, as you may recall from earlier discussions in this book, are four internal VGA registers that hold the last bytes read from the VGA’s four planes; every read from VGA memory loads the latches with the values stored at that display memory address across the four planes. Whenever a write is performed to VGA memory, the latches can provide some, none, or all of the bits written to memory, depending on the bit mask, which selects between the latched data and the drawing data on a bit-by-bit basis. The latches solve half our problem; we can fill the latches with the background color, then use them to draw the background box. The trick now is drawing the text pixels in the foreground color at the same time.

    @@ -105,10 +97,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/55-04.html b/55-04.html index 07143c9..005b6d1 100644 --- a/55-04.html +++ b/55-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Color Modeling in 256-Color Mode - - @@ -39,15 +32,14 @@

    This is where it gets a little complicated. In write mode 3 (which incidentally is not available on the EGA), each byte value that the CPU writes to the VGA does not get written to display memory. Instead, it turns into the bit mask. (Actually, it’s ANDed with the Bit Mask register, and the result becomes the bit mask, but we’ll leave the Bit Mask register set to 0xFF, so the CPU value will become the bit mask.) The bit mask selects, on a bit-by-bit basis, between the data in the latches for each plane (the previously loaded background color, in this case) and the foreground color. Where does the foreground color come from, if not from the CPU? From the Set/Reset register, as shown in Figure 55.3. Thus, each byte written by the CPU (font data, presumably) selects foreground or background color for each of eight pixels, all done with a single write to display memory.

    -


    - Figure 55.3
      The data path in write mode 3.

    +


    + Figure 55.3
      The data path in write mode 3.

    I know this sounds pretty esoteric, but think of it this way: The latches hold the background color in a form suitable for writing eight background pixels (one full byte) at a pop. Write mode 3 allows each CPU byte to punch holes in the background color provided by the latches, holes through which the foreground color from the Set/Reset register can flow. The result is that a single write draws exactly the combination of foreground and background pixels described by each font byte written by the CPU. It may help to look at Listing 55.4, which shows The BitMan’s technique in action. And yes, this technique is absolutely worth the trouble; it’s about three times faster than the fill-then-draw approach described above, and about twice as fast as transparent text. So far as I know, there is no faster way to draw text on a VGA.

    It’s important to note that the BitMan’s technique only works on full bytes of display memory. There’s no way to clip to finer precision; the background color will inevitably flood all of the eight destination pixels that aren’t selected as foreground pixels. This makes The BitMan’s technique most suitable for monospaced fonts with characters that are multiples of eight pixels in width, and for drawing to byte-aligned addresses; the technique can be used in other situations, but is considerably more difficult to apply.

    -

    LISTING 55.4 L55-4.ASM

    +

    LISTING 55.4 L55-4.ASM

      ; Demonstrates drawing solid text on the VGA, using the BitMan’s write mode
      ; 3-based, one-pass technique.
    @@ -218,7 +210,7 @@ DrawCharLoop:                    ;draw all lines of the character
             ret
     DrawTextString   endp
             end   start
    -
    +


    @@ -237,10 +229,6 @@ DrawTextString endp
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/56-01.html b/56-01.html index ec81518..53b0ef1 100644 --- a/56-01.html +++ b/56-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pooh and the Space Station - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pooh and the Space Station - - @@ -37,10 +30,10 @@


    -

    Chapter 56
    +

    Chapter 56
    Pooh and the Space Station

    -

    Using Fast Texture Mapping to Place Pooh on a Polygon

    +

    Using Fast Texture Mapping to Place Pooh on a Polygon

    So, here’s where Winnie the Pooh lives: in a space station orbiting Saturn. No, really; I have it straight from my daughter, and an eight-year-old wouldn’t make up something that important, would she? One day she wondered aloud, “Where is the Hundred Acre Wood, exactly?” and before I could give one of those boring parental responses about how it was imaginary—but A.A. Milne probably imagined it to be somewhere near London—my daughter announced that the Hundred Acre Wood was in a space station orbiting Saturn, and there you have it.

    @@ -54,21 +47,19 @@

    The rest is history.

    -

    Principles of Quick-and-Dirty Texture Mapping

    +

    Principles of Quick-and-Dirty Texture Mapping

    The key to our texture-mapping approach will be to quickly determine what pixel value to draw for each pixel in the transformed destination polygon. These polygon pixel values will be determined by mapping each destination pixel in the transformed polygon back to the image bitmap, via a reverse transformation, and seeing what color resides at the corresponding location in the image bitmap, as shown in Figure 56.1. It might seem more intuitive to map pixels the other way, from the image bitmap to the transformed polygon, but in fact it’s crucial that the mapping proceed backward from the destination to avoid gaps in the final image. With the approach of finding the right value for each destination pixel in turn, via a backward mapping, there’s no way we can miss any destination pixels. On the other hand, with the forward-mapping method, some destination pixels may be skipped or double-drawn, because this is not necessarily a one-to-one or one-to-many mapping. Although we’re not going to take advantage of it now, mapping back to the source makes it possible to average several neighboring image pixels together to calculate the value for each destination pixel; that is, to antialias the image. This can greatly improve texture quality, although it is slower.

    -


    - Figure 56.1
      Using reverse transformation to find the source pixel color.

    +


    + Figure 56.1
      Using reverse transformation to find the source pixel color.

    -

    Mapping Textures Made Easy

    +

    Mapping Textures Made Easy

    To understand how we’re going to map textures, consider Figure 56.2, which maps a bitmapped image directly onto an untransformed polygon. Here, we simply map the origin of the polygon’s untransformed coordinate system somewhere within the image, then map the vertices to the corresponding image pixels. (For simplicity, I’ll assume in this discussion that the polygon’s coordinate system is in units of pixels, but scaling images to polygons is eminently doable. This will become clearer when we look at mapping images onto transformed polygons, next.) Mapping the image to the polygon is then a simple matter of stepping one scan line at a time in both the image and the polygon, each time advancing the X coordinates of the edges according to the slopes of the lines, just as is normally done when filling a polygon. Since the polygon is untransformed, the stepping is identical in both the image and the polygon, and the pixel mapping is one-to-one, so the appropriate part of each scan line of the image can simply be block copied to the destination.

    -


    - Figure 56.2
      Mapping a texture onto an untransformed polygon.

    +


    + Figure 56.2
      Mapping a texture onto an untransformed polygon.

    Now, matters get more complicated. What if the destination polygon is rotated in two dimensions? We no longer have a neat direct mapping from image scan lines to destination polygon scan lines. We still want to draw across each destination scan line, but the proper source pixels for each destination scan line may now track across the source bitmap at an angle, as shown in Figure 56.3. What can we do?

    @@ -91,10 +82,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/56-02.html b/56-02.html index 327b019..d87f427 100644 --- a/56-02.html +++ b/56-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pooh and the Space Station - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pooh and the Space Station - - @@ -39,23 +32,20 @@

    Ah, but what is an “equivalent amount”? Think of it this way. If a destination edge is 100 scan lines high, it will be stepped 100 times. Then, we’ll divide the SourceXWidth and SourceYHeight lengths of the source edge by 100, and add those amounts to the source edge’s coordinates each time the destination is stepped one scan line. Put another way, we have, as usual, arranged things so that in the destination polygon we step DestYHeight times, where DestYHeight is the height of the destination edge. The this approach arranges to step the source image edge DestYHeight times also, to match what the destination is doing.

    -


    - Figure 56.3
      Mapping a texture onto a 2-D rotated polygon.

    +


    + Figure 56.3
      Mapping a texture onto a 2-D rotated polygon.

    Now we’re able to track the coordinates of the polygon edges through the source image in tandem with the destination edges. Stepping across each destination scan line uses precisely the same technique, as shown in Figure 56.4. In the destination, we step DestXWidth times across each scan line of the polygon, once for each pixel on the scan line. (DestXWidth is the horizontal distance between the two edges being scanned on any given scan line.) To match this, we divide SourceXWidth and SourceYHeight (the lengths of the scan line in the source image, as determined by the source edge points we’ve been tracking, as just described) by the width of the destination scan line, DestXWidth, to produce SourceXStep and SourceYStep. Then, we just step DestXWidth times, adding SourceXStep and SourceYStep to SourceX and SourceY each time, and choose the nearest image pixel to (SourceX,SourceY) to copy to (DestX, DestY). (Note that the names used above, such as SourceXWidth, are used for descriptive purposes, and don’t necessarily correspond to the actual variable names used in Listing 56.2.)

    That’s a workable approach for 2-D rotated polygons—but what about 3-D rotated polygons, where the visible dimensions of the polygon can vary with 3-D rotation and perspective projection? First, I’d like to make it clear that texture mapping takes place from the source image to the destination polygon after the destination polygon is projected to the screen. That is, the image will be mapped after the destination polygon is in its final, drawable form. Given that, it should be apparent that the above approach automatically compensates for all changes in the dimensions of a polygon. You see, this approach divides source edges and scan lines into however many steps the destination polygon requires. If the destination polygon is much narrower than the source polygon, as a result of 3-D rotation and perspective projection, we just end up taking bigger steps through the source image and skipping a lot of source image pixels, as shown in Figure 56.5. The upshot is that the above approach handles all transformations and projections effortlessly. It could also be used to scale source images up to fit in larger polygons; all that’s needed is a list of where the polygon’s vertices map into the source image, and everything else happens automatically. In fact, mapping from any polygonal area of a bitmap to any destination polygon will work, given only that the two polygons have the same number of vertices.

    -


    - Figure 56.4
      Mapping a horizontal destination scan line back to the source image.

    +


    + Figure 56.4
      Mapping a horizontal destination scan line back to the source image.

    -


    - Figure 56.5
      Mapping a texture onto a narrower polygon.

    +


    + Figure 56.5
      Mapping a texture onto a narrower polygon.

    -

    Notes on DDA Texture Mapping
    +

    Notes on DDA Texture Mapping

    That’s all there is to quick-and-dirty texture mapping. This technique basically uses a two-stage digital differential analyzer (DDA) approach to step through the appropriate part of the source image in tandem with the normal scan-line stepping through the destination polygon, so I’ll call it “DDA texture mapping.” It’s worth noting that there is no need for any trigonometric functions at all, and only two divides are required per scan line.

    @@ -85,10 +75,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/56-03.html b/56-03.html index d5f9576..18f9eb9 100644 --- a/56-03.html +++ b/56-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Pooh and the Space Station - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Pooh and the Space Station - - @@ -37,7 +30,7 @@


    -

    Fast Texture Mapping: An Implementation

    +

    Fast Texture Mapping: An Implementation

    As you might expect, I’ve implemented DDA texture mapping in X-Sharp, and the changes are reflected in the X-Sharp archive in this chapter’s subdirectory on the listings disk. Listing 56.1 shows the new header file entries, and Listing 56.2 shows the actual texture-mapped polygon drawer. The set-pixel routine that Listing 56.2 calls is a slight modification of the Mode X set-pixel routine from Chapter 47. In addition, INITBALL.C has been modified to create three texture-mapped polygons and define the texture bitmaps, and modifications have been made to allow the user to flip the axis of rotation. You will of course need the complete X-Sharp library to see texture mapping in action, but Listings 56.1 and 56.2 are the actual texture mapping code in its entirety.

    @@ -49,7 +42,7 @@ -

    LISTING 56.1 L56-1.C

    +

    LISTING 56.1 L56-1.C

     /* New header file entries related to texture-mapped polygons */
     
    @@ -95,9 +88,9 @@ typedef struct {
                             normal endpoint in VertNums) */
     } Face;
     extern void DrawTexturedPolygon(PointListHeader *, Point *, TextureMap *);
    -
    + -

    LISTING 56.2 L56-2.C

    +

    LISTING 56.2 L56-2.C

     /* Draws a bitmap, mapped to a convex polygon (draws a texture-mapped polygon).
        “Convex” means that every horizontal line drawn through the polygon at any
    @@ -351,7 +344,7 @@ void ScanOutLine(EdgeScan * LeftEdge, EdgeScan * RightEdge)
           SourceY += SourceYStep;
        }
     }
    -
    +

    No matter how you slice it, DDA texture mapping beats boring, single-color polygons nine ways to Sunday. The big downside is that it’s much slower than a normal polygon fill; move the ball close to the screen in DEMO1, and watch things slow down when one of those big texture maps comes around. Of course, that’s partly because the code is all in C; some well-chosen optimizations would work wonders. In the next chapter we’ll discuss texture mapping further, crank up the speed of our texture mapper, and attend to some rough spots that remain in the DDA texture mapping implementation, most notably in the area of exactly which texture pixels map to which destination pixels as a polygon rotates.

    @@ -374,10 +367,6 @@ void ScanOutLine(EdgeScan * LeftEdge, EdgeScan * RightEdge)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/57-01.html b/57-01.html index ea0b04c..e8861fa 100644 --- a/57-01.html +++ b/57-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - @@ -37,16 +30,16 @@


    -

    Chapter 57
    +

    Chapter 57
    10,000 Freshly Sheared Sheep on the Screen

    -

    The Critical Role of Experience in Implementing Fast, Smooth Texture Mapping

    +

    The Critical Role of Experience in Implementing Fast, Smooth Texture Mapping

    I recently spent an hour or so learning how to shear a sheep. Among other things, I learned—in great detail—about the importance of selecting the proper comb for your shears, heard about the man who holds the world’s record for sheep sheared in a day (more than 600, if memory serves), and discovered, Lord help me, the many and varied ways in which the New Zealand Sheep Shearing Board improves the approved sheep-shearing method every year. The fellow giving the presentation did his best, but let’s face it, sheep just aren’t very interesting. If you have children, you’ll know why I was there; if you don’t, there’s no use explaining.

    The chap doing the shearing did say one thing that stuck with me, although it may not sound particularly profound. (Actually, it sounds pretty silly, but bear with me.) He said, “You don’t get really good at sheep shearing for 10 years, or 10,000 sheep.” I’ll buy that. In fact, to extend that morsel of wisdom to the greater, non-ovine-centric universe, it actually takes a good chunk of experience before you get good at anything worthwhile—especially graphics, for a couple of reasons. First, performance matters a lot in graphics, and performance programming is largely a matter of experience. You can’t speed up PC graphics simply by looking in a book for a better algorithm; you have to understand the code C compilers generate, assembly language optimization, VGA hardware, and the performance implications of various graphics-programming approaches and algorithms. Second, computer graphics is a matter of illusion, of convincing the eye to see what you want it to see, and that’s very much a black art based on experience.

    -

    Visual Quality: A Black Hole ... Er, Art

    +

    Visual Quality: A Black Hole ... Er, Art

    Pleasing the eye with realtime computer animation is something less than a science, at least at the PC level, where there’s a limited color palette and no time for antialiasing; in fact, sometimes it can be more than a little frustrating. As you may recall, in the previous chapter I implemented texture mapping in X-Sharp. There was plenty of experience involved there, some of which I didn’t mention. My first implementation was disappointing; the texture maps shimmied and sheared badly, like a loosely affiliated flock of pixels, each marching to its own drummer. Then, I added a control key to speed up the rotation; what a difference! The aliasing problems were still there, but with the faster rotation, the pixels moved too quickly for the eye to pick up on the aliasing; the rotating texture maps, and the rotating ball as a whole, crossed the threshold into being accepted by the eye as a viewed object, rather than simply a collection of pixels.

    @@ -60,7 +53,7 @@ -

    Fixed-Point Arithmetic, Redux

    +

    Fixed-Point Arithmetic, Redux

    In the previous chapter I added texture mapping to X-Sharp, but lacked space to explain some of its finer points. I’ll pick up the thread now and cover some of those points here, and discuss the visual and performance enhancements that previous chapter’s code needed—and which are now present in the version of X-Sharp in this chapter’s subdirectory on the CD-ROM.

    @@ -93,10 +86,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/57-02.html b/57-02.html index ef456e4..11a8d7c 100644 --- a/57-02.html +++ b/57-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - @@ -37,19 +30,18 @@


    -

    Texture Mapping: Orientation Independence

    +

    Texture Mapping: Orientation Independence

    The double-DDA texture-mapping code presented in the previous chapter worked adequately, but there were two things about it that left me less than satisfied. One flaw was performance; I’ll address that shortly. The other flaw was the way textures shifted noticeably as the orientations of the polygons onto which they were mapped changed.

    The previous chapter’s code followed the standard polygon inside/outside rule for determining which pixels in the source texture map were to be mapped: Pixels that mapped exactly to the left and top destination edges were considered to be inside, and pixels that mapped exactly to the right and bottom destination edges were considered to be outside. That’s fine for filling polygons, but when copying texture maps, it causes different edges of the texture map to be omitted, depending on the destination orientation, because different edges of the texture map correspond to the right and bottom destination edges, depending on the current rotation. Also, the previous chapter’s code truncated to get integer source coordinates. This, together with the orientation problem, meant that when a texture turned upside down, it slowed one new row and one new column of pixels from the next row and column of the texture map. This asymmetry was quite visible, and not at all the desired effect.

    -


    - Figure 57.1
      Gaps caused by mixing fixed-point and all-integer math.

    +


    + Figure 57.1
      Gaps caused by mixing fixed-point and all-integer math.

    Listing 57.1 is one solution to these problems. This code, which replaces the equivalently named function presented in the previous chapter (and, of course, is present in the X-Sharp archive in this chapter’s subdirectory of the listings disk), makes no attempt to follow the standard polygon inside/outside rules when mapping the source. Instead, it advances a half-step into the texture map before drawing the first pixel, so pixels along all edges are half included. Rounding rather than truncation to texture-map coordinates is also performed. The result is that the texture map stays pretty much centered within the destination polygon as the destination rotates, with a much-reduced level of orientation-dependent asymmetry.

    -

    LISTING 57.1 L57-1.C

    +

    LISTING 57.1 L57-1.C

     /* Texture-map-draw the scan line between two edges. Uses approach of
        pre-stepping 1/2 pixel into the source image and rounding to the nearest
    @@ -121,13 +113,13 @@ void ScanOutLine(EdgeScan * LeftEdge, EdgeScan * RightEdge)
           SourceY += SourceStepY;
        }
     }
    -
    + -

    Mapping Textures across Multiple Polygons

    +

    Mapping Textures across Multiple Polygons

    One of the truly nifty things about double-DDA texture mapping is that it is not limited to mapping a texture onto a single polygon. A single texture can be mapped across any number of adjacent polygons simply by having polygons that share vertices in 3-space also share vertices in the texture map. In fact, the demonstration program DEMO1 in the X-Sharp archive maps a single texture across two polygons; this is the blue-on-green pattern that stretches across two panels of the spinning ball. This capability makes it easy to produce polygon-based objects with complex surfaces (such as banding and insignia on spaceships, or even human figures). Just map the desired texture onto the underlying polygonal framework of an object, and let double-DDA texture mapping do the rest.

    -

    Fast Texture Mapping

    +

    Fast Texture Mapping

    Of course, there’s a problem with mapping a texture across many polygons: Texture mapping is slow. If you run DEMO1 and move the ball up close to the screen, you’ll see that the ball slows considerably whenever a texture swings around into view. To some extent that can’t be helped, because each pixel of a texture-mapped polygon has to be calculated and drawn independently. Nonetheless, we can certainly improve the performance of texture mapping a good deal over what I presented in the previous chapter.

    @@ -152,10 +144,6 @@ void ScanOutLine(EdgeScan * LeftEdge, EdgeScan * RightEdge)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/57-03.html b/57-03.html index cf977ec..342c95b 100644 --- a/57-03.html +++ b/57-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - @@ -37,7 +30,7 @@


    -

    LISTING 57.2 L57-2.ASM

    +

    LISTING 57.2 L57-2.ASM

     ; Draws all pixels in the specified scan line, with the pixel colors
     ; taken from the specified texture map.  Uses approach of pre-stepping
    @@ -338,7 +331,7 @@ ScanDone:
     -ScanOutLine    endp
             end
     
    -
    +


    @@ -357,10 +350,6 @@ ScanDone:
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/57-04.html b/57-04.html index 076fb48..7a37ef8 100644 --- a/57-04.html +++ b/57-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 10,000 Freshly Sheared Sheep on the Screen - - @@ -64,10 +57,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/58-01.html b/58-01.html index fb4f9dd..93687a4 100644 --- a/58-01.html +++ b/58-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - @@ -37,10 +30,10 @@


    -

    Chapter 58
    +

    Chapter 58
    Heinlein’s Crystal Ball, Spock’s Brain, and the 9-Cycle Dare

    -

    Using the Whole-Brain Approach to Accelerate Texture Mapping

    +

    Using the Whole-Brain Approach to Accelerate Texture Mapping

    I’ve had the pleasure recently of rereading several of the works of Robert A. Heinlein, and I’m as impressed as I was as a teenager—but in a different way. The first time around, I was wowed by the sheer romance of technology married to powerful stories; this time, I’m struck most of all by The Master’s remarkable prescience. “Blowups Happen” is about the risks of nuclear power, and their effects on human psychology—written before a chain reaction had ever happened on this planet. “Solution Unsatisfactory” is about the unsolvable dilemma—ultimate offense, no defense—posed by atomic weapons; this in 1941. And in Between Planets (1951), consider this minor bit of action:

    @@ -58,7 +51,7 @@

    As Exhibit #1, I present my experience with speeding up the texture mapper in X-Sharp.

    -

    Texture Mapping Redux

    +

    Texture Mapping Redux

    We’ve spent the previous several chapters exploring the X Sharp graphics library, something I built over time as a serious exercise in 3-D graphics. When X-Sharp reached the point at which we left it at the end of the previous chapter, I was rather pleased with it—with one exception.

    @@ -66,15 +59,14 @@

    It was the “Hmph” that really got to me.

    -

    Left-Brain Optimization

    +

    Left-Brain Optimization

    That was the first shot of juice for my optimizer (or at least blow to my ego, which can be just as productive). John went on to say he had gotten texture mapping down to 9 cycles per pixel and one jump per scanline on a 486 (all cycle times will be for the 486 unless otherwise noted); given that my code took, on average, about 44 cycles and 2 taken jumps (plus 1 not taken) per pixel, I had a long way to go.

    The inner loop of my original texture-mapping code is shown in Listing 58.1. All this code does is draw a single texture-mapped scanline, as shown in Figure 58.1; an outer loop runs through all the scanlines in whatever polygon is being drawn. I immediately saw that I could eliminate nearly 10 percent of the cycles by unrolling the loop; obviously, John had done that, else there’s no way he could branch only once per scanline. (By the way, branching only once per scanline via a fully unrolled loop is not generally recommended. A branch every few pixels costs relatively little, and the cache effects of fully unrolled code are not good.) I quickly came up with several other ways to speed up the code, but soon realized that all the clever coding in the world wasn’t going to get me within 100 percent of John’s performance so long as I had to cycle from one plane to the next for every pixel.

    -


    - Figure 58.1
      Texture mapping a single horizontal scanline.

    +


    + Figure 58.1
      Texture mapping a single horizontal scanline.


    @@ -93,10 +85,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/58-02.html b/58-02.html index bddee49..89e655d 100644 --- a/58-02.html +++ b/58-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - @@ -37,7 +30,7 @@


    -

    LISTING 58.1 L58-1.ASM

    +

    LISTING 58.1 L58-1.ASM

     ; Inner loop to draw a single texture-mapped horizontal scanline in
     ; Mode X, the VGA’s page-flipped 256-color mode. Because adjacent
    @@ -86,13 +79,12 @@ NoExtraYAdvance:
             dec     si
             jnz     TexScanLoop
     
    -
    +

    Figure 58.2 shows why this cycling is necessary. In Mode X, the page-flipped 256-color mode of the VGA, each successive pixel across a scanline is stored in a different hardware plane, and an OUT to the VGA’s hardware is needed to select the plane being drawn to. (See Chapters 47, 48, and 49 for details.) An OUT instruction by itself takes 16 cycles (and in the neighborhood of 30 cycles in virtual-86 or non-privileged protected mode), and an ROL takes 2 more, for a total of 18 cycles, double John’s 9 cycles, just to handle plane management. Clearly, getting plane control out of the inner loop was absolutely necessary.

    -


    - Figure 58.2
      Display memory organization in Mode X.

    +


    + Figure 58.2
      Display memory organization in Mode X.

    I must confess, with some embarrassment, that at this point I threw myself into designing a solution that involved executing the texture mapping code up to four times per scanline, once for the pixels in each plane. It’s hard to overstate the complexity of this approach, which involves quadrupling the normal pixel-to-pixel increments, adjusting the start value for each of the passes, and dealing with some nasty boundary cases. Make no mistake, the code was perfectly doable, and would in fact have gotten plane control out of the inner loop, but would have been very difficult to get exactly right, and would have suffered from substantial overhead.

    @@ -102,11 +94,11 @@ NoExtraYAdvance:

    Why indeed?

    -

    A 90-Degree Shift in Perspective

    +

    A 90-Degree Shift in Perspective

    As I said earlier, how you look at an optimization problem defines how you’ll be able to solve it. In order to boost performance, sometimes it’s necessary to look at things from a different angle—and for texture mapping this was literally as well as figuratively true. Chris suggested nothing more nor less than scanning out polygons at a 90-degree angle to normal, starting, say, at the left edge of the polygon, and texture-mapping vertically along each column of pixels, as shown in Figure 58.3. That way, all the pixels in each texture-mapped column would be in the same plane, and I would need to change planes only between columns—outside the inner loop. A trivial change, not fundamental in any sense—and yet just that one change, plus unrolling the loop, reduced the inner loop to the 22-cycles-per-pixel version shown in Listing 58.2. That’s exactly twice as fast as Listing 58.1—and given how incredibly slow most VGAs are at completing OUTs, the real-world speedup should be considerably greater still. (The fastest byte OUT I’ve ever measured for a VGA is 29 cycles, the slowest more than 60 cycles; in the latter case, Listing 58.2 would be on the order of four times faster than Listing 58.1.)

    -

    LISTING 58.2 L58-2.ASM

    +

    LISTING 58.2 L58-2.ASM

     ; Inner loop to draw a single texture-mapped vertical column, rather
     ; than a horizontal scanline. This allows all pixels handled
    @@ -150,7 +142,7 @@ NoExtraYAdvance:
     
     ENDM
     
    -
    +


    @@ -169,10 +161,6 @@ ENDM
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/58-03.html b/58-03.html index ef96217..69405f5 100644 --- a/58-03.html +++ b/58-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - @@ -47,57 +40,53 @@ -


    - Figure 58.3
      Texture mapping a single vertical column.

    +


    + Figure 58.3
      Texture mapping a single vertical column.

    There are a few complications with Chris’s approach, not least that X-Sharp’s polygon-filling convention (top and left edges included, bottom and right edges excluded) is hard to reproduce for column-oriented texture mapping. I solved this in X-Sharp version 22 by tweaking the edge-scanning code to allow column-oriented texture mapping to match the current convention. (You’ll find X-Sharp 22 on the listings diskette in the directory for this chapter.)

    Chris also illustrated another important principle of optimization: A second pair of eyes is invaluable. Even the best of us have blind spots and get caught up in particular implementations; if you bounce your ideas off someone, you may well find them coming back with an unexpected—and welcome—spin.

    -

    That’s Nice—But it Sure as Heck Ain’t 9 Cycles

    +

    That’s Nice—But it Sure as Heck Ain’t 9 Cycles

    Excellent as Chris’s suggestion was, I still had work to do: Listing 58.2 is still more than twice as slow as John Miles’s code. Traditionally, I start the optimization process with algorithmic optimization, then try to tie the algorithm and the hardware together for maximum efficiency, and finish up with instruction-by-instruction, take-no-prisoners optimization. We’ve already done the first two steps, so it’s time to get down to the bare metal.

    Listing 58.2 contains three functional parts: Drawing the pixel, advancing the destination pointer, and advancing the source texture pointer. Each of the three parts is amenable to further acceleration.

    -

    Drawing the pixel is difficult to speed up, given that it consists of only two instructions—difficult, but not impossible. True, the instructions themselves are indeed irreducible, but if we can get rid of the ES: prefix (and, as we shall see, we can), we can rearrange the code to make it run faster on the Pentium. Without a prefix, the instructions execute as follows on the Pentium:

    +

    Drawing the pixel is difficult to speed up, given that it consists of only two instructions—difficult, but not impossible. True, the instructions themselves are indeed irreducible, but if we can get rid of the ES: prefix (and, as we shall see, we can), we can rearrange the code to make it run faster on the Pentium. Without a prefix, the instructions execute as follows on the Pentium:

     
     MOV  AH,[BX]    ;cycle 1 U-pipe
                     ;cycle 1 V-pipe idle; reg contention
     MOV  [DI],AH    ;cycle 2 U-pipe
     
    -
    +

    The second MOV, being dependent on the value loaded into AH by the first MOV, can’t execute until the first MOV is finished, so the Pentium’s second pipe, the V-pipe, lies idle for a cycle. We can reclaim that cycle simply by shuffling another instruction between the two MOVs.

    -

    Advancing the destination pointer is easy to speed up: Just build the offset from one scanline to the next into each pixel-drawing instruction as a constant, as in

    +

    Advancing the destination pointer is easy to speed up: Just build the offset from one scanline to the next into each pixel-drawing instruction as a constant, as in

     
     MOV [EDI+SCANOFFSET],AH
     
    -
    +

    and advance EDI only once per unrolled loop iteration.

    Advancing the source texture pointer is more complex, but correspondingly more rewarding. Listing 58.2 uses a variant form of 32-bit fixed-point arithmetic to advance the source pointer, with the source texture coordinates and increments stored in 16.16 (16 bits of integer, 16 bits of fraction) format. The source coordinates are stored in a slightly unusual format, whereby the fractional X and Y coordinates are stored and advanced separately, but a single integer value, the source pointer, is used to reflect both the X and Y coordinates. In Listing 58.2, the integer and fractional parts are added into the current coordinates with four separate 16-bit operations, and carries from fractional to integer parts are detected via conditional jumps, as shown in Figure 58.4. There’s quite a lot we can do to improve this.

    -


    - Figure 58.4
      Original method for advancing the source texture pointer.

    +


    + Figure 58.4
      Original method for advancing the source texture pointer.

    First, we can sum the X and Y integer advance amounts outside the loop, then add them both to the source pointer with a single instruction. Second, we can recognize that X advances exactly one extra byte when its fractional part carries, and use ADC to account for X carries, as shown in Figure 58.5. That single ADC can add in not only any X carry, but both the X and Y integer advance amounts as well, thereby eliminating a good chunk of the source-advance code in Listing 58.2. Furthermore, we should somehow be able to use 32-bit registers and instructions to help with the 32-bit fixed-point arithmetic; true, the size override prefix (because we’re in a 16-bit segment) will cost a cycle per 32-bit instruction, but that’s better than the 3 cycles it takes to do 32-bit arithmetic with 16-bit instructions. It isn’t obvious, but there’s a nifty trick we can use here, again courtesy of Chris Hecker (who, as you can tell, has done a fair amount of thinking about the complexities of texture mapping).

    We can store the current fractional parts of both the X and Y source coordinates in a single 32-bit register, EDX, as shown in Figure 58.6. It’s important to note that the Y fraction is actually only 15 bits, with bit 15 of EDX always kept at zero; this allows bit 15 to store the carry status from each Y advance. We can similarly store the fractional X and Y advance amounts in ECX, and can store the sum of the integer parts of the X and Y advance amounts in BP. With this arrangement, the single instruction ADD EDX,ECX advances the fractional parts of both X and Y, and the following instruction ADC SI,BP finishes advancing the source pointer in X. That’s a mere 3 cycles, and all that remains is to finish advancing the source pointer in Y.

    -


    - Figure 58.5
      Efficient method for advancing source texture pointer.

    +


    + Figure 58.5
      Efficient method for advancing source texture pointer.

    -


    - Figure 58.6
      Storing both X and Y fractional coordinates in one register.

    +


    + Figure 58.6
      Storing both X and Y fractional coordinates in one register.

    Actually, we also advanced the source pointer by the Y integer amount back when we added BP to SI; all that’s left is to detect whether our addition to the Y fractional current coordinate produced a carry. That’s easily done by testing bit 15 of EDX; if it’s zero, there was no carry and we’re done; otherwise, Y carried, so we have to reset bit 15 and advance the source pointer by one scanline. The resulting program flow is shown in Figure 58.7. Note that unlike the X fractional addition, we can’t get away with just adding in the carry from the Y fractional addition, because when the Y fraction carries, it indicates a move not from one pixel to the next on a scanline (a single byte), but rather from one scanline to the next (a full scanline width).

    @@ -118,10 +107,6 @@ MOV [EDI+SCANOFFSET],AH
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/58-04.html b/58-04.html index 772e2e8..bca3007 100644 --- a/58-04.html +++ b/58-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - @@ -41,7 +34,7 @@

    Or can we?

    -

    LISTING 58.3 L58-3.ASM

    +

    LISTING 58.3 L58-3.ASM

     ; Inner loop to draw a single texture-mapped vertical column,
     ; rather than a horizontal scanline. Maxed-out 16-bit version.
    @@ -51,11 +44,10 @@
     ;       ECX = fractional Y advance in lower 15 bits of CX,
     ;             fractional X advance in high word of ECX, bit
     ;             15 set to 0
    -
    + -


    - Figure 58.7
      Final method for advancing source texture pointer.

    +


    + Figure 58.7
      Final method for advancing source texture pointer.

     ;       EDX = fractional source texture Y coordinate in lower
     ;             15 bits of CX, fractional source texture X coord
    @@ -85,9 +77,9 @@ SCANOFFSET=0
     SCANOFFSET = SCANOFFSET + SCANWIDTH
     
          ENDM
    -
    + -

    Don’t Stop Thinking about Those Cycles

    +

    Don’t Stop Thinking about Those Cycles

    Remember what I said at the outset, that knowing something has been done makes it much easier to do? A corollary is that pushing past that point, once attained, is very difficult. It’s only natural to want to relax in the satisfaction of a job well done; then, too, the very nature of the work changes. Getting from 44 cycles down to John’s 9 cycles was a huge leap, but we knew it could be done—therefore the nature of the problem was to figure out how it was done; in cases like this, if we’re sharp enough (and of course we are!), we’re guaranteed eventual gratification. Now that we’ve reached John’s level of performance, the problem becomes whether the code can be made faster yet, and that’s a different kettle of fish altogether, for it may well be that after thinking about it for a while, we’ll conclude that it can’t. Not only will we have wasted time, but we’ll also never be sure we were right; we’ll know only that we couldn’t find a solution. That way lies madness.

    @@ -110,10 +102,6 @@ SCANOFFSET = SCANOFFSET + SCANWIDTH
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/58-05.html b/58-05.html index 4352cb9..0819e83 100644 --- a/58-05.html +++ b/58-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Heinlein's Crystal Ball, Spock's Brain, and the 9-Cycle Dare - - @@ -37,7 +30,7 @@


    -

    LISTING 58.4 L58-4.ASM

    +

    LISTING 58.4 L58-4.ASM

     ; Inner loop to draw a single texture-mapped vertical column,
     ; rather than a horizontal scanline. Maxed-out 32-bit version.
    @@ -78,13 +71,13 @@ SCANOFFSET = SCANOFFSET + SCANWIDTH
     
          ENDM
     
    -
    +

    And there you have it: A five to 10-times speedup of a decent assembly language texture mapper. All it took was some help from my friends, a good, stiff jolt of right-brain thinking, and some solid left-brain polishing—plus the knowledge that such a speedup was possible. Treat every optimization task as if John Miles has just written to inform you that he’s made it faster than your wildest dreams, and you’ll be amazed at what you can do!

    -

    Texture Mapping Notes

    +

    Texture Mapping Notes

    -

    Listing 58.3 contains no 486 pipeline stalls; it has Pentium stalls, but not much can be done for them because of the size prefix on ADD EDX,ECX, which takes 1 cycle to go through the U-pipe, and shuts down the V-pipe for that cycle. Listing 58.4, on the other hand, has been rearranged to eliminate all Pentium stalls save one. When the Y coordinate fractional part carries and ESI advances, the code executes as follows:

    +

    Listing 58.3 contains no 486 pipeline stalls; it has Pentium stalls, but not much can be done for them because of the size prefix on ADD EDX,ECX, which takes 1 cycle to go through the U-pipe, and shuts down the V-pipe for that cycle. Listing 58.4, on the other hand, has been rearranged to eliminate all Pentium stalls save one. When the Y coordinate fractional part carries and ESI advances, the code executes as follows:

     
     ADD ESI,ECX     ;cycle 1 U-pipe
    @@ -93,7 +86,7 @@ AND DH,NOT 80H  ;cycle 1 V-pipe
     MOV BL,[ESI]    ;cycle 3 U-pipe
     ADD EDX,EBP     ;cycle 3 V-pipe
     
    -
    +

    However, I don’t see any way to eliminate this last AGI, which happens about half the time; even with it, the Pentium execution time for Listing 58.4 is 5.5 cycles. That’s 61 nanoseconds—a highly respectable 16 million texture-mapped pixels per second—on a 90 MHz Pentium.

    @@ -120,10 +113,6 @@ ADD EDX,EBP ;cycle 3 V-pipe
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/59-01.html b/59-01.html index 1a63580..2799845 100644 --- a/59-01.html +++ b/59-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - @@ -37,10 +30,10 @@


    -

    Chapter 59
    +

    Chapter 59
    The Idea of BSP Trees

    -

    What BSP Trees Are and How to Walk Them

    +

    What BSP Trees Are and How to Walk Them

    The answer is: Wendy Tucker.

    @@ -68,13 +61,13 @@

    Before we begin, I’d like to thank John Carmack, the technical wizard behind DOOM, for generously sharing his knowledge of BSP trees with me.

    -

    BSP Trees

    +

    BSP Trees

    A BSP tree is, at heart, nothing more than a tree that subdivides space in order to isolate features of interest. Each node of a BSP tree splits an area or a volume (in 2-D or 3-D, respectively) into two parts along a line or a plane; thus the name “Binary Space Partitioning.” The subdivision is hierarchical; the root node splits the world into two subspaces, then each of the root’s two children splits one of those two subspaces into two more parts. This continues with each subspace being further subdivided, until each component of interest (each line segment or polygon, for example) has been assigned its own unique subspace. This is, admittedly, a pretty abstract description, but the workings of BSP trees will become clearer shortly; it may help to glance ahead to this chapter’s figures.

    Building a tree that subdivides space doesn’t sound particularly profound, but there’s a lot that can be done with such a structure. BSP trees can be used to represent shapes, and operating on those shapes is a simple matter of combining trees as needed; this makes BSP trees a powerful way to implement Constructive Solid Geometry (CSG). BSP trees can also be used for hit testing, line-of-sight determination, and collision detection.

    -

    Visibility Determination

    +

    Visibility Determination

    For the time being, I’m going to discuss only one of the many uses of BSP trees: The ability of a BSP tree to allow you to traverse a set of line segments or polygons in back-to-front or front-to-back order as seen from any arbitrary viewpoint. This sort of traversal can be very helpful in determining which parts of each line segment or polygon are visible and which are occluded from the current viewpoint in a 3-D scene. Thus, a BSP tree makes possible an efficient implementation of the painter’s algorithm, whereby polygons are drawn in back-to-front order, with closer polygons overwriting more distant ones that overlap, as shown in Figure 59.1. (The line segments in Figure 1(a) and in other figures in this chapter, represent vertical walls, viewed from directly above.) Alternatively, visibility determination can be performed by front-to-back traversal working in conjunction with some method for remembering which pixels have already been drawn. The latter approach is more complex, but has the potential benefit of allowing you to early-out from traversal of the scene database when all the pixels on the screen have been drawn.

    @@ -95,10 +88,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/59-02.html b/59-02.html index 5ba9197..feb89e9 100644 --- a/59-02.html +++ b/59-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - @@ -41,11 +34,10 @@

    It’s hard to get cheaper sorting than linear time, and BSP-based rendering stacks up well against alternatives such as z-buffering, octrees, z-scan sorting, and polygon sorting. Better yet, a scene database represented as a BSP tree can be clipped to the view pyramid very efficiently; huge chunks of a BSP tree can be lopped off when clipping to the view pyramid, because if the entire area or volume of a node lies entirely outside the view volume, then all nodes and leaves that are children of that node must likewise be outside the view volume, for reasons that will become clear as we delve into the workings of BSP trees.

    -


    - Figure 59.1
      The painter’s algorithm.

    +


    + Figure 59.1
      The painter’s algorithm.

    -

    Limitations of BSP Trees

    +

    Limitations of BSP Trees

    Powerful as they are, BSP trees aren’t perfect. By far the greatest limitation of BSP trees is that they’re time-consuming to build, enough so that, for all practical purposes, BSP trees must be precalculated, and cannot be built dynamically at runtime. In fact, a BSP-tree compiler that attempts to perform some optimization (limiting the number of surfaces that need to be split, for example) can easily take minutes or even hours to process large world databases.

    @@ -57,7 +49,7 @@

    The only other drawbacks of BSP trees that I know of are the memory required to store the tree, which amounts to a few pointers per node, and the relative complexity of debugging BSP-tree compilation and usage; debugging a large data set being processed by recursive code (which BSP code tends to be) can be quite a challenge. Tools like the BSP compiler I’ll present in the next chapter, which visually depicts the process of spatial subdivision as a BSP tree is constructed, help a great deal with BSP debugging.

    -

    Building a BSP Tree

    +

    Building a BSP Tree

    Now that we know a good bit about what a BSP tree is, how it helps in visible surface determination, and what its strengths and weaknesses are, let’s take a look at how a BSP tree actually works to provide front-to-back or back-to-front ordering. This chapter’s discussion will be at a conceptual level, with plenty of figures; in the next chapter we’ll get into mechanisms and implementation details.

    @@ -65,9 +57,8 @@

    First, let’s construct a simple BSP tree. Figure 59.2 shows a set of four lines that will constitute our sample world. I’ll refer to these as walls, because that’s one easily-visualized context in which a 2-D BSP tree would be useful in a game. Think of Figure 59.2 as depicting vertical walls viewed from directly above, so they’re lines for the purpose of the BSP tree. Note that each wall has a front side, denoted by a normal (perpendicular) vector, and a back side. To make a BSP tree for this sample set, we need to split the world in two, then each part into two again, and so on, until each wall resides in its own unique subspace. An obvious question, then, is how should we carve up the world of Figure 59.2?

    -


    - Figure 59.2
      A sample set of walls, viewed from above.

    +


    + Figure 59.2
      A sample set of walls, viewed from above.

    There are infinitely valid ways to carve up Figure 59.2, but the simplest is just to carve along the lines of the walls themselves, with each node containing one wall. This is not necessarily optimal, in the sense of producing the smallest tree, but it has the virtue of generating the splitting lines without expensive analysis. It also saves on data storage, because the data for the walls can do double duty in describing the splitting lines as well. (Putting one wall on each splitting line doesn’t actually create a unique subspace for each wall, but it does create a unique subspace boundary for each wall; as we’ll see, that spatial organization provides for the same unambiguous visibility ordering as a unique subspace would.)

    @@ -88,10 +79,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/59-03.html b/59-03.html index fecec62..9cf7064 100644 --- a/59-03.html +++ b/59-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - @@ -39,25 +32,22 @@

    Creating a BSP tree is a recursive process, so we’ll perform the first split and go from there. Figure 59.3 shows the world carved along the line of wall C into two parts: walls that are in front of wall C, and walls that are behind. (Any of the walls would have been an equally valid choice for the initial split; we’ll return to the issue of choosing splitting walls in the next chapter.) This splitting into front and back is the essential dualism of BSP trees.

    -


    - Figure 59.3
      Initial split along the line of wall C.

    +


    + Figure 59.3
      Initial split along the line of wall C.

    Next, in Figure 59.4, the front subspace of wall C is split by wall D. This is the only wall in that subspace, so we’re done with wall C’s front subspace.

    Figure 59.5 shows the back subspace of wall C being split by wall B. There’s a difference here, though: Wall A straddles the splitting line generated from wall B. Does wall A belong in the front or back subspace of wall B?

    -


    - Figure 59.4
      Split of wall C’s front subspace along the line of wall D.

    +


    + Figure 59.4
      Split of wall C’s front subspace along the line of wall D.

    -


    - Figure 59.5
      Split of wall C’s back subspace along the line of wall B.

    +


    + Figure 59.5
      Split of wall C’s back subspace along the line of wall B.

    Both, actually. Wall A gets split into two pieces, which I’ll call wall A and wall E; each piece is assigned to the appropriate subspace and treated as a separate wall. As shown in Figure 59.6, each of the split pieces then has a subspace to itself, and each becomes a leaf of the tree. The BSP tree is now complete.

    -

    Visibility Ordering

    +

    Visibility Ordering

    Now that we’ve successfully built a BSP tree, you might justifiably be a little puzzled as to how any of this helps with visibility ordering. The answer is that each BSP node can definitively determine which of its child trees is nearer and which is farther from any and all viewpoints; applied throughout the tree, this principle makes it possible to establish visibility ordering for all the line segments or planes in a BSP tree, no matter what the viewing angle.

    @@ -65,25 +55,22 @@

    Of course, we need more ordering information than wall C alone can give us, but we get that by traversing the tree recursively, making the same far-near decision at each node. Figure 59.8 shows the painter’s algorithm (back-to-front) traversal order of the tree for the viewpoint of Figure 59.7. At each node, we decide whether we’re seeing the front or back side of that node’s wall, then visit whichever of the wall’s children is on the far side from the viewpoint, draw the wall, and then visit the node’s nearer child, in that order. Visiting a child is recursive, involving the same far-near visiting order.

    -


    - Figure 59.6
      The final BSP tree.

    +


    + Figure 59.6
      The final BSP tree.

    -


    - Figure 59.7
      Viewing the BSP tree from an arbitrary angle.

    +


    + Figure 59.7
      Viewing the BSP tree from an arbitrary angle.

    The key is that each BSP splitting line separates all the walls in the current subspace into two groups relative to the viewpoint, and every single member of the farther group is guaranteed not to occlude every single member of the nearer. By applying this ordering recursively, the BSP tree can be traversed to provide back-to-front or front-to-back ordering, with each node being visited only once.

    -


    - Figure 59.8
      Back-to-front traversal of the BSP tree as viewed in Figure 59.7.

    +


    + Figure 59.8
      Back-to-front traversal of the BSP tree as viewed in Figure 59.7.

    The type of tree walk used to produce front-to-back or back-to-front BSP traversal is known as an inorder walk. More on this very shortly; you’re also likely to find a discussion of inorder walking in any good data structures book. The only special aspect of BSP walks is that a decision has to be made at each node about which way the node’s wall is facing relative to the viewpoint, so we know which child tree is nearer and which is farther.

    Listing 59.1 shows a function that draws a BSP tree back-to-front. The decision whether a node’s wall is facing forward, made by WallFacingForward() in Listing 59.1, can, in general, be made by generating a normal to the node’s wall in screenspace (perspective-corrected space as seen from the viewpoint) and checking whether the z component of the normal is positive or negative, or by checking the sign of the dot product of a viewspace (non-perspective corrected space as seen from the viewpoint) normal and a ray from the viewpoint to the wall. In 2-D, the decision can be made by enforcing the convention that when a wall is viewed from the front, the start vertex is leftmost; then a simple screenspace comparison of the x coordinates of the left and right vertices indicates which way the wall is facing.

    -

    Listing 59.1 L59_1.C

    +

    Listing 59.1 L59_1.C

     void WalkBSPTree(NODE *pNode)
     {
    @@ -105,7 +92,7 @@ void WalkBSPTree(NODE *pNode)
           }
        }
     }
    -
    + @@ -132,10 +119,6 @@ void WalkBSPTree(NODE *pNode)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/59-04.html b/59-04.html index 81fe215..c9c5d72 100644 --- a/59-04.html +++ b/59-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - @@ -37,7 +30,7 @@


    -

    Inorder Walks of BSP Trees

    +

    Inorder Walks of BSP Trees

    It was implementing BSP trees that got me to thinking about inorder tree traversal. In inorder traversal, the left subtree of each node gets visited first, then the node, and then the right subtree. You apply this sequence recursively to each node and its children until the entire tree has been visited, as shown in Figure 59.9. Walking a BSP tree is basically an inorder tree walk; the only difference is that with a BSP tree a decision is made before each descent as to which subtree to visit first, rather than simply visiting whatever’s pointed to by the left-subtree pointer. Conceptually, however, an inorder walk is what’s used to traverse a BSP tree; from now on I’ll discuss normal inorder walking, with the understanding that the same principles apply to BSP trees.

    @@ -45,11 +38,10 @@

    First, I ask for an implementation of a function WalkTree() that visits each node in a passed-in tree in inorder sequence. Each candidate unhesitatingly writes something like the perfectly good code in Listings 59.2 and 59.3 shown next.

    -


    - Figure 59.9
      An inorder walk of a BSP tree.

    +


    + Figure 59.9
      An inorder walk of a BSP tree.

    -

    Listing 59.2 L59_2.C

    +

    Listing 59.2 L59_2.C

     // Function to inorder walk a tree, using code recursion.
     // Tested with 32-bit Visual C++ 1.10.
    @@ -75,9 +67,9 @@ void WalkTree(NODE *pNode)
           }
        }
     }
    -
    + -

    Listing 59.3 L59_3.H

    +

    Listing 59.3 L59_3.H

     // Header file TREE.H for tree-walking code.
     typedef struct _NODE {
    @@ -85,7 +77,7 @@ struct _NODE *pLeftChild;
     struct _NODE *pRightChild;
     } NODE;
     
    -
    +

    Then I ask if they have any idea how to make the code faster; some don’t, but most point out that function calls are pretty expensive. Either way, I then ask them to rewrite the function without code recursion.

    @@ -95,11 +87,11 @@ struct _NODE *pRightChild;

    And yet, a data-recursive inorder walk implementation has exactly the same flowchart and exactly the same functionality as the code-recursive version they’ve already written. They already have a fully functional model to follow, with all the problems solved, but they can’t make the connection between that model and the code they’re trying to implement. Why is this?

    -

    Know It Cold

    +

    Know It Cold

    The problem is that these people don’t understand inorder walking through and through. They understand the concepts of visiting left and right subtrees, and they have a general picture of how traversal moves about the tree, but they do not understand exactly what the code-recursive version does. If they really comprehended everything that happens in each iteration of WalkTree()—how each call saves the state, and what that implies for the order in which operations are performed—they would simply and without fuss implement code like that in Listing 59.4, working with the code-recursive version as a model.

    -

    Listing 59.4 L59_4.C

    +

    Listing 59.4 L59_4.C

     // Function to inorder walk a tree, using data recursion.
     // No stack overflow testing is performed.
    @@ -171,7 +163,7 @@ void WalkTree(NODE *pNode)
        }
     }
     
    -
    +


    @@ -190,10 +182,6 @@ void WalkTree(NODE *pNode)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/59-05.html b/59-05.html index 12a34b9..fde8a9f 100644 --- a/59-05.html +++ b/59-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - @@ -55,11 +48,11 @@
    -

    Measure and Learn

    +

    Measure and Learn

    How much difference does all this fuss make, anyway? Listing 59.5 is a sample program that builds a tree, then calls WalkTree () to walk it 1,000 times, and times how long this takes. Using 32-bit Visual C++ 1.10 running on Windows NT, with default optimization selected, Listing 59.5 reports that Listing 59.4 is about 20 percent faster than Listing 59.2 on a 486/33, a reasonable return for a little code rearrangement, especially when you consider that the speedup is diluted by calling the Visit() function and by the cache miss that happens on virtually every node access. (Listing 59.5 builds a rather unique tree, one in which every node has exactly two children. Different sorts of trees can and do produce different performance results. Always know what you’re measuring!)

    -

    Listing 59.5 L59_5.C

    +

    Listing 59.5 L59_5.C

     // Sample program to exercise and time the performance of
     // implementations of WalkTree().
    @@ -128,7 +121,7 @@ void Visit(NODE *pNode)
        VisitCount++;
     }
     
    -
    +


    @@ -147,10 +140,6 @@ void Visit(NODE *pNode)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/59-06.html b/59-06.html index 40494bd..ae62af0 100644 --- a/59-06.html +++ b/59-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: The Idea of BSP Trees - - @@ -63,11 +56,11 @@

    Depths within depths indeed!

    -

    Surfing Amidst the Trees

    +

    Surfing Amidst the Trees

    In the next chapter, we’ll build a BSP-tree compiler, and after that, we’ll put together a rendering system built around the BSP trees the compiler generates. If the subject of BSP trees really grabs your fancy (as it should if you care at all about performance graphics) there is at this writing (February 1996) a World Wide Web page on BSP trees that you must investigate at http://www.qualia.com/bspfaq/. It’s set up in the familiar Internet Frequently Asked Questions (FAQ) style, and is very good stuff.

    -

    Related Reading

    +

    Related Reading

    Foley, J., A. van Dam, S. Feiner, and J. Hughes, Computer Graphics: Principles and Practice (Second Edition), Addison Wesley, 1990, pp. 555-557, 675-680.

    @@ -94,10 +87,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/60-01.html b/60-01.html index 91be286..bc2a7b9 100644 --- a/60-01.html +++ b/60-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - @@ -37,10 +30,10 @@


    -

    Chapter 60
    +

    Chapter 60
    Compiling BSP Trees

    -

    Taking BSP Trees from Concept to Reality

    +

    Taking BSP Trees from Concept to Reality

    As long-time readers of my columns know, I tend to move my family around the country quite a bit. Change doesn’t come out of the blue, so there’s some interesting history to every move, but the roots of the latest move go back even farther than usual. To wit:

    @@ -62,7 +55,7 @@

    Onward to compiling BSP trees.

    -

    Compiling BSP Trees

    +

    Compiling BSP Trees

    As you’ll recall from the previous chapter, a BSP tree is nothing more than a series of binary subdivisions that partion space into ever-smaller pieces. That’s a simple data structure, and a BSP compiler is a correspondingly simple tool. First, it groups all the surfaces (lines in 2-D, or polygons in 3-D) together into a single subspace that encompasses the entire world of the database. Then, it chooses one of the surfaces as the root node, and uses its line or plane to divide the remaining surfaces into two subspaces, splitting surfaces into two parts if they cross the line or plane of the root. Each of the two resultant subspaces is then processed in the same fashion, and so on, recursively, until the point is reached where all surfaces have been assigned to nodes, and each leaf surface subdivides a subspace that is empty except for that surface. Put another way, the root node carves space into two parts, and the root’s children carve each of those parts into two more parts, and so on, with each surface carving ever smaller subspaces, until all surfaces have been used. (Actually, there are many other lines or planes that a BSP tree can use to carve up space, but this is the approach we’ll use in the current discussion.)

    @@ -83,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/60-02.html b/60-02.html index 68155fd..9286d40 100644 --- a/60-02.html +++ b/60-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - @@ -41,7 +34,7 @@

    So there are really only two interesting operations in building a BSP tree: choosing a root node for the current subspace (a “splitter”) and assigning surfaces to one side or another of the current root node, splitting any that straddle the splitter. We’ll get to the issue of choosing splitters shortly, but first let’s look at the process of splitting and assigning. To do that, we need to understand parametric lines.

    -

    Parametric Lines

    +

    Parametric Lines

    We’re all familiar with lines described in slope-intercept form, with y as a function of x

    @@ -64,15 +57,13 @@

    What does that do for us? For one thing, it keeps clipping errors from creeping in, because clipped line segments are always based on the original line segment, not derived from clipped versions. Also, it’s potentially a more compact format, because we need to store the endpoints only for the original line segments; for clipped line segments, we can just store pairs of t values, along with a pointer to the original line segment. The biggest win, however, is that it allows us to use parametric line clipping, a very clean form of clipping, indeed.

    -


    - Figure 60.1
      A sample parametric line.

    +


    + Figure 60.1
      A sample parametric line.

    -


    - Figure 60.2
      Line segment storage in the BSP compiler.

    +


    + Figure 60.2
      Line segment storage in the BSP compiler.

    -

    Parametric Line Clipping

    +

    Parametric Line Clipping

    In order to assign a line segment to one subspace or the other of a splitter, we must somehow figure out whether the line segment straddles the splitter or falls on one side or the other. In order to determine that, we first plug the line segment and splitter into the following parametric line intersection equation

    @@ -84,13 +75,12 @@

    If the denominator is zero, we know that the lines are parallel and don’t intersect, so we don’t divide, but rather check the sign of the numerator, which tells us which side of the splitter the line segment is on. Otherwise, we do the division, and the result is the t value for the intersection point, as shown in Figure 60.3. We then simply compare the t value to the t values of the endpoints of the line segment being split. If it’s between them, that’s where we split the line segment, otherwise, we can tell which side of the splitter the line segment is on by which side of the line segment’s t range it’s on. Simple comparisons do all the work, and there’s no need to do the work of generating actual x and y values. If you look closely at Listing 60.1, the core of the BSP compiler, you’ll see that the parametric clipping code itself is exceedingly short and simple.

    -


    - Figure 60.3
      How line intersection is calculated.

    +


    + Figure 60.3
      How line intersection is calculated.

    One interesting point about Listing 60.1 is that it generates normals to splitting surfaces simply by exchanging the x and y lengths of the splitting line segment and negating the resultant y value, thereby rotating the line 90 degrees. In 3-D, it’s not that simple to come by a normal; you could calculate the normal as the cross-product of two of the polygon’s edges, or precalculate it when you build the world database.

    -

    The BSP Compiler

    +

    The BSP Compiler

    Listing 60.1 shows the core of a BSP compiler—the code that actually builds the BSP tree. (Note that Listing 60.1 is excerpted from a C++ .CPP file, but in fact what I show here is very close to straight C. It may even compile as a .C file, though I haven’t checked.) The compiler begins by setting up an empty tree, then passes that tree and the complete set of line segments from which a BSP tree is to be generated to SelectBSPTree(), which chooses a root node and calls BuildBSPTree() to add that node to the tree and generate child trees for each of the node’s two subspaces. BuildBSPTree() calls SelectBSPTree() recursively to select a root node for each of those child trees, and this continues until all lines have been assigned nodes. SelectBSP() uses parametric clipping to decide on the splitter, as described below, and BuildBSPTree() uses parametric clipping to decide which subspace of the splitter each line belongs in, and to split lines, if necessary.

    @@ -111,10 +101,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/60-03.html b/60-03.html index 1f2af27..7ad4e56 100644 --- a/60-03.html +++ b/60-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - @@ -37,7 +30,7 @@


    -

    Listing 60.1 L60_1.CPP

    +

    Listing 60.1 L60_1.CPP

     #define MAX_NUM_LINESEGS 1000
     #define MAX_INT          0x7FFFFFFF
    @@ -296,7 +289,7 @@ LINESEG * BuildBSPTree(LINESEG * plineseghead, LINESEG * prootline,
         }
         return(prootline);
     }
    -
    +


    @@ -315,10 +308,6 @@ LINESEG * BuildBSPTree(LINESEG * plineseghead, LINESEG * prootline,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/60-04.html b/60-04.html index 169e15e..2df1b8f 100644 --- a/60-04.html +++ b/60-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Compiling BSP Trees - - @@ -39,7 +32,7 @@

    Listing 60.1 isn’t very long or complex, but it’s somewhat more complicated than it could be because it’s structured to allow visual display of the ongoing compilation process. That’s because Listing 60.1 is actually just a part of a BSP compiler for Win32 that visually depicts the progressive subdivision of space as the BSP tree is built. (Note that Listing 60.1 might not compile as printed; I may have missed copying some global variables that it uses.) The complete code is too large to print here in its entirety, but it’s on the CD-ROM in file DDJBSP.ZIP.

    -

    Optimizing the BSP Tree

    +

    Optimizing the BSP Tree

    In the previous chapter, I promised that I’d discuss how to go about deciding which wall to use as the splitter at each node in constructing a BSP tree. That turns out to be a far more difficult problem than one might think, but we can’t ignore it, because the choice of splitter can make a huge difference.

    @@ -53,7 +46,7 @@

    In Listing 60.1, I’ve applied the popular heuristic of choosing as the splitter at each node the surface that splits the fewest of the other surfaces that are being considered for that node. In other words, I choose the wall that splits the fewest of the walls in the subspace it’s subdividing.

    -

    BSP Optimization: an Undiscovered Country

    +

    BSP Optimization: an Undiscovered Country

    Although BSP trees have been around for at least 15 years now, they’re still only partially understood and are a ripe area for applied research and general ingenuity. You might want to try your hand at inventing new BSP optimization approaches; it’s an interesting problem, and you might strike paydirt. There are many things that BSP trees can’t do well, because it takes so long to build them—but what they do, they do exceedingly well, so a better compilation approach that allowed BSP trees to be used for more purposes would be valuable, indeed.

    @@ -74,10 +67,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/61-01.html b/61-01.html index d59964d..26e52d9 100644 --- a/61-01.html +++ b/61-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - @@ -37,10 +30,10 @@


    -

    Chapter 61
    +

    Chapter 61
    Frames of Reference

    -

    The Fundamentals of the Math behind 3-D Graphics

    +

    The Fundamentals of the Math behind 3-D Graphics

    Several years ago, I opened a column in Dr. Dobb’s Journal with a story about singing my daughter to sleep with Beatles’ songs. Beatles’ songs, at least the earlier ones, tend to be bouncy and pleasant, which makes them suitable goodnight fodder—and there are a lot of them, a useful hedge against terminal boredom. So for many good reasons, “Can’t Buy Me Love” and “A Hard Day’s Night” and “Help!” and the rest were evening staples for years.

    @@ -54,7 +47,7 @@

    Before we can talk about transforming between coordinate spaces, however, we need two building blocks: dot products and cross products.

    -

    3-D Math

    +

    3-D Math

    At this point in the book, I was originally going to present a BSP-based renderer, to complement the BSP compiler I presented in the previous chapter. What changed my plans was the considerable amount of mail about 3-D math that I’ve gotten in recent months. In every case, the writer has bemoaned his/her lack of expertise with 3-D math, and has asked what books about 3-D math I’d recommend, and how else he/she could learn more.

    @@ -62,7 +55,7 @@

    The other thing the mail made clear was that there are a lot of people out there who don’t understand either type of product, at least insofar as they apply to 3-D. Since much or even most advanced 3-D graphics machinery relies to a greater or lesser extent on dot products and cross products (even the line intersection formula I discussed in the last chapter is actually a quotient of dot products), I’m going to spend this chapter examining these basic tools and some of their 3-D applications. If this is old hat to you, my apologies, and I’ll return to BSP-based rendering in the next chapter.

    -

    Foundation Definitions

    +

    Foundation Definitions

    The dot and cross products themselves are straightforward and require almost no context to understand, but I need to define some terms I’ll use when describing applications of the products, so I’ll do that now, and then get started with dot products.

    @@ -95,10 +88,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/61-02.html b/61-02.html index a45c330..9266d16 100644 --- a/61-02.html +++ b/61-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - @@ -45,7 +38,7 @@

    For additional information, you might want to check out Foley & van Dam’s Computer Graphics (ISBN 0-201-12110-7), or the chapters in this book dealing with my X-Sharp 3-D graphics library.

    -

    The Dot Product

    +

    The Dot Product

    Now we’re ready to move on to the dot product. Given two vectors U = [u1 u2 u3] and V = [v1 v2 v3], their dot product, denoted by the symbol •, is calculated as:

    @@ -67,7 +60,7 @@

    where q is the angle between the two vectors, and the other two terms are the lengths of the vectors, as shown in Figure 61.1. Although it’s not immediately obvious, equation 3 has a wide variety of applications in 3-D graphics.

    -

    Dot Products of Unit Vectors

    +

    Dot Products of Unit Vectors

    The simplest case of the dot product is when both vectors are unit vectors; that is, when their lengths are both one, as calculated as in Equation 1. In this case, equation 3 simplifies to:

    @@ -87,9 +80,8 @@

    (eq. 5)

    -


    - Figure 61.1
      The dot product.

    +


    + Figure 61.1
      The dot product.

    where Is is the intensity of illumination of the surface, Il is the intensity of the light, and q is the angle between -Dl (where Dl is the light direction vector) and the surface normal. If the inverse light vector and the surface normal are both unit vectors, then this calculation can be performed with four multiplies and three additions—and no explicit cosine calculations—as

    @@ -101,15 +93,14 @@

    where Ns is the surface unit normal and Dl is the light unit direction vector, as shown in Figure 61.2.

    -

    Cross Products and the Generation of Polygon Normals

    +

    Cross Products and the Generation of Polygon Normals

    One question equation 6 begs is where the surface unit normal comes from. One approach is to store the end of a surface normal as an extra data point with each polygon (with the start being some point that’s already in the polygon), and transform it along with the rest of the points. This has the advantage that if the normal starts out as a unit normal, it will end up that way too, if only rotations and translations (but not scaling and shears) are performed.

    The problem with having an explicit normal is that it will remain a normal—that is, perpendicular to the surface—only through viewspace. Rotation, translation, and scaling preserve right angles, which is why normals are still normals in viewspace, but perspective projection does not preserve angles, so vectors that were surface normals in viewspace are no longer normals in screenspace.

    -


    - Figure 61.2
      The dot product as used in calculating lighting intensity.

    +


    + Figure 61.2
      The dot product as used in calculating lighting intensity.


    @@ -128,10 +119,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/61-03.html b/61-03.html index 90cde25..26bc3b0 100644 --- a/61-03.html +++ b/61-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - @@ -39,9 +32,8 @@

    Why does this matter? It matters because, on average, half the polygons in any scene are facing away from the viewer, and hence shouldn’t be drawn. One way to identify such polygons is to see whether they’re facing toward or away from the viewer; that is, whether their normals have negative z values (so they’re visible) or positive z values (so they should be culled). However, we’re talking about screenspace normals here, because the perspective projection can shift a polygon relative to the viewpoint so that although its viewspace normal has a negative z, its screenspace normal has a positive z, and vice-versa, as shown in Figure 61.3. So we need screenspace normals, but those can’t readily be generated by transformation from worldspace.

    -


    - Figure 61.3
      A problem with determining front/back visibility.

    +


    + Figure 61.3
      A problem with determining front/back visibility.

    The solution is to use the cross product of two of the polygon’s edges to generate a normal. The formula for the cross product is:

    @@ -61,15 +53,14 @@ -


    - Figure 61.4
      How the cross product of polygon edge vectors generates a polygon normal.

    +


    + Figure 61.4
      How the cross product of polygon edge vectors generates a polygon normal.

    Perhaps the most often asked question about cross products is “Which way do normals generated by cross products go?” In a left-handed coordinate system, curl the fingers of your left hand so the fingers curl through an angle of less than 180 degrees from the first vector in the cross product to the second vector. Your thumb now points in the direction of the normal.

    If you take the cross product of two orthogonal (right-angle) unit vectors, the result will be a unit vector that’s orthogonal to both of them. This means that if you’re generating a new coordinate space—such as a new viewing frame of reference—you only need to come up with unit vectors for two of the axes for the new coordinate space, and can then use their cross product to generate the unit vector for the third axis. If you need unit normals, and the two vectors being crossed aren’t orthogonal unit vectors, you’ll have to normalize the resulting vector; that is, divide each of the vector’s components by the length of the vector, to make it a unit long.

    -

    Using the Sign of the Dot Product

    +

    Using the Sign of the Dot Product

    The dot product is the cosine of the angle between two vectors, scaled by the magnitudes of the vectors. Magnitudes are always positive, so the sign of the cosine determines the sign of the result. The dot product is positive if the angle between the vectors is less than 90 degrees, negative if it’s greater than 90 degrees, and zero if the angle is exactly 90 degrees. This means that just the sign of the dot product suffices for tests involving comparisons of angles to 90 degrees, and there are more of those than you’d think.

    @@ -79,9 +70,8 @@

    Backface culling with the dot product is just a special case of determining which side of a plane any point (in this case, the viewpoint) is on. The same trick can be applied whenever you want to determine whether a point is in front of or behind a plane, where a plane is described by any point that’s on the plane (which I’ll call the plane origin), plus a plane normal. One such application is in clipping a line (such as a polygon edge) to a plane. Just do a dot product between the plane normal and the vector from one line endpoint to the plane origin, and repeat for the other line endpoint. If the signs of the dot products are the same, no clipping is needed; if they differ, clipping is needed. And yes, the dot product is also the way to do the actual clipping; but before we can talk about that, we need to understand the use of the dot product for projection.

    -


    - Figure 61.5
      Backface culling with the dot product.

    +


    + Figure 61.5
      Backface culling with the dot product.


    @@ -100,10 +90,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/61-04.html b/61-04.html index 75e2155..b356a95 100644 --- a/61-04.html +++ b/61-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Frames of Reference - - @@ -37,7 +30,7 @@


    -

    Using the Dot Product for Projection

    +

    Using the Dot Product for Projection

    Consider Equation 3 again, but this time make one of the vectors, say V, a unit vector. Now the equation reduces to:

    @@ -49,20 +42,19 @@

    In other words, the result is the cosine of the angle between the two vectors, scaled by the magnitude of the non-unit vector. Now, consider that cosine is really just the length of the adjacent leg of a right triangle, and think of the non-unit vector as the hypotenuse of a right triangle, and remember that all sides of similar triangles scale equally. What it all works out to is that the value of the dot product of any vector with a unit vector is the length of the first vector projected onto the unit vector, as shown in Figure 61.6.

    -


    - Figure 61.6
      How the dot product with a unit vector performs a projection.

    +


    + Figure 61.6
      How the dot product with a unit vector performs a projection.

    -

    This unlocks all sorts of neat stuff. Want to know the distance from a point to a plane? Just dot the vector from the point P to the plane origin Op with the plane unit normal Np, to project the vector onto the normal, then take the absolute value

    +

    This unlocks all sorts of neat stuff. Want to know the distance from a point to a plane? Just dot the vector from the point P to the plane origin Op with the plane unit normal Np, to project the vector onto the normal, then take the absolute value

     distance = |(P - Op) • Np|
    -
    +

    as shown in Figure 61.7.

    Want to clip a line to a plane? Calculate the distance from one endpoint to the plane, as just described, and dot the whole line segment with the plane normal, to get the full length of the line along the plane normal. The ratio of the two dot products is then how far along the line from the endpoint the intersection point is; just move along the line segment by that distance from the endpoint, and you’re at the intersection point, as shown in Listing 61.1.

    -

    LISTING 61.1 L61_1.C

    +

    LISTING 61.1 L61_1.C

     // Given two line endpoints, a point on a plane, and a unit normal
     // for the plane, returns the point of intersection of the line
    @@ -93,15 +85,14 @@ void LineIntersectPlane (float *linestart, float *lineend,
        intersectpoint[1] = linestart[1] - vec1[1] * scale;
        intersectpoint[2] = linestart[1] - vec1[2] * scale;
     }
    -
    + -

    Rotation by Projection

    +

    Rotation by Projection

    We can use the dot product’s projection capability to look at rotation in an interesting way. Typically, rotations are represented by matrices. This is certainly a workable representation that encapsulates all aspects of transformation in a single object, and is ideal for concatenations of rotations and translations. One problem with matrices, though, is that many people, myself included, have a hard time looking at a matrix of sines and cosines and visualizing what’s actually going on. So when two 3-D experts, John Carmack and Billy Zelsnack, mentioned that they think of rotation differently, in a way that seemed more intuitive to me, I thought it was worth passing on.

    -


    - Figure 61.7
      Using the dot product to get the distance from a point to a plane.

    +


    + Figure 61.7
      Using the dot product to get the distance from a point to a plane.

    Their approach is this: Think of rotation as projecting coordinates onto new axes. That is, given that you have points in, say, worldspace, define the new coordinate space (viewspace, for example) you want to rotate to by a set of three orthogonal unit vectors defining the new axes, and then project each point onto each of the three axes to get the coordinates in the new coordinate space, as shown for the 2-D case in Figure 61.8. In 3-D, this involves three dot products per point, one to project the point onto each axis. Translation can be done separately from rotation by simple addition.

    @@ -115,9 +106,8 @@ void LineIntersectPlane (float *linestart, float *lineend,

    Three things I’ve learned over the years are that it never hurts to learn a new way of looking at things, that it helps to have a clearer, more intuitive model in your head of whatever it is you’re working on, and that new tools, or new ways to use old tools, are Good Things. My experience has been that rotation by projection, and dot product tricks in general, offer those sorts of benefits for 3-D.

    -


    - Figure 61.8
      Rotation to a new coordinate space by projection onto new axes.

    +


    + Figure 61.8
      Rotation to a new coordinate space by projection onto new axes.


    @@ -136,10 +126,6 @@ void LineIntersectPlane (float *linestart, float *lineend,
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/62-01.html b/62-01.html index 8b35b29..13bf7c8 100644 --- a/62-01.html +++ b/62-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - @@ -37,10 +30,10 @@


    -

    Chapter 62
    +

    Chapter 62
    One Story, Two Rules, and a BSP Renderer

    -

    Taking a Compiled BSP Tree from Logical to Visual Reality

    +

    Taking a Compiled BSP Tree from Logical to Visual Reality

    As I’ve noted before, I’m working on Quake, id Software’s follow-up to DOOM. A month or so back, we added page flipping to Quake, and made the startling discovery that the program ran nearly twice as fast with page flipping as it did with the alternative method of drawing the whole frame to system memory, then copying it to the screen. We were delighted by this, but baffled. I did a few tests and came up with several possible explanations, including slow writes through the external cache, poor main memory performance, and cache misses when copying the frame from system memory to video memory. Although each of these can indeed affect performance, none seemed to account for the magnitude of the speedup, so I assumed there was some hidden hardware interaction at work. Anyway, “why” was secondary; what really mattered was that we had a way to double performance, which meant I had a lot of work to do to support page flipping as widely as possible.

    @@ -66,15 +59,14 @@

    Onward to rendering from a BSP tree.

    -

    BSP-based Rendering

    +

    BSP-based Rendering

    For the last several chapters I’ve been discussing the nature of BSP (Binary Space Partitioning) trees, and in Chapter 60 I presented a compiler for 2-D BSP trees. Now we’re ready to use those compiled BSP trees to do realtime rendering.

    As you’ll recall, the BSP compiler took a list of vertical walls and built a 2-D BSP tree from the walls, as viewed from above. The result is shown in Figure 62.1. The world is split into two pieces by the line of the root wall, and each half of the world is then split again by the root’s children, and so on, until the world is carved into subspaces along the lines of all the walls.

    -


    - Figure 62.1
      Vertical walls and a BSP tree to represent them.

    +


    + Figure 62.1
      Vertical walls and a BSP tree to represent them.

    Our objective is to draw the world so that whenever walls overlap we see the nearer wall at each overlapped pixel. The simplest way to do that is with the painter’s algorithm; that is, drawing the walls in back-to-front order, assuming no polygons interpenetrate or form cycles. BSP trees guarantee that no polygons interpenetrate (such polygons are automatically split), and make it easy to walk the polygons in back-to-front (or front-to-back) order.

    @@ -97,10 +89,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/62-02.html b/62-02.html index abec8d8..060f748 100644 --- a/62-02.html +++ b/62-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - @@ -37,7 +30,7 @@


    -

    Listing 62.1 L62_1.C

    +

    Listing 62.1 L62_1.C

     /* Core renderer for Win32 program to demonstrate drawing from a 2-D
        BSP tree; illustrate the use of BSP trees for surface visibility.
    @@ -478,7 +471,7 @@ void UpdateWorld()
        ReleaseDC(hwndOutput, hdcDIBSection);
        iteration++;
     }
    -
    +


    @@ -497,10 +490,6 @@ void UpdateWorld()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/62-03.html b/62-03.html index d2dd2bb..93c720c 100644 --- a/62-03.html +++ b/62-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - @@ -37,7 +30,7 @@


    -

    The Rendering Pipeline

    +

    The Rendering Pipeline

    Conceptually rendering from a BSP tree really is that simple, but the implementation is a bit more complicated. The full rendering pipeline, as coordinated by UpdateWorld(), is this:

    @@ -55,13 +48,13 @@

    Next, we’ll look at each part of the pipeline more closely. The pipeline is too complex for me to be able to discuss each part in complete detail. Some sources for further reading are Computer Graphics, by Foley and van Dam (ISBN 0-201-12110-7), and the DDJ Essential Books on Graphics Programming CD.

    -

    Moving the Viewer

    +

    Moving the Viewer

    The sample BSP program performs first-person rendering; that is, it renders the world as seen from your eyes as you move about. The rate of movement is controlled by key-handling code that’s not shown in Listing 62.1; however, the variables set by the key-handling code are used in UpdateViewPos() to bring the current location up to date.

    Note that the view position can change not only in x and z (movement around the but only viewing horizontally. Although the BSP tree is only 2-D, it is quite possible to support looking up and down to at least some extent, particularly if the world dataset is restricted so that, for example, there are never two rooms stacked on top of each other, or any tilted walls. For simplicity’s sake, I have chosen not to implement this in Listing 62.1, but you may find it educational to add it to the program yourself.

    -

    Transformation into Viewspace

    +

    Transformation into Viewspace

    The viewing angle (which controls direction of movement as well as view direction) can sweep through the full 360 degrees around the viewpoint, so long as it remains horizontal. The viewing angle is controlled by the key handler, and is used to define a unit vector stored in currentorientation that explicitly defines the view direction (the z axis of viewspace), and implicitly defines the x axis of viewspace, because that axis is at right angles to the z axis, where x increases to the right of the viewer.

    @@ -71,7 +64,7 @@

    When this is done the walls are in viewspace, ready to be clipped.

    -

    Clipping

    +

    Clipping

    In viewspace, the walls may be anywhere relative to the viewpoint: in front, behind, off to the side. We only want to draw those parts of walls that properly belong on the screen; that is, those parts that lie in the view pyramid (view frustum), as shown in Figure 62.2. Unclipped walls—walls that lie entirely in the frustum—should be drawn in their entirety, fully clipped walls should not be drawn, and partially clipped walls must be trimmed before being drawn.

    @@ -79,19 +72,17 @@

    After clipping in z, we clip by viewspace x coordinate, to ensure that we draw only wall portions that lie between the left and right edges of the screen. Like z-clipping, x-clipping can be done as a 2-D clip, because the walls and the left and right sides of the frustum are all vertical. We compare both the start and endpoint of each wall to the left and right sides of the frustum, and reject, accept, or clip each wall’s t values accordingly. The test for x clipping is very simple, because the edges of the frustum are defined as the planes where x==z and -x==z.

    -


    - Figure 62.2
      Clipping to the view pyramid.

    +


    + Figure 62.2
      Clipping to the view pyramid.

    The final clip stage is clipping by y coordinate, and this is the most complicated, because vertical walls can be clipped at an angle in y, as shown in Figure 62.3, so true 3-D clipping of all four wall vertices is involved. We handle this in ClipWalls() by detecting trivial rejection in y, using y==z and ==z as the y boundaries of the frustum. However, we leave partial clipping to be handled as a 2-D clipping problem; we are able to do this only because our earlier z-clip to the near clip plane guarantees that no remaining polygon point can have z<=0, ensuring that when we project we’ll always pass valid, y-clippable screenspace vertices to the polygon filler.

    -

    Projection to Screenspace

    +

    Projection to Screenspace

    At this point, we have viewspace vertices for each wall that’s at least partially visible. All we have to do is project these vertices according to z distance—that is, perform perspective projection—and scale the results to the width of the screen, then we’ll be ready to draw. Although this step is logically separate from clipping, it is performed as the last step for visible walls in ClipWalls().

    -


    - Figure 62.3
      Why y clipping is more complex than x or z clipping.

    +


    + Figure 62.3
      Why y clipping is more complex than x or z clipping.


    @@ -110,10 +101,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/62-04.html b/62-04.html index 8f704e0..6381196 100644 --- a/62-04.html +++ b/62-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: On Story, Two Rules, and a BSP Renderer - - @@ -37,7 +30,7 @@


    -

    Walking the Tree, Backface Culling and Drawing

    +

    Walking the Tree, Backface Culling and Drawing

    Now that we have all the walls clipped to the frustum, with vertices projected into screen coordinates, all we have to do is draw them back to front; that’s the job of DrawWallsBackToFront(). Basically, this routine walks the BSP tree, descending recursively from each node to draw the farther children of each node first, then the wall at the node, then the nearer children. In the interests of efficiency, this particular implementation performs a data-recursive walk of the tree, rather than the more familiar code recursion. Interestingly, the performance speedup from data recursion turned out to be more modest than I had expected, based on past experience; see Chapter 59 for further details.

    @@ -51,11 +44,10 @@

    All the visible, front-facing walls are drawn into a buffer by DrawWallsBackToFront(), then UpdateWorld() calls Win32 to copy the new frame to the screen. The frame of animation is complete.

    -


    - Figure 62.4
      Fast backspace culling test in screenspace.

    +


    + Figure 62.4
      Fast backspace culling test in screenspace.

    -

    Notes on the BSP Renderer

    +

    Notes on the BSP Renderer

    Listing 62.1 is far from complete or optimal. There is no such thing as a tiny BSP rendering demo, because 3D rendering, even when based on a 2-D BSP tree, requires a substantial amount of code and complexity. Listing 62.1 is reasonably close to a minimum rendering engine, and is specifically intended to illuminate basic BSP principles, given the space limitations of one chapter in a book that’s already larger than it should be. Think of Listing 62.1 as a learning tool and a starting point.

    @@ -82,10 +74,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/63-01.html b/63-01.html index 47b8203..177f698 100644 --- a/63-01.html +++ b/63-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - @@ -37,10 +30,10 @@


    -

    Chapter 63
    +

    Chapter 63
    Floating-Point for Real-Time 3-D

    -

    Knowing When to Hurl Conventional Math Wisdom Out the Window

    +

    Knowing When to Hurl Conventional Math Wisdom Out the Window

    In a crisis, sometimes it’s best to go with the first solution that comes into your head—but not very often.

    @@ -66,7 +59,7 @@

    For example, consider floating-point math.

    -

    Not Your Father’s Floating-Point

    +

    Not Your Father’s Floating-Point

    Until last year, I had never done any serious floating-point (FP) optimization, for the perfectly good reason that FP math had never been fast enough for any of the code I needed to write. It was an article of faith that FP, while undeniably convenient, because of its automatic support for constant precision over an enormous range of magnitudes, was just not fast enough for real-time programming, so I, like pretty much everyone else doing 3-D, expended a lot of time and effort in making fixed-point do the job.

    @@ -78,7 +71,7 @@

    By way of getting you started with floating-point for real-time 3-D, in this chapter I’ll examine the basics of Pentium FP optimization, then look at how some key mathematical techniques for 3-D—dot product, cross product, transformation, and projection—can be accelerated.

    -

    Pentium Floating-Point Optimization

    +

    Pentium Floating-Point Optimization

    I’m going to assume you’re already familiar with x86 FP code in general; for additional information, check out Intel’s Pentium Processor User’s Manual (order #241430-001; 1-800-548-4725), a book that you should have if you’re doing Pentium programming of any sort. I’d also recommend taking a look around http://www.intel.com.

    @@ -101,10 +94,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/63-02.html b/63-02.html index 2c1a11e..7d63d7b 100644 --- a/63-02.html +++ b/63-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - @@ -37,68 +30,68 @@


    -

    FDIV is a painfully slow instruction, taking 39 cycles at full precision and 33 cycles at double precision, which is the default precision for Visual C++ 2.0. While FDIV executes, the FPU is occupied, and can’t process subsequent FP instructions until FDIV finishes. However, during the cycles while FDIV is executing (with the exception of the one cycle during which FDIV starts), the integer unit can simultaneously execute instructions other than IMUL. (IMUL uses the FPU, and can only overlap with FDIV for a few cycles.) Since the integer unit can execute two instructions per cycle, this means it’s possible to have three instructions, an FDIV and two integer instructions, executing at the same time. That’s exactly what happens, for example, during the second cycle of this code:

    +

    FDIV is a painfully slow instruction, taking 39 cycles at full precision and 33 cycles at double precision, which is the default precision for Visual C++ 2.0. While FDIV executes, the FPU is occupied, and can’t process subsequent FP instructions until FDIV finishes. However, during the cycles while FDIV is executing (with the exception of the one cycle during which FDIV starts), the integer unit can simultaneously execute instructions other than IMUL. (IMUL uses the FPU, and can only overlap with FDIV for a few cycles.) Since the integer unit can execute two instructions per cycle, this means it’s possible to have three instructions, an FDIV and two integer instructions, executing at the same time. That’s exactly what happens, for example, during the second cycle of this code:

     FDIV ST(0),ST(1)
     ADD  EAX,ECX
     INC  EDX
    -
    +

    There’s an important limitation, though; if the instruction stream following the FDIV reaches a FP instruction (or an IMUL), then that instruction and all subsequent instructions, both integer and FP, must wait to execute until FDIV has finished.

    -

    When a FADD, FSUB, or FMUL instruction is executed, it is 3 cycles before the result can be used by another instruction. (There’s an exception: If the instruction that attempts to use the result is an FST to memory, there’s an extra cycle lost, so it’s 4 cycles from the start of an arithmetic instruction until an FST of that value can begin, so

    +

    When a FADD, FSUB, or FMUL instruction is executed, it is 3 cycles before the result can be used by another instruction. (There’s an exception: If the instruction that attempts to use the result is an FST to memory, there’s an extra cycle lost, so it’s 4 cycles from the start of an arithmetic instruction until an FST of that value can begin, so

     FMUL ST(0),ST(1)
     FST  [temp]
    -
    +

    takes 6 cycles in all.) Again, it’s possible to execute integer-unit instructions during the 2 (or 3, for FST) cycles after one of these FP instructions starts. There’s a more exciting possibility here, though: Given properly structured code, the FPU is capable of averaging 1 cycle per FADD, FSUB, or FMUL. The secret is pipelining.

    -

    Pipelining, Latency, and Throughput

    +

    Pipelining, Latency, and Throughput

    -

    The Pentium’s FPU is the first pipelined x86 FPU. Pipelining means that the FPU is capable of starting an instruction every cycle, and can simultaneously handle several instructions in various stages of completion. Only certain x86 FP instructions allow another instruction to start on the next cycle, though: FADD, FSUB, and FMUL are pipelined, but FST and FDIV are not. (FLD executes in a single cycle, so pipelining is not an issue.) Thus, in the code sequence

    +

    The Pentium’s FPU is the first pipelined x86 FPU. Pipelining means that the FPU is capable of starting an instruction every cycle, and can simultaneously handle several instructions in various stages of completion. Only certain x86 FP instructions allow another instruction to start on the next cycle, though: FADD, FSUB, and FMUL are pipelined, but FST and FDIV are not. (FLD executes in a single cycle, so pipelining is not an issue.) Thus, in the code sequence

     FADD1
     FSUB
     FADD2
     FMUL
    -
    + -

    FADD1 can start on cycle N, FSUB can start on cycle N+1, FADD2 can start on cycle N+2, and FMUL can start on cycle N+3. At the start of cycle N+3, the result of FADD1 is available in the destination operand, because it’s been 3 cycles since the instruction started; FSUB is starting the final cycle of calculation; FADD2 is starting its second cycle, with one cycle yet to go after this; and FMUL is about to be issued. Each of the instructions takes 3 cycles to produce a result from the time it starts, but because they’re simultaneously processed at different pipeline stages, one instruction is issued and one instruction completes every cycle. Thus, the latency of these instructions—that is, the time until the result is available—is 3 cycles, but the throughput—the rate at which the FPU can start new instructions—is 1 cycle. An exception is that the FPU is capable of starting an FMUL only every 2 cycles, so between these two instructions

    +

    FADD1 can start on cycle N, FSUB can start on cycle N+1, FADD2 can start on cycle N+2, and FMUL can start on cycle N+3. At the start of cycle N+3, the result of FADD1 is available in the destination operand, because it’s been 3 cycles since the instruction started; FSUB is starting the final cycle of calculation; FADD2 is starting its second cycle, with one cycle yet to go after this; and FMUL is about to be issued. Each of the instructions takes 3 cycles to produce a result from the time it starts, but because they’re simultaneously processed at different pipeline stages, one instruction is issued and one instruction completes every cycle. Thus, the latency of these instructions—that is, the time until the result is available—is 3 cycles, but the throughput—the rate at which the FPU can start new instructions—is 1 cycle. An exception is that the FPU is capable of starting an FMUL only every 2 cycles, so between these two instructions

     FMUL ST(1),ST(0)
     FMUL ST(2),ST(0)
    -
    + -

    there’s a 1-cycle stall, and the following three instructions execute just as fast as the above pair:

    +

    there’s a 1-cycle stall, and the following three instructions execute just as fast as the above pair:

     FMUL ST(1),ST(0)
     FLD  ST(4)
     FMUL ST(0),ST(1)
    -
    + -

    There’s a caveat here, though: A FP instruction can’t be issued until its operands are available. The FPU can reach a throughput of 1 cycle per instruction on this code

    +

    There’s a caveat here, though: A FP instruction can’t be issued until its operands are available. The FPU can reach a throughput of 1 cycle per instruction on this code

     FADD ST(1),ST(0)
     FLD  [temp]
     FSUB ST(1),ST(0)
    -
    + -

    because neither the FLD nor the FSUB needs the result from the FADD. Consider, however

    +

    because neither the FLD nor the FSUB needs the result from the FADD. Consider, however

     FADD ST(0),ST(2)
     FSUB ST(0),ST(1)
    -
    +

    where the ST(0) operand to FSUB is calculated by FADD. Here, FSUB can’t start until FADD has completed, so there are 2 stall cycles between the two instructions. When dependencies like this occur, the FPU runs at latency rather than throughput speeds, and performance can drop by as much as two-thirds.

    -

    FXCH

    +

    FXCH

    One piece of the puzzle is still missing. Clearly, to get maximum throughput, we need to interleave FP instructions, such that at any one time ideally three instructions are in the pipeline at once. Further, these instructions must not depend on one another for operands. But ST(0) must always be one of the operands; worse, FLD can only push into ST(0), and FST can only store from ST(0). How, then, can we keep three independent instructions going?

    The easy answer would be for Intel to change the FP registers from a stack to a set of independent registers. Since they couldn’t do that, thanks to compatibility issues, they did the next best thing: They made the FXCH instruction, which swaps ST(0) and any other FP register, virtually free. In general, if FXCH is both preceded and followed by FP instructions, then it takes no cycles to execute. (Application Note 500, “Optimizations for Intel’s 32-bit Processors,” February 1994, available from http://www.intel.com, describes all .the conditions under which FXCH is free.) This allows you to move the target of a pending operation from ST(0) to another register, at the same time bringing another register into ST(0) where it can be used, all at no cost. So, for example, we can start three multiplications, then use FXCH to swap back to start adding the results of the first two multiplications, without incurring any stalls, as shown in Listing 63.1.

    -

    Listing 63.1 L63-1.ASM

    +

    Listing 63.1 L63-1.ASM

     ; use of fxch to allow addition of first two; products to start while third : multiplication finishes
             fld     [vec0+0]        ;starts & ends on cycle 0
    @@ -109,9 +102,9 @@ FSUB ST(0),ST(1)
             fmul    [vec1+8]        ;starts on cycle 5
             fxch    st(1)           ;no cost
             faddp   st(2),st(0)     ;starts on cycle 6
    -
    + -

    The Dot Product

    +

    The Dot Product

    Now we’re ready to look at fast FP for common 3-D operations; we’ll start by looking at how to speed up the dot product. As discussed in Chapter 30, the dot product is heavily used in 3-D to calculate cosines and to project points along vectors. The dot product is calculated as d = u1v1 + u2v2 + u3v3; with three loads, three multiplies, two adds, and a store, the theoretical minimum time for this calculation is 10 cycles.

    @@ -134,10 +127,6 @@ FSUB ST(0),ST(1)
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/63-03.html b/63-03.html index 66dbc34..4d10c54 100644 --- a/63-03.html +++ b/63-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - @@ -37,7 +30,7 @@


    -

    Listing 63.2 1 L63-2.ASM

    +

    Listing 63.2 1 L63-2.ASM

     ; unoptimized dot product; 17 cycles
             fld     [vec0+0]        ;starts & ends on cycle 0
    @@ -53,9 +46,9 @@
                                     ;stalls for cycles 12-14
             fstp    [dot]           ;starts on cycle 15,
                                     ; ends on cycle 16
    -
    + -

    Listing 63.3 L63-3.ASM

    +

    Listing 63.3 L63-3.ASM

     ;optimized dot product; 15 cycles
             fld     [vec0+0]        ;starts & ends on cycle 0
    @@ -71,13 +64,13 @@
                                     ;stalls for cycles 10-12
             fstp    [dot]           ;starts on cycle 13,
                                     ; ends on cycle 14
    -
    + -

    The Cross Product

    +

    The Cross Product

    When last we looked at the cross product, we found that it’s handy for generating a vector that’s normal to two other vectors. The cross product is calculated as [u2v3-u3v2 u3v1-u1v3 u1v2-u2v1]. The theoretical minimum cycle count for the cross product is 21 cycles. Listing 63.4 shows a straightforward implementation that calculates each component of the result separately, losing 15 cycles to stalls.

    -

    Listing 63.4 L63-4.ASM

    +

    Listing 63.4 L63-4.ASM

     ;unoptimized cross product; 36 cycles
             fld     [vec0+4]         ;starts & ends on cycle 0
    @@ -108,11 +101,11 @@
                                      ;stalls for cycles 31-33
             fstp    [vec2+8]         ;starts on cycle 34,
                                      ; ends on cycle 35
    -
    +

    We couldn’t get rid of many of the stalls in the dot product code because with six inputs and one output, it was impossible to interleave all the operations. However, the cross product, with three outputs, is much more amenable to optimization. In fact, three is the magic number; because we have three calculation streams and the latency of FADD, FSUB, and FMUL is 3 cycles, we can eliminate almost every single stall in the cross-product calculation, as shown in Listing 63.5. Listing 63.5 loses only one cycle to a stall, the cycle before the first FST; the relevant FSUB has just finished on the preceding cycle, so we run into the extra cycle of latency associated with FST. Listing 63.5 is more than 60 percent faster than Listing 63.4, a striking illustration of the power of properly managing the Pentium’s FP pipeline.

    -

    Listing 63.5 L63-5.ASM

    +

    Listing 63.5 L63-5.ASM

     ;optimized cross product; 22 cycles
             fld       [vec0+4]        ;starts & ends on cycle 0
    @@ -139,13 +132,13 @@
                                       ; ends on cycle 19
             fstp      [vec2+8]        ;starts on cycle 20,
                                       ; ends on cycle 21
    -
    + -

    Transformation

    +

    Transformation

    Transforming a point, for example from worldspace to viewspace, is one of the most heavily used FP operations in realtime 3-D. Conceptually, transformation is nothing more than three dot products and three additions, as I will discuss in Chapter 61. (Note that I’m talking about a subset of a general 4x4 transformation matrix, where the fourth row is always implicitly [0 0 0 1]. This limited form suffices for common transformations, and does 25 percent less work than a full 4x4 transformation.)

    -

    Transformation is calculated as:

    +

    Transformation is calculated as:

     -  -    -                -  -   -
      v1      m11 m12 m13 m14    u1
    @@ -153,14 +146,14 @@
      v3      m31 m32 m33 m34    u3
      1       0   0   0   1     1
     -  -    -                -  -   -
    -
    + -

    or

    +

    or

     v1 = m11u1 + m12u2 + m13u3 + m14
     v2 = m21u1 + m22u2 + m23u3 + m24
     v3 = m31u1 + m32u2 + m33u3 + m34.
    -
    +


    @@ -179,10 +172,6 @@ v3 = m31u1 + m32u2 + m Graphics Programming Black Book © 2001 Michael Abrash - - - - - + diff --git a/63-04.html b/63-04.html index 098149f..327476a 100644 --- a/63-04.html +++ b/63-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Floating-Point for Real-Time 3-D - - @@ -41,7 +34,7 @@

    When fully interleaved, however, only a single cycle is lost (again to the extra cycle of FST latency), and the cycle count drops to 34, as shown in Listing 63.6. This means that on a 100 MHz Pentium, it’s theoretically possible to do nearly 3,000,000 transforms per second, although that’s a purely hypothetical number, due to cache effects and set-up costs. Still, more than 1,000,000 transforms per second is certainly feasible; at a frame rate of 30 Hz, that’s an impressive 30,000 transforms per frame.

    -

    Listing 63.6 L63-6.ASM

    +

    Listing 63.6 L63-6.ASM

     ;optimized transformation: 34 cycles
           fld      [vec0+0]           ;starts & ends on cycle 0
    @@ -83,21 +76,21 @@
                                       ; ends on cycle 31
           fstp     [vec1+4]           ;starts on cycle 32,
                                       ; ends on cycle 33
    -
    + -

    Projection

    +

    Projection

    The final optimization we’ll look at is projection to screenspace. Projection itself is basically nothing more than a divide (to get 1/z), followed by two multiplies (to get x/z and y/z), so there wouldn’t seem to be much in the way of FP optimization possibilities there. However, remember that although FDIV has a latency of up to 39 cycles, it can overlap with integer instructions for all but one of those cycles. That means that if we can find enough independent integer work to do before we need the 1/z result, we can effectively reduce the cost of the FDIV to one cycle. Projection by itself doesn’t offer much with which to overlap, but other work such as clamping, window-relative adjustments, or 2-D clipping could be interleaved with the FDIV for the next point.

    Another dramatic speed-up is possible by setting the precision of the FPU down to single precision via FLDCW, thereby cutting the time FDIV takes to a mere 19 cycles. I don’t have the space to discuss reduced precision in detail in this book, but be aware that along with potentially greater performance, it carries certain risks, as well. The reduced precision, which affects FADD, FSUB, FMUL, FDIV, and FSQRT, can cause subtle differences from the results you’d get using compiler defaults. If you use reduced precision, you should be on the alert for precision-related problems, such as clipped values that vary more than you’d expect from the precise clip point, or the need for using larger epsilons in comparisons for point-on-plane tests.

    -

    Rounding Control

    +

    Rounding Control

    Another useful area that I can note only in passing here is that of leaving the FPU in a particular rounding mode while performing bulk operations of some sort. For example, conversion to int via the FIST instruction requires that the FPU be in chop mode. Unfortunately, the FLDCW instruction must be used to get the FPU into and out of chop mode, and each FLDCW takes 7 cycles, meaning that compilers often take at least 14 cycles for each float->int conversion. In assembly, you can just set the rounding state (or, likewise, the precision, for faster FDIVs) once at the start of the loop, and save all those FLDCW cycles each time through the loop. This is even more true for ceil(), which many compilers implement as horrendously inefficient subroutines, even though there are rounding modes for both ceil() and floor(). Again, though, be aware that results of FP calculations will be subtly different from compiler default behavior while chop, ceil, or floor mode is in effect.

    A final note: There are some speed-ups to be had by manipulating FP variables with integer instructions. Check out Chris Hecker’s column in the February/March 1996 issue of Game Developer for details.

    -

    A Farewell to 3-D Fixed-Point

    +

    A Farewell to 3-D Fixed-Point

    As with most optimizations, there are both benefits and hazards to floating-point acceleration, especially pedal-to-the-metal optimizations such as the last few I’ve mentioned. Nonetheless, I’ve found floating-point to be generally both more robust and easier to use than fixed-point even with those maximum optimizations. Now that floating-point is fast enough for real time, I don’t expect to be doing a whole lot of fixed-point 3-D math from here on out.

    @@ -120,10 +113,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/64-01.html b/64-01.html index ae53870..c391098 100644 --- a/64-01.html +++ b/64-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - @@ -37,10 +30,10 @@


    -

    Chapter 64
    +

    Chapter 64
    Quake’s Visible-Surface Determination

    -

    The Challenge of Separating All Things Seen from All Things Unseen

    +

    The Challenge of Separating All Things Seen from All Things Unseen

    Years ago, I was working at Video Seven, a now-vanished video adapter manufacturer, helping to develop a VGA clone. The fellow who was designing Video Seven’s VGA chip, Tom Wilson, had worked around the clock for months to make his VGA run as fast as possible, and was confident he had pretty much maxed out its performance. As Tom was putting the finishing touches on his chip design, however, news came fourth-hand that a competitor, Paradise, had juiced up the performance of the clone they were developing by putting in a FIFO.

    @@ -66,7 +59,7 @@

    Case in point: The evolution of Quake’s 3-D graphics engine.

    -

    VSD: The Toughest 3-D Challenge of All

    +

    VSD: The Toughest 3-D Challenge of All

    I’ve spent most of my waking hours for the last several months working on Quake, id Software’s successor to DOOM, and I suspect I have a few more months to go. The very best things don’t happen easily, nor quickly—but when they happen, all the sweat becomes worthwhile.

    @@ -93,10 +86,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/64-02.html b/64-02.html index 70b8ae1..d4cfd69 100644 --- a/64-02.html +++ b/64-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - @@ -37,23 +30,21 @@


    -

    The Structure of Quake Levels

    +

    The Structure of Quake Levels

    Before diving into VSD, let me note that each Quake level is stored as a single huge 3-D BSP tree. This BSP tree, like any BSP, subdivides space, in this case along the planes of the polygons. However, unlike the BSP tree I presented in Chapter 62, Quake’s BSP tree does not store polygons in the tree nodes, as part of the splitting planes, but rather in the empty (non-solid) leaves, as shown in overhead view in Figure 64.1.

    Correct drawing order can be obtained by drawing the leaves in front-to-back or back-to-front BSP order, again as discussed in Chapter 62. Also, because BSP leaves are always convex and the polygons are on the boundaries of the BSP leaves, facing inward, the polygons in a given leaf can never obscure one another and can be drawn in any order. (This is a general property of convex polyhedra.)

    -

    Culling and Visible Surface Determination

    +

    Culling and Visible Surface Determination

    The process of VSD would ideally work as follows: First, you would cull all polygons that are completely outside the view frustum (view pyramid), and would clip away the irrelevant portions of any polygons that are partially outside. Then, you would draw only those pixels of each polygon that are actually visible from the current viewpoint, as shown in overhead view in Figure 64.2, wasting no time overdrawing pixels multiple times; note how little of the polygon sets in Figure 64.2 actually need to be drawn. Finally, in a perfect world, the tests to figure out what parts of which polygons are visible would be free, and the processing time would be the same for all possible viewpoints, giving the game a smooth visual flow.

    -


    - Figure 64.1
      Quake’s polygons are stored as empty leaves.

    +


    + Figure 64.1
      Quake’s polygons are stored as empty leaves.

    -


    - Figure 64.2
      Pixels visible from the current viewpoint.

    +


    + Figure 64.2
      Pixels visible from the current viewpoint.

    As it happens, it is easy to determine which polygons are outside the frustum or partially clipped, and it’s quite possible to figure out precisely which pixels need to be drawn. Alas, the world is far from perfect, and those tests are far from free, so the real trick is how to accelerate or skip various tests and still produce the desired result.

    @@ -61,19 +52,18 @@

    For relatively simple worlds, it is perfectly acceptable. It doesn’t scale very well, though. One problem is that as you add more polygons in the world, more transformations and tests have to be performed to cull polygons that aren’t visible; at some point, that will bog considerably performance down.

    -

    Nodes Inside and Outside the View Frustum

    +

    Nodes Inside and Outside the View Frustum

    Happily, there’s a good workaround for this particular problem. As discussed earlier, each leaf of a BSP tree represents a convex subspace, with the nodes that bound the leaf delimiting the space. Perhaps less obvious is that each node in a BSP tree also describes a subspace—the subspace composed of all the node’s children, as shown in Figure 64.3. Another way of thinking of this is that each node splits the subspace into two pieces created by the nodes above it in the tree, and the node’s children then further carve that subspace into all the leaves that descend from the node.

    -


    - Figure 64.3
      The substance described by node E.

    +


    + Figure 64.3
      The substance described by node E.

    Since a node’s subspace is bounded and convex, it is possible to test whether it is entirely outside the frustum. If it is, all of the node’s children are certain to be fully clipped and can be rejected without any additional processing. Since most of the world is typically outside the frustum, many of the polygons in the world can be culled almost for free, in huge, node-subspace chunks. It’s relatively expensive to perform a perfect test for subspace clipping, so instead bounding spheres or boxes are often maintained for each node, specifically for culling tests.

    So culling to the frustum isn’t a problem, and the BSP can be used to draw back-to- front. What, then, is the problem?

    -

    Overdraw

    +

    Overdraw

    The problem John Carmack, the driving technical force behind DOOM and Quake, faced when he designed Quake was that in a complex world, many scenes have an awful lot of polygons in the frustum. Most of those polygons are partially or entirely obscured by other polygons, but the painter’s algorithm described earlier requires that every pixel of every polygon in the frustum be drawn, often only to be overdrawn. In a 10,000-polygon Quake level, it would be easy to get a worst-case overdraw level of 10 times or more; that is, in some frames each pixel could be drawn 10 times or more, on average. No rasterizer is fast enough to compensate for an order of such magnitude and more work than is actually necessary to show a scene; worse still, the painter’s algorithm will cause a vast difference between best-case and worst-case performance, so the frame rate can vary wildly as the viewer moves around.

    @@ -100,10 +90,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/64-03.html b/64-03.html index b9bc172..9294f85 100644 --- a/64-03.html +++ b/64-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - @@ -37,49 +30,47 @@


    -

    The Beam Tree

    +

    The Beam Tree

    John’s original Quake design was to draw front-to-back, using a second BSP tree to keep track of what parts of the screen were already drawn and which were still empty and therefore drawable by the remaining polygons. Logically, you can think of this BSP tree as being a 2-D region describing solid and empty areas of the screen, as shown in Figure 64.4, but in fact it is a 3-D tree, of the sort known as a beam tree. A beam tree is a collection of 3-D wedges (beams), bounded by planes, projecting out from some center point, in this case the viewpoint, as shown in Figure 64.5.

    In John’s design, the beam tree started out consisting of a single beam describing the frustum; everything outside that beam was marked solid (so nothing would draw there), and the inside of the beam was marked empty. As each new polygon was reached while walking the world BSP tree front-to-back, that polygon was converted to a beam by running planes from its edges through the viewpoint, and any part of the beam that intersected empty beams in the beam tree was considered drawable and added to the beam tree as a solid beam. This continued until either there were no more polygons or the beam tree became entirely solid. Once the beam tree was completed, the visible portions of the polygons that had contributed to the beam tree were drawn.

    -


    - Figure 64.4
      Partitioning the screen into 2-D regions.

    +


    + Figure 64.4
      Partitioning the screen into 2-D regions.

    -


    - Figure 64.5
      Beams as wedges projecting from the viewpoint to polygon edges.

    +


    + Figure 64.5
      Beams as wedges projecting from the viewpoint to polygon edges.

    The advantage to working with a 3-D beam tree, rather than a 2-D region, is that determining which side of a beam plane a polygon vertex is on involves only checking the sign of the dot product of the ray to the vertex and the plane normal, because all beam planes run through the origin (the viewpoint). Also, because a beam plane is completely described by a single normal, generating a beam from a polygon edge requires only a cross-product of the edge and a ray from the edge to the viewpoint. Finally, bounding spheres of BSP nodes can be used to do the aforementioned bulk culling to the frustum.

    The early-out feature of the beam tree—stopping when the beam tree becomes solid—seems appealing, because it appears to cap worst-case performance. Unfortunately, there are still scenes where it’s possible to see all the way to the sky or the back wall of the world, so in the worst case, all polygons in the frustum will still have to be tested against the beam tree. Similar problems can arise from tiny cracks due to numeric precision limitations. Beam-tree clipping is fairly time-consuming, and in scenes with long view distances, such as views across the top of a level, the total cost of beam processing slowed Quake’s frame rate to a crawl. So, in the end, the beam-tree approach proved to suffer from much the same malady as the painter’s algorithm: The worst case was much worse than the average case, and it didn’t scale well with increasing level complexity.

    -

    3-D Engine du Jour

    +

    3-D Engine du Jour

    Once the beam tree was working, John relentlessly worked at speeding up the 3-D engine, always trying to improve the design, rather than tweaking the implementation. At least once a week, and often every day, he would walk into my office and say “Last night I couldn’t get to sleep, so I was thinking...” and I’d know that I was about to get my mind stretched yet again. John tried many ways to improve the beam tree, with some success, but more interesting was the profusion of wildly different approaches that he generated, some of which were merely discussed, others of which were implemented in overnight or weekend-long bursts of coding, in both cases ultimately discarded or further evolved when they turned out not to meet the design criteria well enough. Here are some of those approaches, presented in minimal detail in the hopes that, like Tom Wilson with the Paradise FIFO, your imagination will be sparked.

    -

    Subdividing Raycast

    +

    Subdividing Raycast

    Rays are cast in an 8x8 screen-pixel grid; this is a highly efficient operation because the first intersection with a surface can be found by simply clipping the ray into the BSP tree, starting at the viewpoint, until a solid leaf is reached. If adjacent rays don’t hit the same surface, then a ray is cast halfway between, and so on until all adjacent rays either hit the same surface or are on adjacent pixels; then the block around each ray is drawn from the polygon that was hit. This scales very well, being limited by the number of pixels, with no overdraw. The problem is dropouts; it’s quite possible for small polygons to fall between rays and vanish.

    -

    Vertex-Free Surfaces

    +

    Vertex-Free Surfaces

    The world is represented by a set of surface planes. The polygons are implicit in the plane intersections, and are extracted from the planes as a final step before drawing. This makes for fast clipping and a very small data set (planes are far more compact than polygons), but it’s time-consuming to extract polygons from planes.

    -

    The Draw-Buffer

    +

    The Draw-Buffer

    Like a z-buffer, but with 1 bit per pixel, indicating whether the pixel has been drawn yet. This eliminates overdraw, but at the cost of an inner-loop buffer test, extra writes and cache misses, and, worst of all, considerable complexity. Variations include testing the draw-buffer a byte at a time and completely skipping fully-occluded bytes, or branching off each draw-buffer byte to one of 256 unrolled inner loops for drawing 0-8 pixels, in the process possibly taking advantage of the ability of the x86 to do the perspective floating-point divide in parallel while 8 pixels are processed.

    -

    Span-Based Drawing

    +

    Span-Based Drawing

    Polygons are rasterized into spans, which are added to a global span list and clipped against that list so that only the nearest span at each pixel remains. Little sorting is needed with front-to-back walking, because if there’s any overlap, the span already in the list is nearer. This eliminates overdraw, but at the cost of a lot of span arithmetic; also, every polygon still has to be turned into spans.

    -

    Portals

    +

    Portals

    The holes where polygons are missing on surfaces are tracked, because it’s only through such portals that line-of-sight can extend. Drawing goes front-to-back, and when a portal is encountered, polygons and portals behind it are clipped to its limits, until no polygons or portals remain visible. Applied recursively, this allows drawing only the visible portions of visible polygons, but at the cost of a considerable amount of portal clipping.

    -

    Breakthrough!

    +

    Breakthrough!

    In the end, John decided that the beam tree was a sort of second-order structure, reflecting information already implicitly contained in the world BSP tree, so he tackled the problem of extracting visibility information directly from the world BSP tree. He spent a week on this, as a byproduct devising a perfect DOOM (2-D) visibility architecture, whereby a single, linear walk of a DOOM BSP tree produces zero-overdraw 2-D visibility. Doing the same in 3-D turned out to be a much more complex problem, though, and by the end of the week John was frustrated by the increasing complexity and persistent glitches in the visibility code. Although the direct-BSP approach was getting closer to working, it was taking more and more tweaking, and a simple, clean design didn’t seem to be falling out. When I left work one Friday, John was preparing to try to get the direct-BSP approach working properly over the weekend.

    @@ -100,10 +91,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/64-04.html b/64-04.html index 898e614..3488b36 100644 --- a/64-04.html +++ b/64-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Visible-Surface Determination - - @@ -45,7 +38,7 @@

    John says precalculating the PVS was a logical evolution of the approaches he had been considering, that there was no moment when he said “Eureka!” Nonetheless, it was clearly a breakthrough to a brand-new, superior design, a design that, together with a still-in-development sorted-edge rasterizer that completely eliminates overdraw, comes remarkably close to meeting the “perfect-world” specifications we laid out at the start.

    -

    Simplify, and Keep on Trying New Things

    +

    Simplify, and Keep on Trying New Things

    What does it all mean? Exactly what I said up front: Simplify, and keep trying new things. The precalculated PVS is simpler than any of the other schemes that had been considered (although precalculating the PVS is an interesting task that I’ll discuss another time). In fact, at runtime the precalculated PVS is just a constrained version of the painter’s algorithm. Does that mean it’s not particularly profound?

    @@ -65,7 +58,7 @@

    So far, it seems to have worked out pretty well for him.

    -

    Learn Now, Pay Forward

    +

    Learn Now, Pay Forward

    There’s one other thing I’d like to mention before I close this chapter. Much of what I’ve learned, and a great deal of what I’ve written, has been in the pages of Dr. Dobb’s Journal. As far back as I can remember, DDJ has epitomized the attitude that sharing programming information is A Good Thing. I know a lot of programmers who were able to leap ahead in their development because of Hendrix’s Tiny C, or Stevens’ D-Flat, or simply by browsing through DDJ’s annual collections. (Me, for one.) Understandably, most companies understandably view sharing information in a very different way, as potential profit lost—but that’s what makes DDJ so valuable to the programming community.

    @@ -73,7 +66,7 @@

    So remember, when it’s legally possible, sharing information benefits us all in the long run. You can pay forward the debt for the information you gain here and elsewhere by sharing what you know whenever you can, by writing an article or book or posting on the Net. None of us learns in a vacuum; we all stand on the shoulders of giants such as Wirth and Knuth and thousands of others. Lend your shoulders to building the future!

    -

    References

    +

    References

    Foley, James D., et al., Computer Graphics: Principles and Practice, Addison Wesley, 1990, ISBN 0-201-12110-7 (beams, BSP trees, VSD).

    @@ -98,10 +91,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/65-01.html b/65-01.html index 1220f78..9e3b84a 100644 --- a/65-01.html +++ b/65-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - @@ -37,10 +30,10 @@


    -

    Chapter 65
    +

    Chapter 65
    3-D Clipping and Other Thoughts

    -

    Determining What’s Inside Your Field of View

    +

    Determining What’s Inside Your Field of View

    Our part of the world is changing, and I’m concerned. By way of explanation, three anecdotes.

    @@ -64,7 +57,7 @@

    Things aren’t changing everywhere, though; over the past year, I’ve circulated a good bit of info about 3-D graphics, and plan to keep on doing it as long as I can. Next, we’re going to take a look at 3-D clipping.

    -

    3-D Clipping Basics

    +

    3-D Clipping Basics

    Before I got deeply into 3-D, I kept hearing how difficult 3-D clipping was, so I was pleasantly surprised when I actually got around to doing it and found that it was quite straightforward, after all. At heart, 3-D clipping is nothing more than evaluating whether and where a line intersects a plane; in this context, the plane is considered to have an “inside” (a side on which points are to be kept) and an “outside” (a side on which points are to be removed or clipped). We can easily extend this single operation to polygon clipping, working with the line segments that form the edges of a polygon.

    @@ -89,10 +82,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/65-02.html b/65-02.html index c63970f..b59ac32 100644 --- a/65-02.html +++ b/65-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - @@ -37,13 +30,13 @@


    -

    Intersecting a Line Segment with a Plane

    +

    Intersecting a Line Segment with a Plane

    The fundamental 3-D clipping operation is clipping a line segment to a plane. There are two parts to this operation: determining if the line is clipped by (intersects) the plane at all and, if it is clipped, calculating the point of intersection.

    Before we can intersect a line segment with a plane, we must first define how we’ll represent the line segment and the plane. The segment will be represented in the obvious way by the (x,y,z) coordinates of its two endpoints; this extends well to polygons, where each vertex is an (x,y,z) point. Planes can be described in many ways, among them are three points on the plane, a point on the plane and a unit normal, or a unit normal and a distance from the origin along the normal; we’ll use the latter definition. Further, we’ll define the normal to point to the inside (unclipped side) of the plane. The structures for points, polygons, and planes are shown in Listing 65.1.

    -

    LISTING 65.1 L65_1.h

    +

    LISTING 65.1 L65_1.h

     typedef struct {
         double v[3];
    @@ -77,7 +70,7 @@ typedef struct {
         double  distance;
         point_t normal;
     } plane_t;
    -
    +

    Given a line segment, and a plane to which to clip the segment, the first question is whether the segment is entirely on the inside or the outside of the plane, or intersects the plane. If the segment is on the inside, then the segment is not clipped by the plane, and we’re done. If it’s on the outside, then it’s entirely clipped, and we’re likewise done. If it intersects the plane, then we have to remove the clipped portion of the line by replacing the endpoint on the outside of the plane with the point of intersection between the line and the plane.

    @@ -89,19 +82,17 @@ typedef struct {

    From our earlier tests, we already know the length from the plane, measured along the normal, to the inside endpoint; that’s just the distance, along the normal, of the inside endpoint from the origin (the dot product of the endpoint with the normal), minus the plane distance, as shown in Figure 65.1. We also know the length of the line segment, again measured as projected onto the normal; that’s the difference between the distances along the normal of the inside and outside endpoints from the origin. The ratio of these two lengths is the fraction of the segment that remains after clipping. If we scale the x, y, and z lengths of the line segment by that fraction, and add the results to the inside endpoint, we get a new, clipped endpoint at the point of intersection.

    -

    Polygon Clipping

    +

    Polygon Clipping

    Line clipping is fine for wireframe rendering, but what we really want to do is polygon rendering of solid models, which requires polygon clipping. As with line segments, the clipping process with polygons is to determine if they’re inside, outside, or partially inside the clip volume, lopping off any vertices that are outside the clip volume and substituting vertices at the intersection between the polygon and the clip plane, as shown in Figure 65.2.

    An easy way to clip a polygon is to decompose it into a set of edges, and clip each edge separately as a line segment. Let’s define a polygon as a set of vertices that wind clockwise around the outside of the polygonal area, as viewed from the front side of the polygon; the edges are implicitly defined by the order of the vertices. Thus, an edge is the line segment described by the two adjacent vertices that form its endpoints. We’ll clip a polygon by clipping each edge individually, emitting vertices for the resulting polygon as appropriate, depending on the clipping state of the edge. If the start point of the edge is inside, that point is added to the output polygon. Then, if the start and end points are in different states (one inside and one outside), we clip the edge to the plane, as described above, and add the point at which the line intersects the clip plane as the next polygon vertex, as shown in Figure 65.3. Listing 65.2 shows a polygon-clipping function.

    -


    - Figure 65.1
      The distance from the plane to the inside endpoint, measured along the normal.

    +


    + Figure 65.1
      The distance from the plane to the inside endpoint, measured along the normal.

    -


    - Figure 65.2
      Clipping a polygon.

    +


    + Figure 65.2
      Clipping a polygon.


    @@ -120,10 +111,6 @@ typedef struct {
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/65-03.html b/65-03.html index f17ee0b..88427b7 100644 --- a/65-03.html +++ b/65-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - @@ -37,7 +30,7 @@


    -

    LISTING 65.2 L65_2.c

    +

    LISTING 65.2 L65_2.c

     int ClipToPlane(polygon_t *pin, plane_t *pplane, polygon_t *pout)
     {
    @@ -89,21 +82,20 @@ int ClipToPlane(polygon_t *pin, plane_t *pplane, polygon_t *pout)
         pout->color = pin->color;
         return 1;
     }
    -
    +

    Believe it or not, this technique, applied in turn to each edge, is all that’s needed to clip a polygon to a plane. Better yet, a polygon can be clipped to multiple planes by repeating the above process once for each clip plane, with each interation trimming away any part of the polygon that’s clipped by that particular plane.

    One particularly useful aspect of 3-D clipping is that if you’re drawing texture mapped polygons, texture coordinates can be clipped in exactly the same way as (x,y,z) coordinates. In fact, the very same fraction that’s used to advance x, y, and z from the inside point to the point of intersection with the clip plane can be used to advance the texture coordinates as well, so only one extra multiply and one extra add are required for each texture coordinate.

    -

    Clipping to the Frustum

    +

    Clipping to the Frustum

    Given a polygon-clipping function, it’s easy to clip to the frustum: set up the four planes for the sides of the frustum, with another one or two planes for near and far clipping, if desired; next, clip each potentially visible polygon to each plane in turn; then draw whatever polygons emerge from the clipping process. Listing 65.3 is the core code for a simple 3-D clipping example that allows you to move around and look at polygonal models from any angle. The full code for this program is available on the CD-ROM in the file DDJCLIP.ZIP.

    -


    - Figure 65.3
      Clipping a polygon edge.

    +


    + Figure 65.3
      Clipping a polygon edge.

    -

    LISTING 65.3 L65_3.c

    +

    LISTING 65.3 L65_3.c

     int DIBWidth, DIBHeight;
     int DIBPitch;
    @@ -411,7 +403,7 @@ void UpdateWorld()
         SelectObject(hdcDIBSection, holdbitmap);
         ReleaseDC(hwndOutput, hdcDIBSection);
     }
    -
    +


    @@ -430,10 +422,6 @@ void UpdateWorld()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/65-04.html b/65-04.html index 6d7d449..b64c231 100644 --- a/65-04.html +++ b/65-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: 3-D Clipping and Other Thoughts - - @@ -37,7 +30,7 @@


    -

    The Lessons of Listing 65.3

    +

    The Lessons of Listing 65.3

    There are several interesting points to Listing 65.3. First, floating-point arithmetic is used throughout the clipping process. While it is possible to use fixed-point, doing so requires considerable care regarding range and precision. Floating-point is much easier—and, with the Pentium generation of processors, is generally comparable in speed. In fact, for some operations, such as multiplication in general and division when the floating-point unit is in single-precision mode, floating-point is much faster. Check out Chris Hecker’s column in the February 1996 Game Developer for an interesting discussion along these lines.

    @@ -49,7 +42,7 @@

    Finally, clipping in Listing 65.3 is performed in worldspace, rather than in viewspace. The frustum is backtransformed from viewspace (where it is defined, since it exists relative to the viewer) to worldspace for this purpose. Worldspace clipping allows us to transform only those vertices that are visible, rather than transforming all vertices into viewspace, then clipping them. However, the decision whether to clip in worldspace or viewspace is not clear-cut and is affected by several factors.

    -

    Advantages of Viewspace Clipping

    +

    Advantages of Viewspace Clipping

    Although viewspace clipping requires transforming vertices that may not be drawn, it has potential performance advantages. For example, in worldspace, near and far clip planes are just additional planes that have to be tested and clipped to, using dot products. In viewspace, near and far clip planes are typically planes with constant z coordinates, so testing whether a vertex is near or far-clipped can be performed with a single z compare, and the fractional distance along a line segment to a near or far clip intersection can be calculated with a couple of z subtractions and a divide; no dot products are needed.

    @@ -59,7 +52,7 @@

    I didn’t implement normalized clipping in Listing 65.3 because I wanted to illustrate the general 3-D clipping mechanism without additional complications, and because for many applications the dot product (which, after all, takes only 10-20 cycles on a Pentium) is sufficient. However, the more frustum clipping you’re doing, especially if most of the polygons are trivially visible, the more attractive the performance advantages of normalized clipping become.

    -

    Further Reading

    +

    Further Reading

    You now have the basics of 3-D clipping, but because fast clipping is central to high-performance 3-D, there’s a lot more to be learned. One good place for further reading is Foley and van Dam; another is Procedural Elements of Computer Graphics, by David F. Rogers. Read and understand either of these books, and you’ll know everything you need for world-class clipping.

    @@ -82,10 +75,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/66-01.html b/66-01.html index c2cef78..9bb0242 100644 --- a/66-01.html +++ b/66-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - @@ -37,10 +30,10 @@


    -

    Chapter 66
    +

    Chapter 66
    Quake’s Hidden-Surface Removal

    -

    Struggling with Z-Order Solutions to the Hidden Surface Problem

    +

    Struggling with Z-Order Solutions to the Hidden Surface Problem

    Okay, I admit it: I’m sick and tired of classic rock. Admittedly, it’s been a while, about 20 years, since I was last excited to hear anything by the Cars or Boston, and I was never particularly excited in the first place about Bob Seger or Queen, to say nothing of Elvis, so some things haven’t changed. But I knew something was up when I found myself changing the station on the Allman Brothers and Steely Dan and Pink Floyd and, God help me, the Beatles (just stuff like “Hello Goodbye” and “I’ll Cry Instead,” though, not “Ticket to Ride” or “A Day in the Life”; I’m not that far gone). It didn’t take long to figure out what the problem was; I’d been hearing the same songs for a quarter-century, and I was bored.

    @@ -54,21 +47,21 @@

    Not that I should have needed any reminding, considering the ever-evolving nature of Quake.

    -

    Creative Flux and Hidden Surfaces

    +

    Creative Flux and Hidden Surfaces

    Back in Chapter 64, I described the creative flux that led to John Carmack’s decision to use a precalculated potentially visible set (PVS) of polygons for each possible viewpoint in Quake, the game we’re developing here at id Software. The precalculated PVS meant that instead of having to spend a lot of time searching through the world database to find out which polygons were visible from the current viewpoint, we could simply draw all the polygons in the PVS from back-to-front (getting the ordering courtesy of the world BSP tree) and get the correct scene drawn with no searching at all; letting the back-to-front drawing perform the final stage of hidden-surface removal (HSR). This was a terrific idea, but it was far from the end of the road for Quake’s design.

    -

    Drawing Moving Objects

    +

    Drawing Moving Objects

    For one thing, there was still the question of how to sort and draw moving objects properly; in fact, this is the single technical question I’ve been asked most often in recent months, so I’ll take a moment to address it here. The primary problem is that a moving model can span multiple BSP leaves, with the leaves that are touched varying as the model moves; that, together with the possibility of multiple models in one leaf, means there’s no easy way to use BSP order to draw the models in correctly sorted order. When I wrote Chapter 64, we were drawing sprites (such as explosions), moveable BSP models (such as doors), and polygon models (such as monsters) by clipping each into all the leaves it touched, then drawing the appropriate parts as each BSP leaf was reached in back-to-front traversal. However, this didn’t solve the issue of sorting multiple moving models in a single leaf against each other, and also left some ugly sorting problems with complex polygon models.

    John solved the sorting issue for sprites and polygon models in a startlingly low-tech way: We now z-buffer them. (That is, before we draw each pixel, we compare its distance, or z, value with the z value of the pixel currently on the screen, drawing only if the new pixel is nearer than the current one.) First, we draw the basic world, walls, ceilings, and the like. No z-buffer testing is involved at this point (the world visible surface determination is done in a different way, as we’ll see soon); however, we do fill the z-buffer with the z values (actually, 1/z values, as discussed below) for all the world pixels. Z-filling is a much faster process than z-buffering the entire world would be, because no reads or compares are involved, just writes of z values. Once the drawing and z-filling of the world is done, we can simply draw the sprites and polygon models with z-buffering and get perfect sorting all around.

    -

    Performance Impact

    +

    Performance Impact

    Whenever a z-buffer is involved, the questions inevitably are: What’s the memory footprint and what’s the performance impact? Well, the memory footprint at 320x200 is 128K, not trivial but not a big deal for a game that requires 8 MB to run. The performance impact is about 10 percent for z-filling the world, and roughly 20 percent (with lots of variation) for drawing sprites and polygon models. In return, we get a perfectly sorted world, and also the ability to do additional effects, such as particle explosions and smoke, because the z-buffer lets us flawlessly sort such effects into the world. All in all, the use of the z-buffer vastly improved the visual quality and flexibility of the Quake engine, and also simplified the code quite a bit, at an acceptable memory and performance cost.

    -

    Leveling and Improving Performance

    +

    Leveling and Improving Performance

    As I said above, in the Quake architecture, the world itself is drawn first, without z-buffer reads or compares, but filling the z-buffer with the world polygons’ z values, and then the moving objects are drawn atop the world, using full z-buffering. Thus far, I’ve discussed how to draw moving objects. For the rest of this chapter, I’m going to talk about the other part of the drawing equation; that is, how to draw the world itself, where the entire world is stored as a single BSP tree and never moves.

    @@ -89,10 +82,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/66-02.html b/66-02.html index 547e6c5..5ae779f 100644 --- a/66-02.html +++ b/66-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - @@ -51,15 +44,14 @@

    And indeed there is.

    -

    Sorted Spans

    +

    Sorted Spans

    The ideal final HSR stage for Quake would reject all the polygons in the PVS that are actually invisible, and draw only the visible pixels of the remaining polygons, with no overdraw, that is, with every pixel drawn exactly once, all at no performance cost, of course. One way to do that (although certainly not at zero cost) would be to draw the polygons from front-to-back, maintaining a region describing the currently occluded portions of the screen and clipping each polygon to that region before drawing it. That sounds promising, but it is in fact nothing more or less than the beam tree approach I described in Chapter 64, an approach that we found to have considerable overhead and serious leveling problems.

    We can do much better if we move the final HSR stage from the polygon level to the span level and use a sorted-spans approach. In essence, this approach consists of turning each polygon into a set of spans, as shown in Figure 66.1, and then sorting and clipping the spans against each other until only the visible portions of visible spans are left to be drawn, as shown in Figure 66.2. This may sound a lot like z-buffering (which is simply too slow for use in drawing the world, although it’s fine for smaller moving objects, as described earlier), but there are crucial differences.

    -


    - Figure 66.1
      Span generation.

    +


    + Figure 66.1
      Span generation.

    By contrast with z-buffering, only visible portions of visible spans are scanned out pixel by pixel (although all polygon edges must still be rasterized). Better yet, the sorting that z-buffering does at each pixel becomes a per-span operation with sorted spans, and because of the coherence implicit in a span list, each edge is sorted only against some of the spans on the same line and is clipped only to the few spans that it overlaps horizontally. Although complex scenes still take longer to process than simple scenes, the worst case isn’t as bad as with the beam tree or back-to-front approaches, because there’s no overdraw or scanning of hidden pixels, because complexity is limited to pixel resolution and because span coherence tends to limit the worst-case sorting in any one area of the screen. As a bonus, the output of sorted spans is in precisely the form that a low-level rasterizer needs, a set of span descriptors, each consisting of a start coordinate and a length.

    @@ -67,15 +59,14 @@

    So we’ve found the approach we need; now it’s just a matter of writing some code and we’re on our way, right? Well, yes and no. Conceptually, the sorted-spans approach is simple, but it’s surprisingly difficult to implement, with a couple of major design choices to be made, a subtle mathematical element, and some tricky gotchas that I’ll have to defer until Chapter 67. Let’s look at the design choices first.

    -

    Edges versus Spans

    +

    Edges versus Spans

    The first design choice is whether to sort spans or edges (both of which fall into the general category of “sorted spans”). Although the results are the same both ways, a list of spans to be drawn, with no overdraw, the implementations and performance implications are quite different, because the sorting and clipping are performed using very different data structures.

    With span-sorting, spans are stored in x-sorted, linked list buckets, typically with one bucket per scan line. Each polygon in turn is rasterized into spans, as shown in Figure 66.1, and each span is sorted and clipped into the bucket for the scan line the span is on, as shown in Figure 66.2, so that at any time each bucket contains the nearest spans encountered thus far, always with no overlap. This approach involves generating all spans for each polygon in turn, with each span immediately being sorted, clipped, and added to the appropriate bucket.

    -


    - Figure 66.2
      Two sets of spans sorted and clipped against one another.

    +


    + Figure 66.2
      Two sets of spans sorted and clipped against one another.

    With edge-sorting, edges are stored in x-sorted, linked list buckets according to their start scan line. Each polygon in turn is decomposed into edges, cumulatively building a list of all the edges in the scene. Once all edges for all polygons in the view frustum have been added to the edge list, the whole list is scanned out in a single top-to-bottom, left-to-right pass. An active edge list (AEL) is maintained. With each step to a new scan line, edges that end on that scan line are removed from the AEL, active edges are stepped to their new x coordinates, edges starting on the new scan line are added to the AEL, and the edges are sorted by current x coordinate.

    @@ -96,10 +87,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/66-03.html b/66-03.html index 4fbd681..2325293 100644 --- a/66-03.html +++ b/66-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - @@ -45,19 +38,17 @@

    Both span-sorting and edge-sorting work well, and both have been employed successfully in commercial projects. We’ve chosen to use edge-sorting in Quake partly because it seems inherently more efficient, with excellent horizontal coherence that makes for minimal time spent sorting, in contrast with the potentially costly sorting into linked lists that span-sorting can involve. A more important reason, though, is that with edge-sorting we’re able to share edges between adjacent polygons, and that cuts the work involved in sorting, clipping, and rasterizing edges nearly in half, while also shrinking the world database quite a bit due to the sharing.

    -


    - Figure 66.3
      Activating a polygon when a leading edge is encountered in the AEL.

    +


    + Figure 66.3
      Activating a polygon when a leading edge is encountered in the AEL.

    One final advantage of edge-sorting is that it makes no distinction between convex and concave polygons. That’s not an important consideration for most graphics engines, but in Quake, edge clipping, transformation, projection, and sorting have become a major bottleneck, so we’re doing everything we can to get the polygon and edge counts down, and concave polygons help a lot in that regard. While it’s possible to handle concave polygons with span-sorting, that can involve significant performance penalties.

    -


    - Figure 66.4
      Deactivating a polygon when a trailing edge is encountered in the AEL.

    +


    + Figure 66.4
      Deactivating a polygon when a trailing edge is encountered in the AEL.

    Nonetheless, there’s no cut-and-dried answer as to which approach is better. In the end, span-sorting and edge-sorting amount to the same functionality, and the choice between them is a matter of whatever you feel most comfortable with. In Chapter 67, I’ll go into considerable detail about edge-sorting, complete with a full implementation. I’m going the spend the rest of this chapter laying the foundation for Chapter 67 by discussing sorting keys and 1/z calculation. In the process, I’m going to have to make a few forward references to aspects of edge-sorting that I haven’t yet covered in detail; my apologies, but it’s unavoidable, and all should become clear by the end of Chapter 67.

    -

    Edge-Sorting Keys

    +

    Edge-Sorting Keys

    Now that we know we’re going to sort edges, using them to emit spans for the polygons nearest the viewer, the question becomes: How can we tell which polygons are nearest? Ideally, we’d just store a sorting key in each polygon, and whenever a new edge came along, we’d compare its surface’s key to the keys of other currently active polygons, and could easily tell which polygon was nearest.

    @@ -86,10 +77,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/66-04.html b/66-04.html index fc1d6dc..3b71764 100644 --- a/66-04.html +++ b/66-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Hidden-Surface Removal - - @@ -47,7 +40,7 @@

    The full 1/z calculation requires two multiplies and two adds, all of which should be floating-point to avoid range errors. That much floating-point math sounds expensive but really isn’t, especially on a Pentium, where a plane’s 1/z value at any point can be calculated in as little as six cycles in assembly language.

    -

    Where That 1/Z Equation Comes From

    +

    Where That 1/Z Equation Comes From

    For those who are interested, here’s a quick derivation of the 1/z equation. The plane equation for a plane is

    @@ -63,7 +56,7 @@

    We’ll see 1/z sorting in action in Chapter 67.

    -

    Quake and Z-Sorting

    +

    Quake and Z-Sorting

    I mentioned earlier that Quake no longer uses BSP order as the sorting key; in fact, it uses 1/z as the key now. Elegant as the gradients are, calculating 1/z from them is clearly slower than just doing a compare on a BSP-ordered key, so why have we switched Quake to 1/z?

    @@ -71,7 +64,7 @@

    Another advantage of 1/z sorting is that it solves the sorting issues I mentioned at the start involving moving models that are themselves small BSP trees. Sorting in world BSP order wouldn’t work here, because these models are separate BSPs, and there’s no easy way to work them into the world BSP’s sequence order. We don’t want to use z-buffering for these models because they’re often large objects such as doors, and we don’t want to lose the overdraw-reduction benefits that closed doors provide when drawn through the edge list. With sorted spans, the edges of moving BSP models are simply placed in the edge list (first clipping polygons so they don’t cross any solid world surfaces, to avoid complications associated with interpenetration), along with all the world edges, and 1/z sorting takes care of the rest.

    -

    Decisions Deferred

    +

    Decisions Deferred

    There is, without a doubt, an awful lot of information in the preceding pages, and it may not all connect together yet in your mind. The code and accompanying explanation in the next chapter should help; if you want to peek ahead, the code is available on the CD-ROM as DDJZSORT.ZIP in the directory for Chapter 67. You may also want to take a look at Foley and van Dam’s Computer Graphics or Rogers’ Procedural Elements for Computer Graphics.

    @@ -94,10 +87,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/67-01.html b/67-01.html index e3acdcd..68cbb68 100644 --- a/67-01.html +++ b/67-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - @@ -37,10 +30,10 @@


    -

    Chapter 67
    +

    Chapter 67
    Sorted Spans in Action

    -

    Implementing Independent Span Sorting for Rendering without Overdraw

    +

    Implementing Independent Span Sorting for Rendering without Overdraw

    In Chapter 66, we dove headlong into the intricacies of hidden surface removal by way of z-sorted (actually, 1/z-sorted) spans. At the end of that chapter, I noted that we were currently using 1/z-sorted spans in Quake, but it was unclear whether we’d switch back to BSP order. Well, some time after that writing, it’s become clear: We’re back to sorting spans by BSP order.

    @@ -48,7 +41,7 @@

    1/z-sorted spans in Quake turned out pretty much the same way, as we’ll see in a moment. First, though, I’d like to note up front that this chapter is very technical and builds heavily on material I covered earlier in this section of the book; if you haven’t already read Chapters 59 through 66, you really should. Make no mistake about it, this is commercial-quality stuff; in fact, the code in this chapter uses the same sorting technique as the test version of Quake, QTEST1.ZIP, that id Software placed on the Internet in early March 1996. This material is the Real McCoy, true reports from the leading edge, and I trust that you’ll be patient if careful rereading and some occasional catch-up reading of earlier chapters are required to absorb everything contained herein. Besides, the ultimate reference for any design is working code, which you’ll find, in part, in Listing 67.1, and in its entirety in the file DDJZSORT.ZIP on the CD-ROM.

    -

    Quake and Sorted Spans

    +

    Quake and Sorted Spans

    As you’ll recall from Chapter 66, Quake uses sorted spans to get zero overdraw while rendering the world, thereby both improving overall performance and leveling frame rates by speeding up scenes that would otherwise experience heavy overdraw. Our original design used spans sorted by BSP order; because we traverse the world BSP tree from front-to-back relative to the viewpoint, the order in which BSP nodes are visited is a guaranteed front-to-back sorting order. We simply gave each node an increasing BSP sequence number as it was visited, set each polygon’s sort key to the BSP sequence number of the node (BSP splitting plane) it lay on, and used those sort keys when generating spans.

    @@ -81,10 +74,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/67-02.html b/67-02.html index cc8f49b..3fbccfa 100644 --- a/67-02.html +++ b/67-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - @@ -45,31 +38,29 @@

    For the remainder of this chapter, I’m going to look at the three main types of 1/z span sorting, then discuss a sample 3-D app built around 1/z span sorting.

    -

    Types of 1/z Span Sorting

    +

    Types of 1/z Span Sorting

    As a quick refresher: With 1/z span sorting, all the polygons in a scene are treated as sets of screenspace pixel spans, and 1/z (where z is distance from the viewpoint in viewspace, as measured along the viewplane normal) is used to sort the spans so that the nearest span overlapping each pixel is drawn. As I discussed in Chapter 66, in the sample program we’re actually going to do all our sorting with polygon edges, which represent spans in an implicit form.

    There are three types of 1/z span sorting, each requiring a different implementation. In order of increasing speed and decreasing complexity, they are: intersecting, abutting, and independent. (These are names of my own devising; I haven’t come across any standard nomenclature in the literature.)

    -

    Intersecting Span Sorting

    +

    Intersecting Span Sorting

    Intersecting span sorting occurs when polygons can interpenetrate. Thus, two spans may cross such that part of each span is visible, in which case the spans have to be split and drawn appropriately, as shown in Figure 67.1.

    -


    - Figure 67.1
      Intersecting span sorting.

    +


    + Figure 67.1
      Intersecting span sorting.

    Intersecting is the slowest and most complicated type of span sorting, because it is necessary to compare 1/z values at two points in order to detect interpenetration, and additional work must be done to split the spans as necessary. Thus, although intersecting span sorting certainly works, it’s not the first choice for performance.

    -

    Abutting Span Sorting

    +

    Abutting Span Sorting

    Abutting span sorting occurs when polygons that are not part of a continuous surface can butt up against one another, but don’t interpenetrate, as shown in Figure 67.2. This is the sorting used in Quake, where objects like doors often abut walls and floors, and turns out to be more complicated than you might think. The problem is that when an abutting polygon starts on a given scan line, as with polygon B in Figure 67.2, it starts at exactly the same 1/z value as the polygon it abuts, in this case, polygon A, so additional sorting is needed when these ties happen. Of course, the two-point sorting used for intersecting polygons would work, but we’d like to find something faster.

    As it turns out, the additional sorting for abutting polygons is actually quite simple; whichever polygon has a greater 1/z gradient with respect to screen x (that is, whichever polygon is heading fastest toward the viewer along the scan line) is the front one. The hard part is identifying when ties—that is, abutting polygons—occur; due to floating-point imprecision, as well as fixed-point edge-stepping imprecision that can move an edge slightly on the screen, calculations of 1/z from the combination of screen coordinates and 1/z gradients (as discussed last time) can be slightly off, so most tie cases will show up as near matches, not exact matches. This imprecision makes it necessary to perform two comparisons, one with an adjust-up by a small epsilon and one with an adjust-down, creating a range in which near-matches are considered matches. Fine-tuning this epsilon to catch all ties, without falsely reporting close-but-not-abutting edges as ties, proved to be troublesome in Quake, and the epsilon calculations and extra comparisons slowed things down.

    -


    - Figure 67.2
      Abutting span sorting.

    +


    + Figure 67.2
      Abutting span sorting.

    I do think that abutting 1/z span sorting could have been made reliable enough for production use in Quake, were it not that we share edges between adjacent polygons in Quake, so that the world is a large polygon mesh. When a polygon ends and is followed by an adjacent polygon that shares the edge that just ended, we simply assume that the adjacent polygon sorts relative to other active polygons in the same place as the one that ended (because the mesh is continuous and there’s no interpenetration), rather than doing a 1/z sort from scratch. This speeds things up by saving a lot of sorting, but it means that if there is a sorting error, a whole string of adjacent polygons can be sorted incorrectly, pulled in by the one missorted polygon. Missorting is a very real hazard when a polygon is very nearly perpendicular to the screen, so that the 1/z calculations push the limits of numeric precision, especially in single-precision floating point.

    @@ -92,10 +83,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/67-03.html b/67-03.html index af4659a..3cddb11 100644 --- a/67-03.html +++ b/67-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - @@ -37,17 +30,17 @@


    -

    Independent Span Sorting

    +

    Independent Span Sorting

    Finally, we come to independent span sorting, the simplest and fastest of the three, and the type the sample code in Listing 67.1 uses. Here, polygons never intersect or touch any other polygons except adjacent polygons with which they form a continuous mesh. This means that when a polygon starts on a scan line, a single 1/z comparison between that polygon and the polygons it overlaps on the screen is guaranteed to produce correct sorting, with no extra calculations or tricky cases to worry about.

    Independent span sorting is ideal for scenes with lots of moving objects that never actually touch each other, such as a space battle. Next, we’ll look at an implementation of independent 1/z span sorting.

    -

    1/z Span Sorting in Action

    +

    1/z Span Sorting in Action

    Listing 67.1 is a portion of a program that demonstrates independent 1/z span sorting. This program is based on the sample 3-D clipping program from Chapter 65; however, the earlier program did hidden surface removal (HSR) by simply z-sorting whole objects and drawing them back-to-front, while Listing 67.1 draws all polygons by way of a 1/z-sorted edge list. Consequently, where the earlier program worked only so long as object centers correctly described sorting order, Listing 67.1 works properly for all combinations of non-intersecting and non-abutting polygons. In particular, Listing 67.1 correctly handles concave polyhedra; a new L-shaped object (the data for which is not included in Listing 67.1) has been added to the sample program to illustrate this capability. The ability to handle complex shapes makes Listing 67.1 vastly more useful for real-world applications than the 3-D clipping demo from Chapter 65.

    -

    Listing 67.1 L67_1.C

    +

    Listing 67.1 L67_1.C

     // Part of Win32 program to demonstrate z-sorted spans. Whitespace
     // removed for space reasons. Full source code, with whitespace,
    @@ -481,7 +474,7 @@ void UpdateWorld()
         SelectObject(hdcDIBSection, holdbitmap);
         DeleteDC(hdcDIBSection);
     }
    -
    +


    @@ -500,10 +493,6 @@ void UpdateWorld()
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/67-04.html b/67-04.html index 2320e28..6f8b992 100644 --- a/67-04.html +++ b/67-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - @@ -68,10 +61,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/67-05.html b/67-05.html index 7f88a54..de48c91 100644 --- a/67-05.html +++ b/67-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Sorted Spans in Action - - @@ -37,7 +30,7 @@


    -

    Implementation Notes

    +

    Implementation Notes

    Finally, a few notes on Listing 67.1. First, you’ll notice that although we clip all polygons to the view frustum in worldspace, we nonetheless later clamp them to valid screen coordinates before adding them to the edge list. This catches any cases where arithmetic imprecision results in clipped polygon vertices that are a bit outside the frustum. I’ve only found such imprecision to be significant at very small z distances, so clamping would probably be unnecessary if there were a near clip plane, and might not even be needed in Listing 67.1, because of the slight nudge inward that we give the frustum planes, as described in Chapter 65. However, my experience has consistently been that relying on worldspace or viewspace clipping to produce valid screen coordinates 100 percent of the time leads to sporadic and hard-to-debug errors.

    @@ -66,10 +59,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/68-01.html b/68-01.html index 73f6425..2f9a2a2 100644 --- a/68-01.html +++ b/68-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - @@ -37,10 +30,10 @@


    -

    Chapter 68
    +

    Chapter 68
    Quake’s Lighting Model

    -

    A Radically Different Approach to Lighting Polygons

    +

    A Radically Different Approach to Lighting Polygons

    It was during my senior year in college that I discovered computer games. Not Wizardry, or Choplifter, or Ultima, because none of those existed yet—the game that hooked me was the original Star Trek game, in which you navigated from one 8x8 quadrant to another in search of starbases, occasionally firing phasers or photon torpedoes. This was less exciting than it sounds; after each move, the current quadrant had to be reprinted from scratch, along with the current stats—and the output device was a 10 cps printball console. A typical game took over an hour, during which nothing particularly stimulating ever happened (Klingons appeared periodically, but they politely waited for your next move before attacking, and your photon torpedoes never missed, so the outcome was never in doubt), but none of that mattered; nothing could detract from the sheer thrill of being in a computer-simulated universe.

    @@ -50,13 +43,13 @@

    It’s important to keep it that way. I’ve seen far too many people start to treat programming like a job, forgetting the joy of doing it, and burn out. So keep an eye on how you feel about the programming you’re doing, and if it’s getting stale, it’s time to learn something new; there’s plenty of interesting programming of all sorts to be done. Follow your interests—and don’t forget to have fun!

    -

    The Lighting Conundrum

    +

    The Lighting Conundrum

    I spent about two years working with John Carmack on Quake’s 3-D graphics engine. John faced several fundamental design issues while architecting Quake. I’ve written in earlier chapters about some of those issues, including eliminating non-visible polygons quickly via a precalculated potentially visible set (PVS), and improving performance by inserting potentially visible polygons into a global edge list and scanning out only the nearest polygon at each pixel.

    In this chapter, I’m going to talk about another, equally crucial design issue: how we developed our lighting approach for the part of the Quake engine that draws the world itself, the static walls and floors and ceilings. Monsters and players are drawn using completely different rendering code, with speed the overriding factor. A primary goal for the world, on the other hand, was to be as precise as possible, getting everything right so that polygons, textures, and sophisticated lighting would be pegged in place, with no visible shifting or distortion under all viewing conditions, for maximum player immersion—all with good performance, of course. As I’ll discuss, the twin goals of performance and rock-solid, complex lighting proved to be difficult to achieve with traditional lighting approaches; ultimately, a dramatically different approach was required.

    -

    Gouraud Shading

    +

    Gouraud Shading

    The traditional way to do realistic lighting in polygon pipelines is Gouraud shading (also known as smooth shading). Gouraud shading involves generating a lighting value at each polygon vertex by applying all relevant world lighting, linearly interpolating between lighting values down the edges of the polygon, and then linearly interpolating between the edges of the polygon across each span. If texture mapping is desired (and all polygons are texture mapped in Quake), then at each pixel in each span, the pixel’s corresponding texture map location (texel) is determined, and the interpolated lighting is applied to the texel to generate a final, lit pixel. Texels are generally taken from a 32x32 or 64x64 texture that’s tiled repeatedly across the polygon, for several reasons: performance (a 64x64 texture sits nicely in the 486 or Pentium cache), database size, and less artwork.

    @@ -64,7 +57,7 @@

    Gouraud shading allows for decent lighting effects with a relatively small amount of calculation and a compact data set that’s a simple extension of the basic polygon model. However, there are several important drawbacks to Gouraud shading, as well.

    -

    Problems with Gouraud Shading

    +

    Problems with Gouraud Shading

    The quality of Gouraud shading depends heavily on the average size of the polygons being drawn. Linear interpolation is used, so highlights can only occur at vertices, and color gradients are monotonic across the face of each polygon. This can make for bland lighting effects if polygons are large, and makes it difficult to do spotlights and other detailed or dramatic lighting effects. After John brought the initial, primitive Quake engine up using Gouraud shading for lighting, the first thing he tried to improve lighting quality was adding a single vertex and creating new polygons wherever a spotlight was directly overhead a polygon, with the new vertex added directly underneath the light, as shown in Figure 68.1. This produced fairly attractive highlights, but simultaneously made evident several problems.

    @@ -85,10 +78,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/68-02.html b/68-02.html index 78701b4..0bc7cad 100644 --- a/68-02.html +++ b/68-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - @@ -41,11 +34,10 @@

    Similar problems occur with overlapping lights, and with shadows, where additional polygons are required in order to approximate lighting detail well. In particular, good shadow edges need small polygons, because otherwise the gradient between light and dark gets spread across too wide an area. Worse still, the rate of lighting change across a shadow edge can vary considerably as a function of the geometry the edge crosses; wider polygons stretch and diffuse the transition between light and shadow. A related problem is that lighting discontinuities can be very visible at t-junctions (although ultimately we had to add edges to eliminate t-junctions anyway, because otherwise dropouts can occur along polygon edges). These problems can be eased by adding extra edges, but that increases the rasterization load.

    -


    - Figure 68.1
      Adding an extra vertex directly beneath a light.

    +


    + Figure 68.1
      Adding an extra vertex directly beneath a light.

    -

    Perspective Correctness

    +

    Perspective Correctness

    Another problem is that Gouraud shading isn’t perspective-correct. With Gouraud shading, lighting varies linearly across the face of a polygon, in equal increments per pixel—but unless the polygon is parallel to the screen, the same sort of perspective correction is needed to step lighting across the polygon properly as is required for texture mapping. Lack of perspective correction is not as visibly wrong for lighting as it is for texture mapping, because smooth lighting gradients can tolerate considerably more warping than can the detailed bitmapped images used in texture mapping, but it nonetheless shows up in several ways.

    @@ -57,21 +49,20 @@

    The obvious solution to rotational variance is to use only triangles, but that brings with it a new set of problems. It takes twice as many triangles as quads to describe the same scene, increasing the size of the world database and requiring extra rasterization, at a performance cost. Triangles still don’t provide perspective lighting; their lighting is rotationally invariant, but it’s still wrong—just wrong in a more consistant way. Gouraud-shaded triangles still result in odd lighting patterns, and require lots of triangles to support shadowing and other lighting detail. Finally, triangles don’t solve clipping or viewing variance.

    -


    - Figure 68.2
      How Gouraud shading varies with polygon screen orientation.

    +


    + Figure 68.2
      How Gouraud shading varies with polygon screen orientation.

    Yet another problem is that while it may work well to add extra geometry so that spotlights and shadows show up well, that’s feasible only for static lighting. Dynamic lighting—light cast by sources that move—has to work with whatever geometry the world has to offer, because its needs are constantly changing.

    These issues led us to conclude that if we were going to use Gouraud shading, we would have to build Quake levels from many small triangles, with sufficiently finely detailed geometry so that complex lighting could be supported and the inaccuracies of Gouraud shading wouldn’t be too noticeable. Unfortunately, that line of thinking brought us back to the problem of a much larger world database and a much heavier rasterization load (all the worse because Gouraud shading requires an additional interpolant, slowing the inner rasterization loop), so that not only would the world still be less than totally solid, because of the limitations of Gouraud shading, but the engine would also be too slow to support the complex worlds we had hoped for in Quake.

    -

    The Quest for Alternative Lighting

    +

    The Quest for Alternative Lighting

    None of which is to say that Gouraud shading isn’t useful in general. Descent uses it to excellent effect, and in fact Quake uses Gouraud shading for moving entities, because these consist of small triangles and are always in motion, which helps hide the relatively small lighting errors. However, Gouraud shading didn’t seem capable of meeting our design goals for rendering quality and speed for drawing the world as a whole, so it was time to look for alternatives.

    There are many alternative lighting approaches, most of them higher-quality than Gouraud, starting with Phong shading, in which the surface normal is interpolated across the polygon’s surface, and going all the way up to ray-tracing lighting techniques in which full illumination calculations are performed for all direct and reflected paths from each light source for each pixel. What all these approaches have in common is that they’re slower than Gouraud shading, too slow for our purposes in Quake. For weeks, we kicked around and rejected various possibilities and continued working with Gouraud shading for lack of a better alternative—until the day John came into work and said, “You know, I have an idea....”

    -

    Decoupling Lighting from Rasterization

    +

    Decoupling Lighting from Rasterization

    John’s idea came to him while was looking at a wall that had been carved into several pieces because of a spotlight, with an ugly lighting glitch due to a t-junction. He thought to himself that if only there were some way to treat it as one surface, it would look better and draw faster—and then he realized that there was a way to do that.

    @@ -94,10 +85,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/68-03.html b/68-03.html index 38a73d8..434d8a9 100644 --- a/68-03.html +++ b/68-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - @@ -41,19 +34,18 @@

    So what does surface-based lighting buy us? First and foremost, it provides consistent, perspective-correct lighting, eliminating all rotational, viewing, and clipping variance, because lighting is done in surface space rather than in screen space. By lighting in surface space, we bind the lighting to the texels in an invariant way, and then the lighting gets a free ride through the perspective texture mapper and ends up perfectly matched to the texels. Surface-based lighting also supports good, although not perfect, detail for overlapping lights and shadows. The 16-texel grid has a resolution of two feet in the Quake frame of reference, and this relatively fine resolution, together with the filtering performed when the light map is built, is sufficient to support complex shadows with smoothly fading edges. Additionally, surface-based lighting eliminates lighting glitches at t-junctions, because lighting is unrelated to vertices. In short, surface-based lighting meets all of Quake’s visual quality goals, which leaves only one question: How does it perform?

    -

    Size and Speed

    +

    Size and Speed

    As it turns out, the raw speed of surface-based lighting is pretty good. Although an extra step is required to build the surface, moving lighting and tiling into a separate loop from texture mapping allows each of the two loops to be optimized very effectively, with almost all variables kept in registers. The surface-building inner loop is particularly efficient, because it consists of nothing more than interpolating intensity, combining it with a texel and using the result to look up a lit texel color, and storing the results with a dword write every four texels. In assembly language, we got this code down to 2.25 cycles per lit texel in Quake. Similarly, the texture-mapping inner loop, which overlaps an FDIV for floating-point perspective correction with integer pixel drawing in 16-pixel bursts, has been squeezed down to 7.5 cycles per pixel on a Pentium, so the combined inner loop times for building and drawing a surface is roughly in the neighborhood of 10 cycles per pixel. It’s certainly possible to write a Gouraud-shaded perspective-correct texture mapper that’s somewhat faster than 10 cycles, but 10 cycles/pixel is fast enough to do 40 frames/second at 640x400 on a Pentium/100, so the cycle counts of surface-based lighting are acceptable. It’s worth noting that it’s possible to write a one-pass texture mapper that does approximately perspective-correct lighting. However, I have yet to hear of or devise such an inner loop that isn’t complicated and full of special cases, which makes it hard to optimize; worse, this approach doesn’t work well with the procedural and post-processing techniques I’ll discuss shortly.

    -


    - Figure 68.3
      Tiling the texture and lighting the texels from the light map.

    +


    + Figure 68.3
      Tiling the texture and lighting the texels from the light map.

    Moreover, surface-based lighting tends to spend more of its time in inner loops, because polygons can have any number of sides and don’t need to be split into multiple smaller polygons for lighting purposes; this reduces the amount of transformation and projection that are required, and makes polygon spans longer. So the performance of surface-based lighting stacks up very well indeed—except for caching.

    I mentioned earlier that a 64x64 texture tile fits nicely in the processor cache. A typical surface doesn’t. Every texel in every surface is unique, so even at 320x200 resolution, something on the rough order of 64,000 texels must be read in order to draw a single scene. (The number actually varies quite a bit, as discussed below, but 64,000 is in the ballpark.) This means that on a Pentium, we’re guaranteed to miss the cache once every 32 texels, and the number can be considerably worse than that if the texture access patterns are such that we don’t use every texel in a given cache line before that data gets thrown out of the cache. Then, too, when a surface is built, the surface buffer won’t be in the cache, so the writes will be uncached writes that have to go to main memory, then get read back from main memory at texture mapping time, potentially slowing things further still. All this together makes the combination of surface building and unlit texture mapping a potential performance problem, but that never posed a problem during the development of Quake, thanks to surface caching.

    -

    Surface Caching

    +

    Surface Caching

    When he thought of surface-based lighting, John immediately realized that surface building would be relatively expensive. (In fact, he assumed it would be considerably more expensive than it actually turned out to be with full assembly-language optimization.) Consequently, his design included the concept of caching surfaces, so that if the same surface were visible in the next frame, it could be reused without having to be rebuilt.

    @@ -76,10 +68,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/68-04.html b/68-04.html index 7588b60..3a4a792 100644 --- a/68-04.html +++ b/68-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake's Lighting Model - - @@ -39,21 +32,20 @@

    The amount of memory required for surface caching looked forbidding at first. Surfaces are large relative to texture tiles, because every texel of every surface is unique. Also, a surface can contain many texels relative to the number of pixels actually drawn on the screen, because due to perspective foreshortening, distant polygons have only a few pixels relative to the surface size in texels. Surfaces associated with partly hidden polygons must be fully built, even though only part of the polygon is visible, and if polygons are drawn back to front with overdraw, some polygons won’t even be visible, but will still require surface building and caching. What all this meant was that the surface cache initially looked to be very large, on the order of several megabytes, even at 320x200—too much for a game intended to run on an 8 MB machine.

    -

    Mipmapping To The Rescue

    +

    Mipmapping To The Rescue

    Two factors combined to solve this problem. First, polygons are drawn through an edge list with no overdraw, as I discussed a few chapters back, so no surface is ever built unless at least part of it is visible. Second, surfaces are built at four mipmap levels, depending on distance, with each mipmap level having one-quarter as many texels as the preceding level, as shown in Figure 68.4.

    For those whose heads haven’t been basted in 3-D technology for the past several years, mipmapping is 3-D graphics jargon for a process that normalizes the number of texels in a surface to be approximately equal to the number of pixels, reducing calculation time for distant surfaces containing only a few pixels. The mipmap level for a given surface is selected to result in a texel:pixel ratio approximately between 1:1 and 1:2, so texels map roughly to pixels, and more distant surfaces are correspondingly smaller. As a result, the number of surface texels required to draw a scene at 320x200 is on the rough order of 64,000; the number is actually somewhat higher, because of portions of surfaces that are obscured and viewspace-tilted polygons, which have high texel-to-pixel ratios along one axis, but not a whole lot higher. Thanks to mipmapping and the edge list, 600K has proven to be plenty for the surface cache at 320x200, even in the most complex scenes, and at 640x480, a little more than 1 MB suffices.

    -


    - Figure 68.4
      How mipmapping reduces surface caching requirements.

    +


    + Figure 68.4
      How mipmapping reduces surface caching requirements.

    All mipmapped texture tiles are generated as a preprocessing step, and loaded from disk at runtime. One interesting point is that a key to making mipmapping look good turned out to be box-filtering down from one level to the next by averaging four adjacent pixels, then using error diffusion dithering to generate the mipmapped texels.

    Also, mipmapping is done on a per-surface basis; the mipmap level for a whole surface is selected based on the distance from the viewer of the nearest vertex. This led us to limit surface size to a maximum of 256x256. Otherwise, surfaces such as floors would extend for thousands of texels, all at the mipmap level of the nearest vertex, and would require huge amounts of surface cache space while displaying a great deal of aliasing in distant regions due to a high texel:pixel ratio.

    -

    Two Final Notes on Surface Caching

    +

    Two Final Notes on Surface Caching

    Dynamic lighting has a significant impact on the performance of surface caching, because whenever the lighting on a surface changes, the surface has to be rebuilt. In the worst case, where the lighting changes on every visible surface, the surface cache provides no benefit, and rendering runs at the combined speed of surface building and texture mapping. This worst-case slowdown is tolerable but certainly noticeable, so it’s best to design games that use surface caching so only some of the surfaces change lighting at any one time. If necessary, you could alternate surface relighting so that half of the surfaces change on even frames, and half on odd frames, but large-scale, constant relighting is not surface caching’s strongest suit.

    @@ -76,10 +68,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/69-01.html b/69-01.html index bff3da9..203811d 100644 --- a/69-01.html +++ b/69-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - @@ -37,10 +30,10 @@


    -

    Chapter 69
    +

    Chapter 69
    Surface Caching and Quake’s Triangle Models

    -

    Probing Hardware-Assisted Surfaces and Fast Model Animation Without Sprites

    +

    Probing Hardware-Assisted Surfaces and Fast Model Animation Without Sprites

    In the late ’70s, I spent a summer doing contract programming at a government-funded installation called the Northeast Solar Energy Center (NESEC). Those were heady times for solar energy, what with the oil shortages, and there was lots of money being thrown at places like NESEC, which was growing fast.

    @@ -56,7 +49,7 @@

    Of course, most of the time programmers really are rational creatures, and the more information we have, the better. In that spirit, let’s look at more of the stuff that makes Quake tick, starting with what I’ve recently learned about surface caching.

    -

    Surface Caching with Hardware Assistance

    +

    Surface Caching with Hardware Assistance

    In Chapter 68, I discussed in detail the surface caching technique that Quake uses to do detailed, high-quality lighting without lots of polygons. Since writing that chapter, I’ve gone further, and spent a considerable amount of time working on the port of Quake to Rendition’s Verite 3-D accelerator chip. So let me start off this chapter by discussing what I’ve learned about using surface caching in conjunction with hardware.

    @@ -85,10 +78,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/69-02.html b/69-02.html index 9f2c4ed..ca36aa4 100644 --- a/69-02.html +++ b/69-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - @@ -37,11 +30,11 @@


    -

    Letting the Graphics Card Build the Textures

    +

    Letting the Graphics Card Build the Textures

    One obvious solution is to have the accelerator card build the textures, rather than having the CPU build and then download them. This eliminates downloading completely, and lets the accelerator, which should be faster at such things, do the texel manipulation. Whether this is actually faster depends on whether the CPU or the accelerator is doing more of the work overall, but it eliminates download time, which is a big help. This approach retains the ability to composite other effects, such as splatters and dents, onto surfaces, but by the same token retains the high memory requirements and dynamic lighting performance impact of the surface cache. It also requires that the 3-D API and accelerator being used allow drawing into a texture, which is not universally true. Neither do all APIs or accelerators allow applications enough control over the texture heap so that an efficient surface cache can be implemented, a point that favors non-caching approaches. (A similar option that wasn’t open to us due to time limitations is downloading 8-bpp surfaces and having the accelerator expand them to 16-bpp surfaces as it stores them in texture memory. Better yet, some accelerators support 8-bpp palettized hardware textures that are expanded to 16-bpp on the fly during texturing.)

    -

    The Light Map as Alpha Texture

    +

    The Light Map as Alpha Texture

    Another appealing non-caching approach is doing unlit texture-mapping in one pass, then lighting from the light map as a second pass, using the light map as an alpha texture. In other words, the textured polygon is drawn first, with no lighting, then the light map is textured on top of the polygon, with the light map intensity used as an alpha value to determine how brightly to light each texel. The hardware’s texture-mapping circuitry is used for both passes, so the lighting comes out perspective-correct and consistent under all viewing conditions, just as with the surface cache. The lighting polygons don’t even have to match the texture polygons, so they can represent dynamically changing lighting.

    @@ -49,11 +42,11 @@

    The next graphics engine you’ll see from id Software will be oriented heavily toward hardware accelerators, and at this point it’s a tossup whether the engine will use surface caching, Gouraud shading, or two-pass lighting.

    -

    Drawing Triangle Models

    +

    Drawing Triangle Models

    Most of the last group of chapters in this book discuss how Quake works. If you look closely, though, you’ll see that almost all of the information is about drawing the world—the static walls, floors, ceilings, and such. There are several reasons for this, in particular that it’s hard to get a world renderer working well, and that the world is the base on which everything else is drawn. However, moving entities, such as monsters, are essential to a useful game engine. Traditionally, these have been done with sprites, but when we set out to build Quake, we knew that it was time to move on to polygon-based models. (In the case of Quake, the models are composed of triangles.) We didn’t know exactly how we were going to make the drawing of these models fast enough, though, and went through quite a bit of experimentation and learning in the process of doing so. For the rest of this chapter I’ll discuss some interesting aspects of our triangle-model architecture, and present code for one useful approach for the rapid drawing of triangle models.

    -

    Drawing Triangle Models Fast

    +

    Drawing Triangle Models Fast

    We would have liked one rendering model, and hence one graphics pipeline, for all drawing in Quake; this would have simplified the code and tools, and would have made it much easier to focus our optimization efforts. However, when we tried adding polygon models to Quake’s global edge table, edge processing slowed down unacceptably. This isn’t that surprising, because the edge table was designed to handle 200 to 300 large polygons, not the 2,000 to 3,000 tiny triangles that a dozen triangle models in a scene can add. Restructuring the edge list to use trees rather than linked lists would have helped with the larger data sets, but the basic problem is that the edge table requires a considerable amount of overhead per edge per scan line, and triangle models have too few pixels per edge to justify that overhead. Also, the much larger edge table generated by adding triangle models doesn’t fit well in the CPU cache.

    @@ -76,10 +69,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/69-03.html b/69-03.html index 021073f..595161a 100644 --- a/69-03.html +++ b/69-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - @@ -41,11 +34,10 @@

    Early on, we decided to allow lower drawing quality for triangle models than for the world, in the interests of speed. For example, the triangles in the models are small, and usually distant—and generally part of a quickly moving monster that’s trying its best to do you in—so the quality benefits of perspective texture mapping would add little value. Consequently, we chose to draw the triangles with affine texture mapping, avoiding the work required for perspective. Mind you, the models are perspective-correct at the vertices; it’s just the pixels between the vertices that suffer slight warping.

    -


    - Figure 69.1
      Quake’s triangle-model drawing pipeline.

    +


    + Figure 69.1
      Quake’s triangle-model drawing pipeline.

    -

    Trading Subpixel Precision for Speed

    +

    Trading Subpixel Precision for Speed

    Another sacrifice at the altar of performance was subpixel precision. Before each triangle is drawn, we snap its vertices to the nearest integer screen coordinates, rather than doing the extra calculations to handle fractional vertex coordinates. This causes some jumping of triangle edges, but again, is not a problem in normal gameplay, especially for the animation of figures in continuous motion.

    @@ -53,7 +45,7 @@

    Finally, we decided to Gouraud-shade the triangle models, because this makes them look considerably more 3-D. However, we can’t afford to calculate where all the relevant light sources for each model are in each frame, or even which is the primary light source. Instead, we select each model’s lighting level based on how brightly the floor point it was standing on is lit, and use that lighting level for both ambient lighting (so all parts of the model have some illumination) and Gouraud shading—but the lighting vector for Gouraud shading is a fixed vector, so the model is always lit from the same direction. Somewhat surprisingly, in practice this looks considerably better than pure ambient lighting.

    -

    An Idea that Didn’t Work

    +

    An Idea that Didn’t Work

    As we implemented triangle models, we tried several ideas that didn’t work out. One that’s notable because it seems so appealing is caching a model’s image from one frame and reusing it in the next frame as a sprite. Our thinking was that clipping, transforming, projecting, and drawing a several-hundred-triangle model was going to be a lot more expensive than drawing a sprite, too expensive to allow very many models to be visible at once. We wanted to be able to display at least a dozen simultaneous models, so the idea was that for all but the closest models, we’d draw into a sprite, then reuse that sprite at the model’s new locations for the next two or three frames, amortizing the 3-D drawing cost over several frames and boosting overall model-drawing performance. The rendering wouldn’t be exactly right when the sprite was reused, because the view of the model would change from frame to frame as the viewer and model moved, but it didn’t seem likely that that slight inaccuracy would be noticeable for any but the nearest and largest models.

    @@ -61,7 +53,7 @@

    The sprite architecture also introduced considerable code complexity, increased memory footprint because of the need to cache the sprites, and made it difficult to get hidden surfaces exactly right because sprites are unavoidably 2-D. The performance of drawing the sprites dropped sharply as models got closer, and that’s also where the sprites looked worse when they were reused, limiting sprites to use at a considerable distance. All these problems could have been worked out reasonably well if necessary, but the sprite architecture just had the feeling of being fundamentally not the right approach, so we tried thinking along different lines.

    -

    An Idea that Did Work

    +

    An Idea that Did Work

    John Carmack had the notion that it was just way too much effort per pixel to do all the work of scanning out the tiny triangles in distant models. After all, distant models are just indistinct blobs of pixels, suffering heavily from effects such as texture aliasing and pixel quantization, he reasoned, so it should work just as well if we could come up with another way of drawing blobs of approximately equal quality. The trick was to come up with such an alternative approach. We tossed around half-formed ideas like flood-filling the model’s image within its silhouette, or encoding the model as a set of deltas, picking a visible seed point, and working around the visible side of the model according to the deltas. The first approach that seemed practical enough to try was drawing the pixel at each vertex replicated to form a 2x2 box, with all the vertices together forming the approximate shape of the model. Sometimes this worked quite well, but there were gaps where the triangles were large, and the quality was very erratic. However, it did point the way to something that in the end did the trick.

    @@ -84,10 +76,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/69-04.html b/69-04.html index 2e48de0..2c89c61 100644 --- a/69-04.html +++ b/69-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Surface Catching and Quake's Triangle Models - - @@ -43,7 +36,7 @@

    How much does subdivision rasterization help performance? When John originally implemented it, it more than doubled triangle-model drawing speed, because the affine texture mapper was not yet optimized. However, I took it upon myself to see how fast I could make the mapper, so now affine texture mapping is only about 20 percent slower than subdivision rasterization. While 20 percent may not sound impressive, it includes clipping, transform, projection, and backface-culling time, so the rasterization difference alone is more than 50 percent. Besides, 20 percent overall means that we can have 12 monsters now where we could only have had 10 before, so we count subdivision rasterization as a clear success.

    -

    LISTING 69.1 L69-1.C

    +

    LISTING 69.1 L69-1.C

     // Quake’s recursive subdivision triangle rasterizer; draws all
     // pixels in a triangle other than the vertices by splitting an
    @@ -157,13 +150,12 @@ nodraw:
     D_PolysetRecursiveTriangle (lp3, lp1, new);
     D_PolysetRecursiveTriangle (lp3, new, lp2);
     }
    -
    + -


    - Figure 69.2
      One recursive subdivision triangle-drawing step.

    +


    + Figure 69.2
      One recursive subdivision triangle-drawing step.

    -

    More Ideas that Might Work

    +

    More Ideas that Might Work

    Useful as subdivision rasterization proved to be, we by no means think that we’ve maxed out triangle-model drawing, if only because we spent far less design and development time on subdivision than on the affine rasterizer, so it’s likely that there’s quite a bit more performance to be found for drawing small triangles. For example, it could be faster to precalculate drawing masks or even precompile drawing code for all possible small triangles (say, up to 4x4 or 5x5), and the memory footprint looks reasonable. (It’s worth noting that both precalculated drawing and subdivision rasterization are only possible because we snap to integer coordinates; none of this stuff works with fixed-point vertices.)

    @@ -188,10 +180,6 @@ D_PolysetRecursiveTriangle (lp3, new, lp2);
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-01.html b/70-01.html index b47f0a0..f9c31c2 100644 --- a/70-01.html +++ b/70-01.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -37,7 +30,7 @@


    -

    Chapter 70
    +

    Chapter 70
    Quake: A Post-Mortem and a Glimpse into the Future

    Why did not any of the children in the first group think of this faster method of going across the room? It is simple. They looked at what they were given to use for materials and, they are like all of us, they wanted to use everything. But they did not need everything. They could do better with less, in a different way.

    @@ -52,7 +45,7 @@

    Before I begin, I’d like to remind you that all of the Doom and Quake material I’m presenting in this book is presented in the spirit of sharing information to make our corner of the world a better place for everyone. I’d like to thank John Carmack, Quake’s architect and lead programmer, and id Software for allowing me to share this technology with you, and I encourage you to share your own insights by posting on the Internet and writing books and articles whenever you have the opportunity and the right to do so. (Of course, check with your employer first!) We’ve all benefited greatly from the shared wisdom of people like Knuth, Foley and van Dam, Jim Blinn, Jim Kajiya, and hundreds of others—are you ready to take a shot at making your own contribution to the future?

    -

    Preprocessing the World

    +

    Preprocessing the World

    For the most part, I’ll discuss Quake’s 3-D engine in this chapter, although I’ll touch on other areas of interest. For 3-D rendering purposes, Quake consists of two basic sorts of objects: the world, which is stored as a single BSP model and never changes shape or position; and potentially moving objects, called entities, which are drawn in several different ways. I’ll discuss each separately.

    @@ -79,10 +72,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-02.html b/70-02.html index 52a178b..e086422 100644 --- a/70-02.html +++ b/70-02.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -41,7 +34,7 @@

    After the BSP is built, the outer surfaces of the level, which no one can ever see (because levels are sealed spaces), are removed, so the interior of the level, containing all the empty space through which a player can move, is completely surrounded by a solid region. This eliminates a great many irrelevant polygons, and reduces the complexity of the next step, calculating the potentially visible set.

    -

    The Potentially Visible Set (PVS)

    +

    The Potentially Visible Set (PVS)

    After the BSP tree is built, the potentially visible set (PVS) for each leaf is calculated. The PVS for a leaf consists of all the leaves that can be seen from anywhere in that leaf, and is used to reduce to a near-minimum the polygons that have to be considered for drawing from a given viewpoint, as well as the entities that have to be updated over the network (for multiplayer games) and drawn. Calculating the PVS is expensive; Quake levels take 10 to 30 minutes to process on a four-processor Alpha, and even with speedup tweaks to the BSPer (the most effective of which was replacing many calls to malloc() with stack-based structures—beware of malloc() in performance-sensitive code), Quake 2 levels are taking up to an hour to process. (Note, however, that that includes BSPing, PVS calculations, and radiosity lighting, which I’ll discuss later.)

    @@ -72,10 +65,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-03.html b/70-03.html index 1fed53d..0af69e9 100644 --- a/70-03.html +++ b/70-03.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -37,7 +30,7 @@


    -

    Passages: The Last-Minute Change that Didn’t Happen

    +

    Passages: The Last-Minute Change that Didn’t Happen

    Earlier, I mentioned that we almost changed 3-D engines again in the last month of Quake’s development. Here’s what happened: One of the alternatives to the PVS is the use of portals, where the focus is on the places where polygons don’t exist along leaf faces, rather than the more usual focus on the polygons themselves. These “empty” places are themselves polygons, called portals, that describe all the places that visibility can pass from one leaf to another. Portals are used by the PVS generator to determine visibility, and are used in other 3-D engines as the primary mechanism for determining leaf or sector visibility. For example, portals can be projected to screenspace, then used as a 2-D clipping region to restrict drawing of more distant polygons to only those that are visible through the portal. Or, as in Quake’s preprocessor, visibility boundary planes can be constructed from one portal to the next, and 3-D clipping to those planes can be used to determine visible polygons or leaves. Used either way, portals can support more changeable worlds than the PVS, because, unlike the PVS, the portals themselves can easily be changed on the fly.

    @@ -47,7 +40,7 @@

    The more approaches you try, the larger your toolkit and the broader your understanding will be when you tackle your next project.

    -

    Drawing the World

    +

    Drawing the World

    Everything described so far is a preprocessing step. When Quake is actually running, the world is drawn as follows: First, the PVS for the view leaf is decompressed, and each leaf flagged as visible is marked as being in the current frame’s PVS. (The marking is done by storing the current frame’s number in the leaf; this avoids having to clear the PVS marking each frame.) All the parent nodes of each leaf in the PVS are also marked; this information could have been stored as additional PVS flags, but to save space is bubbled up the BSP from each visible leaf.

    @@ -82,10 +75,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-04.html b/70-04.html index ae090d8..d9595b4 100644 --- a/70-04.html +++ b/70-04.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -39,17 +32,17 @@

    The edge list is an atypical technology for John; it’s an extra stage in the engine, it’s complex, and it doesn’t scale well. A Quake level might have a maximum of 500 potentially drawable polygons that get placed into the edge list, and that runs fine, but if you were to try to put 5,000 polygons into the edge list, it would quickly bog down due to edge sorting, link following, and dataset size. Different data structures (like using a tree to store the edges rather than a linear linked list) would help to some degree, but basically the edge list has a relatively small window of applicability; it was appropriate technology for the degree of complexity possible in a Pentium-based game (and even then, only with the reduction in polygons made possible by the PVS), but will probably be poorly suited to more complex scenes. It served well in the Quake engine, but remains an inelegant solution, and, in the end, it feels like there’s something better we didn’t hit on. However, as John says, “I’m pragmatic above all else”—and the edge list did the job.

    -

    Rasterization

    +

    Rasterization

    Once the visible spans are scanned out of the edge list, they must still be drawn, with perspective-correct texture mapping and lighting. This involves hundreds of lines of heavily optimized assembly language, but is fundamentally pretty simple. In order to draw the spans for a given surface, the screenspace equations for 1/z, s/z, and t/z (where s and t are the texture coordinates and z is distance) are calculated for the surface. Then for each span, these values are calculated for the points at each end of the span, the reciprocal of 1/z is calculated with a divide, and s and t are then calculated as (s/z)*z and (t/z)*z. If the span is longer than 16 pixels, s and t are likewise calculated every 16 pixels along the span. Then each stretch of up to 16 pixels is drawn by linearly interpolating between these correctly calculated points. This introduces some slight error, but this is almost never visible, and even then is only a small ripple, well worth the performance improvement gained by doing the perspective-correct math only once every 16 pixels. To speed things up a little more, the FDIV to calculate the reciprocal of 1/z is overlapped with drawing 16 pixels, taking advantage of the Pentium’s ability to perform floating-point in parallel with integer instructions, so the FDIV effectively takes only one cycle.

    -

    Lighting

    +

    Lighting

    Lighting is less simple to explain. The traditional way of doing polygon lighting is to calculate the correct light at the vertices and linearly interpolate between those points (Gouraud shading), but this has several disadvantages; in particular, it makes it hard to get detailed lighting without creating a lot of extra polygons, the lighting isn’t perspective correct, and the lighting varies with viewing angle for polygons other than triangles. To address these problems, Quake uses surface-based lighting instead. In this approach, when it’s time to draw a surface (a world polygon), that polygon’s texture is tiled into a memory buffer. At the same time, the texture is lit according to the surface’s light map, as calculated during preprocessing. Lighting values are linearly interpolated between the light map’s 16-texel grid points, so the lighting effects are smooth, but slightly blurry. Then, the polygon is drawn to the screen using the perspective-correct texture mapping described above, with the prelit surface buffer being the source texture, rather than the original texture tile. No additional lighting is performed during texture mapping; all lighting is done when the surface buffer is created.

    Certainly it takes longer to build a surface buffer and then texture map from it than it does to do lighting and texture mapping in a single pass. However, surface buffers are cached for reuse, so only the texture mapping stage is usually needed. Quake surfaces tend to be big, so texture mapping is slowed by cache misses; however, the Quake approach doesn’t need to interpolate lighting on a pixel-by-pixel basis, which helps speed things up, and it doesn’t require additional polygons to provide sophisticated lighting. On balance, the performance of surface-based drawing is roughly comparable to tiled, Gouraud-shaded texture mapping—and it looks much better, being perspective correct, rotationally invariant, and highly detailed. Surface-based drawing also has the potential to support some interesting effects, because anything that can be drawn into the surface buffer can be cached as well, and is automatically drawn in correct perspective. For instance, paint splattered on a wall could be handled by drawing the splatter image as a sprite into the appropriate surface buffer, so that drawing the surface would draw the splatter as well.

    -

    Dynamic Lighting

    +

    Dynamic Lighting

    Here we come to a feature added to Quake after last year’s Computer Game Developer’s Conference (CGDC). At that time, Quake did not support dynamic lighting; that is, explosions and such didn’t produce temporary lighting effects. We hadn’t thought dynamic lighting would add enough to the game to be worth the trouble; however, at CGDC Billy Zelsnack showed us a demo of his latest 3-D engine, which was far from finished at the time, but did have impressive dynamic lighting effects. This caused us to move dynamic lighting up the priority list, and when I got back to id, I spent several days making the surface-building code as fast as possible (winding up at 2.25 cycles per texel in the inner loop) in anticipation of adding dynamic lighting, which would of course cause dynamically lit surfaces to constantly be rebuilt as the lighting changed. (A significant drawback of dynamic lighting is that it makes surface caching worthless for dynamically lit surfaces, but if most of the surfaces in a scene are not dynamically lit at any one time, it works out fine.) There things stayed for several weeks, while more critical work was done, and it was uncertain whether dynamic lighting would, in fact, make it into Quake.

    @@ -76,10 +69,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-05.html b/70-05.html index d735c37..1ef11e0 100644 --- a/70-05.html +++ b/70-05.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -37,11 +30,11 @@


    -

    Entities

    +

    Entities

    So far, all we’ve drawn is the static, unchanging (apart from dynamic lighting) world. That’s an important foundation, but it’s certainly not a game; now we need to add moving objects. These objects fall into four very different categories: BSP models, polygon models, sprites, and particles.

    -

    BSP Models

    +

    BSP Models

    BSP models are just like the world, except that they can move. Examples include doors, moving bridges, and health and ammo boxes. The way these are rendered is by clipping their polygons into the world BSP tree, so each polygon fragment is in only one leaf. Then these fragments are added to the edge list, just like world polygons, and scanned out, along with the rest of the world, when the edge list is processed. The only trick here is front-to-back ordering. Each BSP model polygon fragment is given the BSP sorting order of the leaf in which it resides, allowing it to sort properly versus the world polygons. If two or more polygons from different BSP models are in the same leaf, however, BSP ordering is no longer useful, so we then sort those polygons by 1/z, calculated from the polygons’ plane equations.

    @@ -49,7 +42,7 @@

    BSP models take some extra time because of the cost of clipping them into the world BSP tree, but render just as fast as the rest of the world, again with no overdraw, so closed doors, for example, block drawing of whatever’s on the other side (although it’s still necessary to transform, project, and add to the edge list the polygons the door occludes, because they’re still in the PVS—they’re potentially visible if the door opens). This makes BSP models most suitable for fairly simple structures, such as boxes, which have relatively few polygons to clip, and cause relatively few edges to be added to the edge list.

    -

    Polygon Models and Z-Buffering

    +

    Polygon Models and Z-Buffering

    Polygon models, such as monsters, weapons, and projectiles, consist of a triangle mesh with front and back skins stretched over the model. For speed, the triangles are drawn with affine texture mapping; the triangles are small enough, and the models are generally distant enough, that affine distortion isn’t visible. (However, it is visible on the player’s weapon; this caused a lot of extra work for the artists, and we will probably implement a perspective-correct polygon-model rasterizer in Quake 2 for this specific purpose.) The triangles are also Gouraud shaded; interestingly, the light vector used to shade the models is always from the same direction, and has no relation to any actual lights in the world (although it does vary in intensity, along with the model’s ambient lighting, to match the brightness of the spot the player is standing above in the world). Even this highly inaccurate lighting works well, though; the Gouraud shading makes models look much more three-dimensional, and varying the lighting in even so crude a way allows hiding in shadows and illumination by explosions and muzzle flashes.

    @@ -78,10 +71,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-06.html b/70-06.html index 38d07ba..43cf894 100644 --- a/70-06.html +++ b/70-06.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -37,25 +30,25 @@


    -

    The Subdivision Rasterizer

    +

    The Subdivision Rasterizer

    This rasterizer, which we call the subdivision rasterizer, first draws all the vertices in the model. Then it takes each front-facing triangle, and determines if it has a side that’s at least two pixels long. If it does, we split that side into two pieces at the pixel nearest to the middle (using adds and shifts to average the endpoints of that side), draw the vertex at the split point, and process each of the two split triangles recursively, until we get down to triangles that have only one-pixel sides and hence have nothing left to draw. This approach is hideously slow and quite ugly (due to inaccuracies from integer quantization) for 100-pixel triangles—but it’s very fast for, say, five-pixel triangles, and is indistinguishable from more accurate rasterization when a model is 25 or 50 feet away. Better yet, the subdivider is ridiculously simple—a few dozen lines of code, far simpler than the affine rasterizer—and was implemented in an evening, immediately making the drawing of distant models about three times as fast, a very good return for a bit of conceptual work. The affine rasterizer got fairly close to the same performance with further optimization—in the range of 10% to 50% slower—but that took weeks of difficult programming.

    We switch between the two rasterizers based on the model’s distance and average triangle size, and in almost any scene, most models are far enough away so subdivision rasterization is used. There are undoubtedly faster ways yet to rasterize distant models adequately well, but the subdivider was clearly a win, and is a good example of how thinking in a radically different direction can pay off handsomely.

    -

    Sprites

    +

    Sprites

    We had hoped to be able to eliminate sprites completely, making Quake 100% 3-D, but sprites—although sometimes very visibly 2-D—were used for a few purposes, most noticeably the cores of explosions. As of CGDC last year, explosions consisted of an exploding spray of particles (discussed below), but there just wasn’t enough visual punch with that representation; adding a series of sprites animating an explosion did the trick. (In hindsight, we probably should have made the explosions polygon models rather than sprites; it would have looked about as good, and the few sprites we used didn’t justify the considerable amount of code and programming time required to support them.) Drawing a sprite is similar to drawing a normal polygon, complete with perspective correction, although of course the inner loop must detect and skip over transparent pixels, and must also perform z-buffering.

    -

    Particles

    +

    Particles

    The last drawing entity type is particles. Each particle is a solid-colored rectangle, scaled by distance from the viewer and drawn with z-buffering. There can be up to 2,000 particles in a scene, and they are used for rocket trails, explosions, and the like. In one sense, particles are very primitive technology, but they allow effects that would be extremely difficult to do well with the other types of entities, and they work well in tandem with other entities, as, for example, providing a trail of fire behind a polygon-model lava ball that flies into the air, or generating an expanding cloud around a sprite explosion core.

    -

    How We Spent Our Summer Vacation: After Shipping Quake

    +

    How We Spent Our Summer Vacation: After Shipping Quake

    Since shipping Quake in the summer of 1996, we’ve extended it in several ways: We’ve worked with Rendition to port it to the Verite accelerator chip, we’ve ported it to OpenGL, we’ve ported it to Win32, we’ve done QuakeWorld, and we’ve added features for Quake 2. I’ll discuss each of these briefly.

    -

    Verite Quake

    +

    Verite Quake

    Verite Quake (VQuake) was the first hardware-accelerated version of Quake. It looks extremely good, due to bilinear texture filtering, which eliminates most pixel aliasing, and because it provides good performance at higher resolutions such as 512x384 and 640x480. Implementing VQuake proved to be an interesting task, for two reasons: The Verite chip’s fill rate was marginal for Quake’s needs, and Verite contains a programmable RISC chip, enabling more sophisticated processing than most 3-D accelerators. The need to squeeze as much performance as possible out of Verite ruled out the use of a standard API such as Direct 3D or OpenGL; instead, VQuake uses Rendition’s proprietary API, Speedy3D, with the addition of some special calls and custom Verite code.

    @@ -82,10 +75,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-07.html b/70-07.html index 658edd4..501b536 100644 --- a/70-07.html +++ b/70-07.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -37,7 +30,7 @@


    -

    GLQuake

    +

    GLQuake

    The second (and, according to current plans, last) port of Quake to a hardware accelerator was an OpenGL version, GLQuake, a native Win32 application. I have no intention of getting into the 3-D API wars currently raging; the observation I want to make here is that GLQuake uses two-pass alpha lighting, and runs very well on fast chips such as the 3Dfx, but rather slowly on most of the current group of accelerators. The accelerators coming out this year should all run GLQuake fine, however. It’s also worth noting that we’ll be using two-pass alpha lighting in the N64 port of Quake; in fact, it looks like the N64’s hardware is capable of performing both texture-tiling and alpha-lighting in a single pass, which is pretty much an ideal hardware-acceleration architecture: It’s as good looking and generally faster than surface caching, without the need to build, download, and cache surfaces, and much better looking and about as fast as Gouraud shading. We hope to see similar capabilities implemented in PC accelerators and exposed by 3-D APIs in the near future.

    @@ -53,13 +46,13 @@

    Both alpha-blending and z-buffering are relatively new to PC games, but are standard equipment on accelerators, and it’s a lot of fun seeing what sorts of previously very difficult effects can now be up and working in a matter of hours.

    -

    WinQuake

    +

    WinQuake

    I’m not going to spend much time on the Win32 port of Quake; most of what I learned doing this consists of tedious details that are doubtless well covered elsewhere, and frankly it wasn’t a particularly interesting task and was harder than I expected, and I’m pretty much tired of the whole thing. However, I will say that Win32 is clearly the future, especially now that NT is coming on strong, and like it or not, you had best learn to write games for Win32. Also, Internet gaming is becoming ever more important, and Win32’s built-in TCP/IP support is a big advantage over DOS; that alone was enough to convince us we had to port Quake. As a last comment, I’d say that it is nice to have Windows take care of device configuration and interfacing—now if only we could get manufacturers to write drivers for those devices that actually worked reliably! This will come as no surprise to veteran Windows programmers, who have suffered through years of buggy 2-D Windows drivers, but if you’re new to Windows programming, be prepared to run into and learn to work around—or at least document in your readme files—driver bugs on a regular basis.

    Still, when you get down to it, the future of gaming is a networked Win32 world, and that’s that, so if you haven’t already moved to Win32, I’d say it’s time.

    -

    QuakeWorld

    +

    QuakeWorld

    QuakeWorld is a native Win32 multiplayer-only version of Quake, and was done as a learning experience; it is not a commercial product, but is freely distributed on the Internet. The idea behind it was to try to improve the multiplayer experience, especially for people linked by modem, by reducing actual and perceived latency. Before I discuss QuakeWorld, however, I should discuss the evolution of Quake’s multiplayer code.

    @@ -80,10 +73,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-08.html b/70-08.html index cf095fb..3426a23 100644 --- a/70-08.html +++ b/70-08.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -59,7 +52,7 @@

    The second way in which QuakeWorld attacks latency is by not interpolating. The player is actually predicted well ahead of the latest server packet (after all, the client has all the information needed to move the player, unless an outside force intervenes), giving very responsive control. The rest of the world is drawn as of the latest server packet; this is jerkier than Quake, again showing that smoothness is often a tradeoff for latency. The player’s prediction may, of course, result in a minor paradox; for example, if an explosion turns out to have knocked the player sideways, the player’s location may suddenly jump without warning as the server packet arrives with the correct location. In the latest version of QuakeWorld, the other players are predicted as well, with consequently more frequent paradoxes, but smoother, more convincing motion. Platforms and doors are still not predicted, and consequently are still pretty jerky. It is, of course, possible to predict more and more objects into the future; it’s a tradeoff of smoothness and perceived low latency for the frustration of paradoxes—and that’s the way it’s going to stay until most people are connected to the Internet by something better than modems.

    -

    Quake 2

    +

    Quake 2

    I can’t talk in detail about Quake 2 as a game, but I can describe some interesting technology features. The Quake 2 rendering engine isn’t going to change that much from Quake; the improvements are largely in areas such as physics, gameplay, artwork, and overall design. The most interesting graphics change is in the preprocessing, where John has added support for radiosity lighting; that is, the ability to put a light source into the world and have the light bounced around the world realistically. This is sometimes terrific—it makes for great glowing light around lava and hanging light panels—but in other cases it’s less spectacular than the effects that designers can get by placing lots of direct-illumination light sources in a room, so the two methods can be used as needed. Also, radiosity is very computationally expensive, approximately as expensive as BSPing. Most of the radiosity demos I’ve seen have been in one or two rooms, and the order of the problem goes up tremendously on whole Quake levels. Here’s another case where the PVS is essential; without it, radiosity processing time would be O(polygons2), but with the PVS it’s O(polygons*average_potentially_visible_polygons), which is over an order of magnitude less (and increases approximately linearly, rather than as a squared function, with greater-level complexity).

    @@ -80,10 +73,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/70-09.html b/70-09.html index c385d4d..cce0fca 100644 --- a/70-09.html +++ b/70-09.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Quake: A Post-Mortem and a Glimpse into the Future - - @@ -45,7 +38,7 @@

    By the way, Quake 2 is currently being developed as a native Win32 app only; no DOS version is planned.

    -

    Looking Forward

    +

    Looking Forward

    In my address to the Computer Game Developer’s Conference in 1996, I said that it wasn’t a bad time to start up a game company aimed at hardware-only rasterization, and trying to make a game that leapfrogged the competition. It looks like I was probably a year early, because hardware took longer to ship than I expected, although there was a good living to be made writing games that hardware vendors could bundle with their boards. Now, though, it clearly is time. By Christmas 1997, there will be several million fast accelerators out there, and by Christmas 1998, there will be tens of millions. At the same time, vastly more people are getting access to the Internet, and it’s from the convergence of these two trends that I think the technology for the next generation of breakthrough real-time games will emerge.

    @@ -78,10 +71,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/about.html b/about.html index 46ce84f..ff97a02 100644 --- a/about.html +++ b/about.html @@ -10,16 +10,7 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Foreword - - - - - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Foreword @@ -37,7 +28,7 @@


    -

    Foreword

    +

    Foreword

    I got my start programming on Apple II computers at school, and almost all of my early work was on the Apple platform. After graduating, it quickly became obvious that I was going to have trouble paying my rent working in the Apple II market in the late eighties, so I was forced to make a very rapid move into the Intel PC environment.

    @@ -93,10 +84,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/about_author.html b/about_author.html index fb0d36f..80d7cb3 100644 --- a/about_author.html +++ b/about_author.html @@ -10,16 +10,8 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: About the Author - - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: About the Author - - @@ -37,7 +29,7 @@


    -

    Acknowledgments

    +

    Acknowledgments

    There are many people to thank—because this book was written over many years, in many different settings, an unusually large number of people have played a part in making this book possible. Thanks to Dan Illowsky for not only contributing ideas and encouragement, but also getting me started writing articles long ago, when I lacked the confidence to do it on my own—and for teaching me how to handle the business end of things. Thanks to Will Fastie for giving me my first crack at writing for a large audience in the long-gone but still-missed PC Tech Journal, and for showing me how much fun it could be in his even longer-vanished but genuinely terrific column in Creative Computing (the most enjoyable single column I have ever read in a computer magazine; I used to haunt the mailbox around the beginning of the month just to see what Will had to say). Thanks to Robert Keller, Erin O’Connor, Liz Oakley, Steve Baker, and the rest of the cast of thousands that made Programmer’s Journal a uniquely fun magazine—especially Erin, who did more than anyone to teach me the proper use of the English language. (To this day, Erin will still patiently explain to me when one should use “that” and when one should use “which,” even though eight years of instruction on this and related topics have left no discernible imprint on my brain.) Thanks to Tami Zemel, Monica Berg, and the rest of the Dr. Dobb’s Journal crew for excellent, professional editing, and for just being great people. Thanks to the Coriolis gang for their tireless hard work: Jeff Duntemann, Kim Eoff, Jody Kent, Robert Clarfield, and Anthony Stock. Thanks to Jack Tseng for teaching me a lot about graphics hardware, and even more about how much difference hard work can make. Thanks to John Cockerham, David Stafford, Terje Mathisen, the BitMan, Chris Hecker, Jim Mackraz, Melvin Lafitte, John Navas, Phil Coleman, Anton Truenfels, John Carmack, John Miles, John Bridges, Jim Kent, Hal Hardenbergh, Dave Miller, Steve Levy, Jack Davis, Duane Strong, Daev Rohr, Bill Weber, Dan Gochnauer, Patrick Milligan, Tom Wilson, Peter Klerings, Dave Methvin, Mick Brown, the people in the ibm.pc/fast.code topic on Bix, and all the rest of you who have been so generous with your ideas and suggestions. I’ve done my best to acknowledge contributors by name in this book, but if your name is omitted, my apologies, and consider yourself thanked; this book could not have happened without you. And, of course, thanks to Shay and Emily for their generous patience with my passion for writing and computers.

    @@ -58,10 +50,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/appendix-a.html b/appendix-a.html index 6019871..4eea79b 100644 --- a/appendix-a.html +++ b/appendix-a.html @@ -8,16 +8,8 @@ - - - - - - - + - - @@ -37,7 +29,7 @@


    -

    Afterword

    +

    Afterword

    If you’ve followed me this far, you might agree that we’ve come through some rough country. Still, I’m of the opinion that hard-won knowledge is the best knowledge, not only because it sticks to you better, but also because winning a hard race makes it easier to win the next one.

    @@ -80,10 +72,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/book-index.html b/book-index.html index 4a26ec3..634b75a 100644 --- a/book-index.html +++ b/book-index.html @@ -10,16 +10,8 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Index - - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Index - - @@ -37,7 +29,7 @@


    -

    Index

    +

    Index

    Numbers
    @@ -9122,10 +9114,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/index.html b/index.html index 0b55e1b..ccbff56 100644 --- a/index.html +++ b/index.html @@ -10,16 +10,9 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Table of Contents - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Table of Contents - - @@ -2078,10 +2071,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - + diff --git a/intro.html b/intro.html index d781cb7..4e3b4f6 100644 --- a/intro.html +++ b/intro.html @@ -10,16 +10,7 @@ - Michael Abrash's Graphics Programming Black Book Special Edition: Introduction - - - - - - - - - + Michael Abrash's Graphics Programming Black Book Special Edition: Introduction @@ -37,7 +28,7 @@


    -

    Introduction

    +

    Introduction

    What was it like working with John Carmack on Quake? Like being strapped onto a rocket during takeoff—in the middle of a hurricane. It seemed like the whole world was watching, waiting to see if id Software could top Doom; every casual e-mail tidbit or conversation with a visitor ended up posted on the Internet within hours. And meanwhile, we were pouring everything we had into Quake’s technology; I’d often come in in the morning to find John still there, working on a new idea so intriguing that he couldn’t bear to sleep until he had tried it out. Toward the end, when I spent most of my time speeding things up, I would spend the day in a trance writing optimized assembly code, stagger out of the Town East Tower into the blazing Texas heat, and somehow drive home on LBJ Freeway without smacking into any of the speeding pickups whizzing past me on both sides. At home, I’d fall into a fitful sleep, then come back the next day in a daze and do it again. Everything happened so fast, and under so much pressure, that sometimes I wonder how any of us made it through that without completely burning out.

    @@ -76,10 +67,6 @@
    Graphics Programming Black Book © 2001 Michael Abrash -
    - - - - +