Compare commits
35 changed files with 536 additions and 535 deletions
4
Makefile
Normal file → Executable file
4
Makefile
Normal file → Executable file
|
|
@ -7,12 +7,12 @@ all: html epub mobi
|
||||||
html:
|
html:
|
||||||
rm -rf out/html && mkdir -p out/html
|
rm -rf out/html && mkdir -p out/html
|
||||||
cp -r images html/book.css out/html/
|
cp -r images html/book.css out/html/
|
||||||
pandoc --to html5+smart -o out/html/black-book.html --section-divs --toc --standalone --template=html/template.html $(FILES)
|
pandoc -S --to html5 -o out/html/black-book.html --section-divs --toc --standalone --template=html/template.html $(FILES)
|
||||||
|
|
||||||
epub:
|
epub:
|
||||||
mkdir -p out
|
mkdir -p out
|
||||||
rm -f out/black-book.epub
|
rm -f out/black-book.epub
|
||||||
pandoc --to epub3+smart -o out/black-book.epub --epub-cover-image images/cover.png --toc --epub-chapter-level=2 --data-dir=epub --template=epub/template.html $(FILES)
|
pandoc -S --to epub3 -o out/black-book.epub --epub-cover-image images/cover.png --toc --epub-chapter-level=2 --data-dir=epub --template=epub/template.html $(FILES)
|
||||||
|
|
||||||
mobi:
|
mobi:
|
||||||
rm -f out/black-book.mobi
|
rm -f out/black-book.mobi
|
||||||
|
|
|
||||||
24
README.md
24
README.md
|
|
@ -2,38 +2,36 @@
|
||||||
|
|
||||||
This is the source for an ebook version of Michael Abrash's Black Book of Graphics Programming (Special Edition), originally published in 1997 and [released online for free in 2001](http://www.drdobbs.com/parallel/graphics-programming-black-book/184404919).
|
This is the source for an ebook version of Michael Abrash's Black Book of Graphics Programming (Special Edition), originally published in 1997 and [released online for free in 2001](http://www.drdobbs.com/parallel/graphics-programming-black-book/184404919).
|
||||||
|
|
||||||
Reproduced with blessing of Michael Abrash, converted and maintained by [James Gregory](mailto:james@jagregory.com).
|
Reproduced with permission of Michael Abrash, converted and maintained by [James Gregory](mailto:james@jagregory.com).
|
||||||
|
|
||||||
The [GitHub releases list](https://github.com/jagregory/abrash-black-book/releases) has an EPUB and Mobi version available for download, and you can find a mirror of the HTML version at [www.jagregory.com/abrash-black-book](http://www.jagregory.com/abrash-black-book/).
|
|
||||||
|
|
||||||
## How does this differ from the previously released versions?
|
## How does this differ from the previously released versions?
|
||||||
|
|
||||||
The book is now out of print, and hard to come by. Last time I checked, it was going for over $200 on eBay.
|
The book is now out of print, and hard to come by. Last time I checked it was going for over $200 on ebay.
|
||||||
|
|
||||||
The version which Michael and Dr. Dobbs released in 2001 was a collection of PDF files. That version is [still available](http://www.drdobbs.com/parallel/graphics-programming-black-book/184404919). However, the structure (multiple files) and the format (PDF) result in a poor user experience on an ebook reader or other mobile device.
|
The version which Michael and Dr. Dobbs released in 2001 was as a collection PDFs. This version is [still available](http://www.drdobbs.com/parallel/graphics-programming-black-book/184404919); however, the structure (multiple files) and the format (PDF) doesn't lend itself well to reading on a ebook reader or other mobile device.
|
||||||
|
|
||||||
This version has been thoroughly cleaned of artifacts and condensed into something which can easily be converted into an ebook-friendly format. You can read this version online at GitHub, or download any of the EPUB or Mobi releases. You can clone the repository and generate your own version with [pandoc](http://johnmacfarlane.net/pandoc/) if necessary.
|
This version has been thoroughly cleaned of artefacts and condensed into something which can be easily converted into a ebook friendly format. You can read this version online at Github, or download any of the Epub or Mobi releases. You can clone the repository and generate your own version with [pandoc](http://johnmacfarlane.net/pandoc/) if necessary.
|
||||||
|
|
||||||
## Contributing
|
## Contributing
|
||||||
|
|
||||||
Changes are welcome, especially conversion-related ones. If you spot any problems while reading, please [submit an issue](https://github.com/jagregory/abrash-black-book/issues) and I'll correct it. Pull requests are always welcome.
|
Changes are welcome, especially conversion related ones. If you spot any issues whilst reading, please submit an issue and I'll correct it. Pull Requests are always welcome.
|
||||||
|
|
||||||
Some larger changes could be made to improve the content. I'd love to see some of the images converted to a vector representation so we can provide higher-resolution versions. Formulas and equations could be typeset with [MathJax](http://www.mathjax.org/).
|
There's some larger changes that could be made to help preserve the content longer term. I'd love to see some of the images converted to a vector representation so we can provide higher-resolution versions, and similarly formulas and maths could be represented in MathML.
|
||||||
|
|
||||||
## Generating your own ebook
|
## Generating your own ebook
|
||||||
|
|
||||||
You need to have the following software installed and on your `PATH` before you begin:
|
You need to have the following software installed and on your `PATH` before you begin:
|
||||||
|
|
||||||
* [pandoc](http://johnmacfarlane.net/pandoc/) version 2.0 or greater for Markdown to HTML and EPUB conversion.
|
* [pandoc](http://johnmacfarlane.net/pandoc/) for Markdown to HTML and Epub conversion.
|
||||||
* [kindlegen](http://www.amazon.com/gp/feature.html?docId=1000765211) for Epub to Mobi conversion.
|
* [kindlegen](http://www.amazon.com/gp/feature.html?docId=1000765211) for Epub to Mobi conversion.
|
||||||
|
|
||||||
To generate an e-reader friendly version of the book, you can use `make` with one of the following options:
|
To generate an e-reader friendly version of the book, you can use `make` with one of the following options:
|
||||||
|
|
||||||
* `html` - build an HTML5 single-page version of the book
|
* `html` - build a HTML5 single-page version of the book
|
||||||
* `epub` - build an EPUB3 ebook
|
* `epub` - build an Epub3 ebook
|
||||||
* `mobi` - build a Kindle-friendly Mobi
|
* `mobi` - build a Kindle-friendly Mobi
|
||||||
* `all` - do all of the above
|
* `all` - do all of the above
|
||||||
|
|
||||||
Once complete, there will be an `out` directory with a `black-book.epub`, a `black-book.mobi` and an `html` directory with a `black-book.html` file.
|
Once complete, there'll be an `out` directory with a `black-book.epub`, a `black-book.mobi` and a `html` directory with a `black-book.html` file.
|
||||||
|
|
||||||
> Note: Generating a Mobi requires an EPUB to already exist. Also, Mobi generation can be *slow* because of compression. If you want a quick Mobi conversion you can just run `kindlegen out/black-book.epub`.
|
> Note: Generating a mobi requires an epub to already exist. Also, mobi generation can be *slow* because of compression. If you want a quick mobi conversion you can just run `kindlegen out/black-book.epub`.
|
||||||
|
|
|
||||||
BIN
images/13-01.jpg
Normal file
BIN
images/13-01.jpg
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 4.8 KiB |
BIN
images/13-01.png
BIN
images/13-01.png
Binary file not shown.
|
Before Width: | Height: | Size: 88 KiB |
|
|
@ -23,7 +23,7 @@ learn in the space of a few months on the PC.
|
||||||
|
|
||||||
The biggest benefit to me of actually making money as a programmer was
|
The biggest benefit to me of actually making money as a programmer was
|
||||||
the ability to buy all the books and magazines I wanted. I bought a lot.
|
the ability to buy all the books and magazines I wanted. I bought a lot.
|
||||||
I was in territory that I knew almost nothing about, so I read
|
I was in territory that I new almost nothing about, so I read
|
||||||
*everything* that I could get my hands on. Feature articles, editorials,
|
*everything* that I could get my hands on. Feature articles, editorials,
|
||||||
even advertisements held information for me to assimilate.
|
even advertisements held information for me to assimilate.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -254,7 +254,7 @@ requires over two and one-half minutes to checksum *one* file!
|
||||||
These results make it clear that it's folly to rely on your compiler's
|
These results make it clear that it's folly to rely on your compiler's
|
||||||
optimization to make your programs fast. Listing 1.1 is simply poorly
|
optimization to make your programs fast. Listing 1.1 is simply poorly
|
||||||
designed, and no amount of compiler optimization will compensate for
|
designed, and no amount of compiler optimization will compensate for
|
||||||
that failing. To drive home the point, Listings 1.2 and 1.3, which
|
that failing. To drive home the point, conListings 1.2 and 1.3, which
|
||||||
together are equivalent to Listing 1.1 except that the entire checksum
|
together are equivalent to Listing 1.1 except that the entire checksum
|
||||||
loop is written in tight assembly code. The assembly language
|
loop is written in tight assembly code. The assembly language
|
||||||
implementation is indeed faster than any of the C versions, as shown in
|
implementation is indeed faster than any of the C versions, as shown in
|
||||||
|
|
@ -500,7 +500,7 @@ your programs that directly affect response time. Notice, for example,
|
||||||
that I haven't bothered to implement a version of the checksum program
|
that I haven't bothered to implement a version of the checksum program
|
||||||
entirely in assembly; Listings 1.2 and 1.6 call assembly subroutines
|
entirely in assembly; Listings 1.2 and 1.6 call assembly subroutines
|
||||||
that handle the time-critical operations, but C is still used for
|
that handle the time-critical operations, but C is still used for
|
||||||
checking command-line parameters, opening files, printing, and the
|
checking command-line parameters, operning files, printing, and the
|
||||||
like.
|
like.
|
||||||
|
|
||||||
> 
|
> 
|
||||||
|
|
@ -521,7 +521,7 @@ Listing 1.4 is good, but let's see if there are other—perhaps less
|
||||||
obvious—ways to get the same results faster. Let's start by considering
|
obvious—ways to get the same results faster. Let's start by considering
|
||||||
why Listing 1.4 is so much better than Listing 1.1. Like `read()`,
|
why Listing 1.4 is so much better than Listing 1.1. Like `read()`,
|
||||||
`getc()` calls DOS to read from the file; the speed improvement of
|
`getc()` calls DOS to read from the file; the speed improvement of
|
||||||
Listing 1.4 over Listing 1.1 occurs because `getc()` reads many bytes
|
Listing 1.4 over Listing 1.1 occurs because `getc()` eads many bytes
|
||||||
at once via DOS, then manages those bytes for us. That's faster than
|
at once via DOS, then manages those bytes for us. That's faster than
|
||||||
reading them one at a time using `read()`—but there's no reason to
|
reading them one at a time using `read()`—but there's no reason to
|
||||||
think that it's faster than having our program read and manage blocks
|
think that it's faster than having our program read and manage blocks
|
||||||
|
|
@ -667,7 +667,7 @@ indeed make a significant difference. Table 1.1 indicates that the
|
||||||
optimized version of Listing 1.5 produced by Microsoft C outperforms an
|
optimized version of Listing 1.5 produced by Microsoft C outperforms an
|
||||||
unoptimized version of the same code by more than 60 percent. What's
|
unoptimized version of the same code by more than 60 percent. What's
|
||||||
more, a mostly-assembly version of Listing 1.5, shown in Listings 1.6
|
more, a mostly-assembly version of Listing 1.5, shown in Listings 1.6
|
||||||
and 1.7, outperforms even the best-optimized C version of Listing 1.5 by 26
|
and 1.7, outperforms even the best-optimized C version of List1.5 by 26
|
||||||
percent. These are considerable improvements, well worth pursuing—once
|
percent. These are considerable improvements, well worth pursuing—once
|
||||||
the design has been maxed out.
|
the design has been maxed out.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -89,7 +89,7 @@ much different from the original, and in fact still contains exactly the
|
||||||
same number of instructions, the performance of the entire subroutine
|
same number of instructions, the performance of the entire subroutine
|
||||||
improved by about 10 percent from just this one change. (Incidentally,
|
improved by about 10 percent from just this one change. (Incidentally,
|
||||||
that wasn't the end of the optimization; I eliminated the `DEC` and
|
that wasn't the end of the optimization; I eliminated the `DEC` and
|
||||||
`JNZ` instructions by expanding the four iterations of the loop—but
|
`JNJ` instructions by expanding the four iterations of the loop—but
|
||||||
that's a tale for another chapter.)
|
that's a tale for another chapter.)
|
||||||
|
|
||||||
The point is this: To write truly superior assembly programs, you need
|
The point is this: To write truly superior assembly programs, you need
|
||||||
|
|
@ -200,7 +200,7 @@ enough.
|
||||||
The single most critical aspect of the hardware, and the one about which
|
The single most critical aspect of the hardware, and the one about which
|
||||||
it is hardest to learn, is the CPU. The x86 family CPUs have a complex,
|
it is hardest to learn, is the CPU. The x86 family CPUs have a complex,
|
||||||
irregular instruction set, and, unlike most processors, they are neither
|
irregular instruction set, and, unlike most processors, they are neither
|
||||||
straightforward nor well-documented true code performance. What's more,
|
straightforward nor wellregarding true code performance. What's more,
|
||||||
assembly is so difficult to learn that most articles and books that
|
assembly is so difficult to learn that most articles and books that
|
||||||
present assembly code settle for code that just works, rather than code
|
present assembly code settle for code that just works, rather than code
|
||||||
that pushes the CPU to its limits. In fact, since most articles and
|
that pushes the CPU to its limits. In fact, since most articles and
|
||||||
|
|
|
||||||
|
|
@ -163,7 +163,7 @@ presented in Chapter K on the companion CD-ROM.
|
||||||
; in when ZTimerOn was called.
|
; in when ZTimerOn was called.
|
||||||
;
|
;
|
||||||
|
|
||||||
Code segment word public 'CODE'
|
Code segment word public ‘CODE'
|
||||||
assumecs: Code, ds:nothing
|
assumecs: Code, ds:nothing
|
||||||
public ZTimerOn, ZTimerOff, ZTimerReport
|
public ZTimerOn, ZTimerOff, ZTimerReport
|
||||||
|
|
||||||
|
|
@ -220,7 +220,7 @@ OriginalFlags db ? ; storage for upper byte of
|
||||||
; ZTimerOn called
|
; ZTimerOn called
|
||||||
TimedCount dw ? ; timer 0 count when the timer
|
TimedCount dw ? ; timer 0 count when the timer
|
||||||
; is stopped
|
; is stopped
|
||||||
ReferenceCount dw ? ; number of counts required to
|
ReferenceCount dw ; number of counts required to
|
||||||
; execute timer overhead code
|
; execute timer overhead code
|
||||||
OverflowFlag db ? ; used to indicate whether the
|
OverflowFlag db ? ; used to indicate whether the
|
||||||
; timer overflowed during the
|
; timer overflowed during the
|
||||||
|
|
@ -229,28 +229,28 @@ OverflowFlag db ? ; used to indicate whether the
|
||||||
; String printed to report results.
|
; String printed to report results.
|
||||||
;
|
;
|
||||||
OutputStr label byte
|
OutputStr label byte
|
||||||
db 0dh, 0ah, 'Timed count: ', 5 dup (?)
|
db 0dh, 0ah, ‘Timed count: ‘, 5 dup (?)
|
||||||
ASCIICountEnd labelbyte
|
ASCIICountEnd labelbyte
|
||||||
db ' microseconds', 0dh, 0ah
|
db ‘ microseconds', 0dh, 0ah
|
||||||
db '$'
|
db ‘$'
|
||||||
;
|
;
|
||||||
; String printed to report timer overflow.
|
; String printed to report timer overflow.
|
||||||
;
|
;
|
||||||
OverflowStr label byte
|
OverflowStr label byte
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '****************************************************'
|
db ‘****************************************************'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '* The timer overflowed, so the interval timed was *'
|
db ‘* The timer overflowed, so the interval timed was *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '* too long for the precision timer to measure. *'
|
db ‘* too long for the precision timer to measure. *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '* Please perform the timing test again with the *'
|
db ‘* Please perform the timing test again with the *'
|
||||||
db0dh, 0ah
|
db0dh, 0ah
|
||||||
db '* long-period timer. *'
|
db ‘* long-period timer. *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '****************************************************'
|
db ‘****************************************************'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '$'
|
db ‘$'
|
||||||
|
|
||||||
; ********************************************************************
|
; ********************************************************************
|
||||||
; * Routine called to start timing. *
|
; * Routine called to start timing. *
|
||||||
|
|
@ -671,7 +671,7 @@ count reaches zero, the timer turns over and starts counting down again
|
||||||
without stopping, and a pulse is generated for a single clock period.
|
without stopping, and a pulse is generated for a single clock period.
|
||||||
While the pulse is not held for nearly as long as in square wave mode,
|
While the pulse is not held for nearly as long as in square wave mode,
|
||||||
it doesn't matter, since the 8259 interrupt controller is configured in
|
it doesn't matter, since the 8259 interrupt controller is configured in
|
||||||
the PC to be edge-triggered and hence cares only about the existence of a pulse
|
the PC to be edgeand hence cares only about the existence of a pulse
|
||||||
from timer 0, not the duration of the pulse. As a result, timer 0
|
from timer 0, not the duration of the pulse. As a result, timer 0
|
||||||
continues to generate timer interrupts in divide-by-N mode, and the
|
continues to generate timer interrupts in divide-by-N mode, and the
|
||||||
system clock continues to maintain good time.
|
system clock continues to maintain good time.
|
||||||
|
|
@ -688,7 +688,7 @@ the Zen timer shown in Listing 3.1 supports.
|
||||||
In fact, the Zen timer shown in Listing 3.1 can only time intervals of
|
In fact, the Zen timer shown in Listing 3.1 can only time intervals of
|
||||||
up to about 54 ms in length, since that is the period of time that can
|
up to about 54 ms in length, since that is the period of time that can
|
||||||
be measured by timer 0 before its count turns over and repeats.
|
be measured by timer 0 before its count turns over and repeats.
|
||||||
Fifty-four ms may not seem like a very long time, but even a CPU as slow
|
fifty-four ms may not seem like a very long time, but even a CPU as slow
|
||||||
as the 8088 can perform more than 1,000 divides in 54 ms, and division
|
as the 8088 can perform more than 1,000 divides in 54 ms, and division
|
||||||
is the single instruction that the 8088 performs most slowly. If a
|
is the single instruction that the 8088 performs most slowly. If a
|
||||||
measured period turns out to be longer than 54 ms (that is, if timer 0
|
measured period turns out to be longer than 54 ms (that is, if timer 0
|
||||||
|
|
@ -730,10 +730,10 @@ restart until the timing interval ends, losing time all the while.
|
||||||
|
|
||||||
The effects on the system time of the Zen timer aren't a matter for
|
The effects on the system time of the Zen timer aren't a matter for
|
||||||
great concern, as they are temporary, lasting only until the next warm
|
great concern, as they are temporary, lasting only until the next warm
|
||||||
or cold boot. System that have battery-backed clocks, (AT-style machines; that
|
or cold boot. System that have batteryclocks, (AT-style machines; that
|
||||||
is, virtually all machines in common use) automatically reset the
|
is, virtually all machines in common use) automatically reset the
|
||||||
correct time whenever the computer is booted, and systems without
|
correct time whenever the computer is booted, and systems without
|
||||||
battery-backed clocks prompt for the correct date and time when booted.
|
battery-clocks prompt for the correct date and time when booted.
|
||||||
Also,repeated use of the Zen timer usually makes the system clock slow
|
Also,repeated use of the Zen timer usually makes the system clock slow
|
||||||
by at most a total of a few seconds, unless code that takes much longer
|
by at most a total of a few seconds, unless code that takes much longer
|
||||||
than 54 ms to run is timed (in which case the Zen timer will notify you
|
than 54 ms to run is timed (in which case the Zen timer will notify you
|
||||||
|
|
@ -789,8 +789,8 @@ from timer counts to microseconds, and prints the resulting time in
|
||||||
microseconds to the standard output.
|
microseconds to the standard output.
|
||||||
|
|
||||||
Note that `ZTimerReport` need not be called immediately after
|
Note that `ZTimerReport` need not be called immediately after
|
||||||
`ZTimerOff`. In fact, after a given call to `ZTimerOff`,
|
`ZTimerOff`. In fact, after a given call to `ZTimerOff,
|
||||||
`ZTimerReport` can be called at any time right up until the next call to
|
ZTimerReport` can be called at any time right up until the next call to
|
||||||
`ZTimerOn`.
|
`ZTimerOn`.
|
||||||
|
|
||||||
You may want to use the Zen timer to measure several portions of a
|
You may want to use the Zen timer to measure several portions of a
|
||||||
|
|
@ -880,7 +880,7 @@ performance will be similar even on different IBM models; in fact, quite
|
||||||
the opposite is true. For example, every PS/2 computer, even the
|
the opposite is true. For example, every PS/2 computer, even the
|
||||||
relatively slow Model 30, executes code much faster than does a PC or
|
relatively slow Model 30, executes code much faster than does a PC or
|
||||||
XT. As another example, I set out to do the timings for my earlier book
|
XT. As another example, I set out to do the timings for my earlier book
|
||||||
*Zen of Assembly Language* on an XT-compatible computer, only to find that the
|
*Zen of Assembly Language* on an XTcomputer, only to find that the
|
||||||
computer wasn't quite IBM-compatible regarding code performance. The
|
computer wasn't quite IBM-compatible regarding code performance. The
|
||||||
differences were minor, mind you, but my experience illustrates the risk
|
differences were minor, mind you, but my experience illustrates the risk
|
||||||
of assuming that a specific make of computer will perform in a certain
|
of assuming that a specific make of computer will perform in a certain
|
||||||
|
|
@ -913,11 +913,11 @@ and should contain calls to `ZTimerOn` and `ZTimerOff` .
|
||||||
;
|
;
|
||||||
; By Michael Abrash
|
; By Michael Abrash
|
||||||
;
|
;
|
||||||
mystack segment para stack 'STACK'
|
mystack segment para stack ‘STACK'
|
||||||
db 512 dup(?)
|
db 512 dup(?)
|
||||||
mystack ends
|
mystack ends
|
||||||
;
|
;
|
||||||
Code segment para public 'CODE'
|
Code segment para public ‘CODE'
|
||||||
assume cs:Code, ds:Code
|
assume cs:Code, ds:Code
|
||||||
extrnZTimerOn:near, ZTimerOff:near, ZTimerReport:near
|
extrnZTimerOn:near, ZTimerOff:near, ZTimerReport:near
|
||||||
Start proc near
|
Start proc near
|
||||||
|
|
@ -1111,8 +1111,8 @@ pztime <filename>
|
||||||
In fact, that's exactly how I timed each of the listings in this book.
|
In fact, that's exactly how I timed each of the listings in this book.
|
||||||
Code fragments you write yourself can be timed in just the same way. If
|
Code fragments you write yourself can be timed in just the same way. If
|
||||||
you wish to time code directly in place in your programs, rather than in
|
you wish to time code directly in place in your programs, rather than in
|
||||||
the test-bed program of Listing 3.2, simply insert calls to `ZTimerOn`,
|
the test-bed program of Listing 3.2, simply insert calls to `ZTimerOn,
|
||||||
`ZTimerOff`, and `ZTimerReport` in the appropriate places and link
|
ZTimerOff`, and `ZTimerReport` in the appropriate places and link
|
||||||
PZTIMER to your program.
|
PZTIMER to your program.
|
||||||
|
|
||||||
### The Long-Period Zen Timer
|
### The Long-Period Zen Timer
|
||||||
|
|
@ -1303,7 +1303,7 @@ computers.
|
||||||
; All registers and all flags are preserved by all routines.
|
; All registers and all flags are preserved by all routines.
|
||||||
;
|
;
|
||||||
|
|
||||||
Code segment word public 'CODE'
|
Code segment word public ‘CODE'
|
||||||
assume cs: Code, ds:nothing
|
assume cs: Code, ds:nothing
|
||||||
public ZTimerOn, ZTimerOff, ZTimerReport
|
public ZTimerOn, ZTimerOff, ZTimerReport
|
||||||
|
|
||||||
|
|
@ -1384,10 +1384,10 @@ ReferenceCount dw ? ;number of counts required to
|
||||||
; String printed to report results.
|
; String printed to report results.
|
||||||
;
|
;
|
||||||
OutputStr labelbyte
|
OutputStr labelbyte
|
||||||
db 0dh, 0ah, 'Timed count: '
|
db 0dh, 0ah, ‘Timed count: ‘
|
||||||
TimedCountStr db10 dup (?)
|
TimedCountStr db10 dup (?)
|
||||||
db' microseconds', 0dh, 0ah
|
db' microseconds', 0dh, 0ah
|
||||||
db '$'
|
db ‘$'
|
||||||
;
|
;
|
||||||
; Temporary storage for timed count as it's divided down by powers
|
; Temporary storage for timed count as it's divided down by powers
|
||||||
; of ten when converting from doubleword binary to ASCII.
|
; of ten when converting from doubleword binary to ASCII.
|
||||||
|
|
@ -1417,7 +1417,7 @@ PowersOfTenEnd label word
|
||||||
;
|
;
|
||||||
TurnOverStrlabelbyte
|
TurnOverStrlabelbyte
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '****************************************************'
|
db ‘****************************************************'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db'* Either midnight passed or an hour or more passed *'
|
db'* Either midnight passed or an hour or more passed *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
|
|
@ -1429,11 +1429,11 @@ TurnOverStr label byte
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db'* run to be timed by the long-period Zen timer. *'
|
db'* run to be timed by the long-period Zen timer. *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '* Suggestions: use the DOS TIME command, the DOS *'
|
db ‘* Suggestions: use the DOS TIME command, the DOS *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '* time function, or a watch. *'
|
db ‘* time function, or a watch. *'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db '****************************************************'
|
db ‘****************************************************'
|
||||||
db 0dh, 0ah
|
db 0dh, 0ah
|
||||||
db'$'
|
db'$'
|
||||||
|
|
||||||
|
|
@ -1896,7 +1896,7 @@ substantially.
|
||||||
|
|
||||||
Finally, please note that the *precision* Zen timer works perfectly well
|
Finally, please note that the *precision* Zen timer works perfectly well
|
||||||
on both PS/2 and non-PS/2 computers. The PS/2 and 8253 considerations
|
on both PS/2 and non-PS/2 computers. The PS/2 and 8253 considerations
|
||||||
we've just discussed apply *only* to the long-period Zen timer.
|
we've just discussed apply *only* to the longZen timer.
|
||||||
|
|
||||||
### Example Use of the Long-Period Zen Timer
|
### Example Use of the Long-Period Zen Timer
|
||||||
|
|
||||||
|
|
@ -1932,11 +1932,11 @@ timing.
|
||||||
;
|
;
|
||||||
; By Michael Abrash
|
; By Michael Abrash
|
||||||
;
|
;
|
||||||
mystack segment para stack 'STACK'
|
mystack segment para stack ‘STACK'
|
||||||
db 512 dup(?)
|
db 512 dup(?)
|
||||||
mystack ends
|
mystack ends
|
||||||
;
|
;
|
||||||
Code segment para public 'CODE'
|
Code segment para public ‘CODE'
|
||||||
assume cs:Code, ds:Code
|
assume cs:Code, ds:Code
|
||||||
extrn ZTimerOn:near, ZTimerOff:near, ZTimerReport:near
|
extrn ZTimerOn:near, ZTimerOff:near, ZTimerReport:near
|
||||||
Startproc near
|
Startproc near
|
||||||
|
|
@ -2126,7 +2126,7 @@ be dealt with here: small code model and large; I'll tackle the simpler
|
||||||
one, the small code model, first.
|
one, the small code model, first.
|
||||||
|
|
||||||
Altering the Zen timer for linking to a small code model C program
|
Altering the Zen timer for linking to a small code model C program
|
||||||
involves the following steps: Change `ZTimerOn` to
|
involves the following steps: `C` hange `ZTimerOn` to
|
||||||
`_ZTimerOn`, change `ZTimerOff` to `_ZTimerOff`, change
|
`_ZTimerOn`, change `ZTimerOff` to `_ZTimerOff`, change
|
||||||
`ZTimerReport` to `_ZTimerReport`, and change `Code` to
|
`ZTimerReport` to `_ZTimerReport`, and change `Code` to
|
||||||
`_TEXT` . Figure 3.2 shows the line numbers and new states of all
|
`_TEXT` . Figure 3.2 shows the line numbers and new states of all
|
||||||
|
|
@ -2190,13 +2190,13 @@ call near ptr ReferenceZTimerOn
|
||||||
(and likewise for `ReferenceZTimerOff` ), which works because
|
(and likewise for `ReferenceZTimerOff` ), which works because
|
||||||
`ReferenceZTimerOn` is in the same segment as the calling code. This
|
`ReferenceZTimerOn` is in the same segment as the calling code. This
|
||||||
is normally a great optimization, being both smaller and faster than a
|
is normally a great optimization, being both smaller and faster than a
|
||||||
far call.
|
far call. However, it's not so great for the Zen
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
However, it's not so great for the Zen timer, because our purpose in calling the reference timing code is to
|
timer, because our purpose in calling the reference timing code is to
|
||||||
determine exactly how much time is taken by overhead code—including the
|
determine exactly how much time is taken by overhead code—including the
|
||||||
far calls to `ZTimerOn` and `ZTimerOf`! By converting the far calls
|
far calls to `ZTimerOn` and `ZTimerOf`f! By converting the far calls
|
||||||
to push/near call pairs within the Zen timer module, TASM makes it
|
to push/near call pairs within the Zen timer module, TASM makes it
|
||||||
impossible to emulate exactly the overhead of the Zen timer, and makes
|
impossible to emulate exactly the overhead of the Zen timer, and makes
|
||||||
timings slightly (about 16 cycles on a 386) less accurate.
|
timings slightly (about 16 cycles on a 386) less accurate.
|
||||||
|
|
@ -2255,7 +2255,7 @@ processor cache at the start of the code being timed, because the timing
|
||||||
code is not necessarily fetched and does not necessarily access memory
|
code is not necessarily fetched and does not necessarily access memory
|
||||||
in exactly the same time sequence as the code immediately preceding the
|
in exactly the same time sequence as the code immediately preceding the
|
||||||
code under measurement normally does. This prefetch effect can introduce
|
code under measurement normally does. This prefetch effect can introduce
|
||||||
as much as 3 to 4 µs of inaccuracy. Similarly, the state of the prefetch
|
as much as 3 to 4 µ of inaccuracy. Similarly, the state of the prefetch
|
||||||
queue at the end of the code being timed affects how long the code that
|
queue at the end of the code being timed affects how long the code that
|
||||||
stops the timer takes to execute. Consequently, the Zen timer tends to
|
stops the timer takes to execute. Consequently, the Zen timer tends to
|
||||||
be more accurate for longer code sequences, since the relative magnitude
|
be more accurate for longer code sequences, since the relative magnitude
|
||||||
|
|
|
||||||
|
|
@ -878,8 +878,8 @@ the PC must be completely refreshed about once every four milliseconds
|
||||||
in order to ensure the integrity of the data it stores. Obviously, it's
|
in order to ensure the integrity of the data it stores. Obviously, it's
|
||||||
highly desirable that the memory in the PC retain the correct data
|
highly desirable that the memory in the PC retain the correct data
|
||||||
indefinitely, so each DRAM chip in the PC *must* always be refreshed
|
indefinitely, so each DRAM chip in the PC *must* always be refreshed
|
||||||
within 4 ms of the last refresh. Since there's no guarantee that a given
|
within 4 µs of the last refresh. Since there's no guarantee that a given
|
||||||
program will access each and every DRAM block once every 4 ms, the PC
|
program will access each and every DRAM block once every 4 µs, the PC
|
||||||
contains special circuitry and programming for providing DRAM refresh.
|
contains special circuitry and programming for providing DRAM refresh.
|
||||||
|
|
||||||
#### How DRAM Refresh Works in the PC
|
#### How DRAM Refresh Works in the PC
|
||||||
|
|
@ -900,8 +900,8 @@ purpose of refreshing the DRAM; the data that is read isn't used.)
|
||||||
The 256 addresses accessed by the refresh DMA accesses are arranged so
|
The 256 addresses accessed by the refresh DMA accesses are arranged so
|
||||||
that taken together they properly refresh all the memory in the PC. By
|
that taken together they properly refresh all the memory in the PC. By
|
||||||
accessing one of the 256 addresses every 15.08 µs, all of the PC's DRAM
|
accessing one of the 256 addresses every 15.08 µs, all of the PC's DRAM
|
||||||
is refreshed in 256 x 15.08 µs, or 3.86 ms, which is just about the
|
is refreshed in 256 x 15.08 µs, or 3.86 µs, which is just about the
|
||||||
desired 4 ms time I mentioned earlier. (Only the first 640K of memory is
|
desired 4 µs time I mentioned earlier. (Only the first 640K of memory is
|
||||||
refreshed in the PC; video adapters and other adapters above 640K
|
refreshed in the PC; video adapters and other adapters above 640K
|
||||||
containing memory that requires refreshing must provide their own DRAM
|
containing memory that requires refreshing must provide their own DRAM
|
||||||
refresh in pre-AT systems.)
|
refresh in pre-AT systems.)
|
||||||
|
|
@ -1223,7 +1223,7 @@ display, and even with the display adapter cycle-eater it just doesn't
|
||||||
take that long to manipulate 4,000 bytes. Even if the display adapter
|
take that long to manipulate 4,000 bytes. Even if the display adapter
|
||||||
cycle-eater were to cause the 8088 to take as much as 5µs per display
|
cycle-eater were to cause the 8088 to take as much as 5µs per display
|
||||||
memory access—more than five times normal—it would still take only
|
memory access—more than five times normal—it would still take only
|
||||||
4,000x 2x 5µs, or 40 ms, to read and write every byte of display memory.
|
4,000x 2x 5µs, or 40 µs, to read and write every byte of display memory.
|
||||||
That's a lot of time as measured in 8088 cycles, but it's less than the
|
That's a lot of time as measured in 8088 cycles, but it's less than the
|
||||||
blink of an eye in human time, and video performance only matters in
|
blink of an eye in human time, and video performance only matters in
|
||||||
human time. After all, the whole point of drawing graphics is to convey
|
human time. After all, the whole point of drawing graphics is to convey
|
||||||
|
|
@ -1261,7 +1261,7 @@ seriously impact code performance, even as measured in human time.
|
||||||
|
|
||||||
For example, if we assume the same 5 µs per display memory access for
|
For example, if we assume the same 5 µs per display memory access for
|
||||||
the EGA's high-resolution graphics mode that we assumed for text mode,
|
the EGA's high-resolution graphics mode that we assumed for text mode,
|
||||||
it would take 26,000 x 2 x 5 µs, or 260 ms, to scroll the screen once in
|
it would take 26,000 x 2 x 5 µs, or 260 µs, to scroll the screen once in
|
||||||
the EGA's high-resolution graphics mode, mode 10H. That's more than
|
the EGA's high-resolution graphics mode, mode 10H. That's more than
|
||||||
one-quarter of a second—noticeable by human standards, an eternity by
|
one-quarter of a second—noticeable by human standards, an eternity by
|
||||||
computer standards.
|
computer standards.
|
||||||
|
|
|
||||||
|
|
@ -346,7 +346,7 @@ For example, you'd certainly expect a sequence such as
|
||||||
pop ax
|
pop ax
|
||||||
ret
|
ret
|
||||||
pop ax
|
pop ax
|
||||||
ret
|
et
|
||||||
:
|
:
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -107,7 +107,7 @@ from the use of DI to address memory (remember, the loop is unrolled, so
|
||||||
the last instruction is followed by the first instruction), but because
|
the last instruction is followed by the first instruction), but because
|
||||||
the intervening instruction takes two cycles, there's no penalty at all.
|
the intervening instruction takes two cycles, there's no penalty at all.
|
||||||
|
|
||||||

|

|
||||||
|
|
||||||
> 
|
> 
|
||||||
> Remember, pipeline penalties diminish with increasing number of cycles,
|
> Remember, pipeline penalties diminish with increasing number of cycles,
|
||||||
|
|
|
||||||
|
|
@ -302,7 +302,8 @@ the pike. The success or failure of the search can then be determined
|
||||||
outside the loop, if necessary, by checking for the tail node's special
|
outside the loop, if necessary, by checking for the tail node's special
|
||||||
pointer—but the inside of the loop is streamlined to just one test, as
|
pointer—but the inside of the loop is streamlined to just one test, as
|
||||||
shown in Listing 15.5. Not all linked lists lend themselves to
|
shown in Listing 15.5. Not all linked lists lend themselves to
|
||||||
sentinels, but the performance benefits are considerable
|
sentinels, but the performance benefits are considerable for those lend
|
||||||
|
themselves to sentinels, but the performance benefits are considerable
|
||||||
for those that do.
|
for those that do.
|
||||||
|
|
||||||

|

|
||||||
|
|
|
||||||
|
|
@ -237,7 +237,7 @@ contention. Such operations, as in
|
||||||
|
|
||||||
```nasm
|
```nasm
|
||||||
mov eax,edx ;U-pipe cycle 1
|
mov eax,edx ;U-pipe cycle 1
|
||||||
sub edx,edx ;V-pipe cycle 1
|
sub edx,edxX ;V-pipe cycle 1
|
||||||
```
|
```
|
||||||
|
|
||||||
are free of charge.
|
are free of charge.
|
||||||
|
|
|
||||||
|
|
@ -37,7 +37,7 @@ the way up to 360x480—and that's with the vanilla IBM VGA!
|
||||||
|
|
||||||
In this chapter, I'm going to focus on one of my favorite 256-color
|
In this chapter, I'm going to focus on one of my favorite 256-color
|
||||||
modes, which provides 320x400 resolution and two graphics pages and can
|
modes, which provides 320x400 resolution and two graphics pages and can
|
||||||
be set up with very little reprogramming of the VGA. In the next chapter, I'll
|
be set up with very little reof the VGA. In the next chapter, I'll
|
||||||
discuss higher-resolution 256-color modes, and starting in Chapter 47,
|
discuss higher-resolution 256-color modes, and starting in Chapter 47,
|
||||||
I'll cover the high-performance "Mode X" 256-color programming that many
|
I'll cover the high-performance "Mode X" 256-color programming that many
|
||||||
games use.
|
games use.
|
||||||
|
|
@ -341,7 +341,7 @@ LinesDone:
|
||||||
;
|
;
|
||||||
call GetNextKey
|
call GetNextKey
|
||||||
mov ax,0003h
|
mov ax,0003h
|
||||||
int 10h ;text mode
|
int 10h text mode
|
||||||
mov ah,4ch
|
mov ah,4ch
|
||||||
int 21h ;done
|
int 21h ;done
|
||||||
;
|
;
|
||||||
|
|
|
||||||
|
|
@ -447,7 +447,7 @@ DrawRectParms ends
|
||||||
mov dh,RightMask[bx] ;set the right-edge clip mask
|
mov dh,RightMask[bx] ;set the right-edge clip mask
|
||||||
mov bx,LeftX[bp]
|
mov bx,LeftX[bp]
|
||||||
and bx,NOT 7 ;intrapixel address of left edge
|
and bx,NOT 7 ;intrapixel address of left edge
|
||||||
sub si,bx
|
su si,bx
|
||||||
shr si,1
|
shr si,1
|
||||||
shr si,1
|
shr si,1
|
||||||
shr si,1 ;# of bytes across spanned by rectangle - 1
|
shr si,1 ;# of bytes across spanned by rectangle - 1
|
||||||
|
|
@ -455,7 +455,7 @@ DrawRectParms ends
|
||||||
and dl,dh ; combine the masks
|
and dl,dh ; combine the masks
|
||||||
MasksSet:
|
MasksSet:
|
||||||
mov bx,BottomY[bp]
|
mov bx,BottomY[bp]
|
||||||
sub bx,TopY[bp] ;# of scan lines to fill - 1
|
su bx,TopY[bp] ;# of scan lines to fill - 1
|
||||||
FillLoop:
|
FillLoop:
|
||||||
push di ;remember line start offset
|
push di ;remember line start offset
|
||||||
mov al,dl ;left edge clip mask
|
mov al,dl ;left edge clip mask
|
||||||
|
|
@ -661,7 +661,7 @@ TextUpDone:
|
||||||
CharUp: ;draws the character in AL at ES:DI
|
CharUp: ;draws the character in AL at ES:DI
|
||||||
lds si,[BIOS8x8Ptr] ;point to the 8x8 font start
|
lds si,[BIOS8x8Ptr] ;point to the 8x8 font start
|
||||||
mov bl,al
|
mov bl,al
|
||||||
sub bh,bh
|
su bh,bh
|
||||||
shl bx,1
|
shl bx,1
|
||||||
shl bx,1
|
shl bx,1
|
||||||
shl bx,1 ;*8 to look up character offset in font
|
shl bx,1 ;*8 to look up character offset in font
|
||||||
|
|
|
||||||
|
|
@ -499,7 +499,7 @@ parms ends
|
||||||
les di,[bp+BufferPtr]
|
les di,[bp+BufferPtr]
|
||||||
mov dx,[bp+RectHeight]
|
mov dx,[bp+RectHeight]
|
||||||
mov bx,[bp+BufferWidth]
|
mov bx,[bp+BufferWidth]
|
||||||
sub bx,[bp+RectWidth] ;distance from end of one dest scan
|
su bx,[bp+RectWidth] ;distance from end of one dest scan
|
||||||
; to start of next
|
; to start of next
|
||||||
mov al,byte ptr [bp+Color]
|
mov al,byte ptr [bp+Color]
|
||||||
mov ah,al ;double the color for REP STOSW
|
mov ah,al ;double the color for REP STOSW
|
||||||
|
|
@ -544,7 +544,7 @@ parms2 ends
|
||||||
mov bx,[bp+Pixels]
|
mov bx,[bp+Pixels]
|
||||||
mov dx,[bp+ImageHeight]
|
mov dx,[bp+ImageHeight]
|
||||||
mov ax,[bp+BufferWidth2]
|
mov ax,[bp+BufferWidth2]
|
||||||
sub ax,[bp+ImageWidth] ;distance from end of one dest scan
|
su ax,[bp+ImageWidth] ;distance from end of one dest scan
|
||||||
mov [bp+BufferWidth2],ax ; to start of next
|
mov [bp+BufferWidth2],ax ; to start of next
|
||||||
RowLoop2:
|
RowLoop2:
|
||||||
mov cx,[bp+ImageWidth]
|
mov cx,[bp+ImageWidth]
|
||||||
|
|
@ -556,7 +556,7 @@ ColumnLoop:
|
||||||
mov es:[di],al
|
mov es:[di],al
|
||||||
SkipPixel:
|
SkipPixel:
|
||||||
inc bx ;point to next source pixel
|
inc bx ;point to next source pixel
|
||||||
inc di ;point to next dest pixel
|
inc d ;point to next dest pixel
|
||||||
dec cx
|
dec cx
|
||||||
jnz ColumnLoop
|
jnz ColumnLoop
|
||||||
add di,[bp+BufferWidth2] ;point to next scan to fill
|
add di,[bp+BufferWidth2] ;point to next scan to fill
|
||||||
|
|
@ -596,9 +596,9 @@ parms3 ends
|
||||||
lds si,[bp+SrcBufferPtr]
|
lds si,[bp+SrcBufferPtr]
|
||||||
mov dx,[bp+CopyHeight]
|
mov dx,[bp+CopyHeight]
|
||||||
mov bx,[bp+DestBufferWidth] ;distance from end of one dest scan
|
mov bx,[bp+DestBufferWidth] ;distance from end of one dest scan
|
||||||
sub bx,[bp+CopyWidth] ; of copy to the next
|
su bx,[bp+CopyWidth] ; of copy to the next
|
||||||
mov ax,[bp+SrcBufferWidth] ;distance from end of one source scan
|
mov ax,[bp+SrcBufferWidth] ;distance from end of one source scan
|
||||||
sub ax,[bp+CopyWidth] ; of copy to the next
|
su ax,[bp+CopyWidth] ; of copy to the next
|
||||||
RowLoop3:
|
RowLoop3:
|
||||||
mov cx,[bp+CopyWidth] ;# of bytes to copy
|
mov cx,[bp+CopyWidth] ;# of bytes to copy
|
||||||
shr cx,1
|
shr cx,1
|
||||||
|
|
|
||||||
|
|
@ -547,7 +547,7 @@ void WalkTree(NODE *pNode)
|
||||||
// Pop the next node from the stack so
|
// Pop the next node from the stack so
|
||||||
// we can visit it and see if it has a
|
// we can visit it and see if it has a
|
||||||
// right subtree to be traversed
|
// right subtree to be traversed
|
||||||
if ((pNode = *--pNodeStack) == NULL)
|
if ((pNode = *—pNodeStack) == NULL)
|
||||||
{
|
{
|
||||||
// Stack is empty and the current node
|
// Stack is empty and the current node
|
||||||
// has no right child; we're done
|
// has no right child; we're done
|
||||||
|
|
|
||||||
|
|
@ -1,5 +1,7 @@
|
||||||
# About this version
|
# About this version
|
||||||
|
|
||||||
|
All rights belong to Michael Abrash. Reproduced with permission.
|
||||||
|
|
||||||
This version was extracted from the PDFs which were [released by Michael Abrash and Dr. Dobbs in 2001](http://www.drdobbs.com/parallel/graphics-programming-black-book/184404919). The intention is to maintain a canonical electronic version of the book, and make it easier to read in other formats and on other devices than were available when the book was released online.
|
This version was extracted from the PDFs which were [released by Michael Abrash and Dr. Dobbs in 2001](http://www.drdobbs.com/parallel/graphics-programming-black-book/184404919). The intention is to maintain a canonical electronic version of the book, and make it easier to read in other formats and on other devices than were available when the book was released online.
|
||||||
|
|
||||||
For comments, suggestions, and improvements contact James Gregory at [james@jagregory.com](mailto:james@jagregory.com).
|
For comments, suggestions, and improvements contact James Gregory at [james@jagregory.com](mailto:james@jagregory.com).
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue