Thursday, February 17, 2011

Python subprocess vs os.popen overhead

Let's say you're writing a Python program and you want to run an external command and read its output. The right thing to do used to be to:

 import os
 p = os.popen("command")
 output = p.read()

But there were a lot of ways to run programs depending on what kind of output you wanted to read (if any) what kind control you wanted (if any) and so on. Thus was born the subprocess module. In the current, 2.7 documentation for the os module, there's a note on popen:
Deprecated since version 2.6: This function is obsolete. Use the subprocess module. Check especially the Replacing Older Functions with the subprocess Module section.
Well, that's pretty definitive, right? Unfortunately, not so much.

Typically, process creation overhead doesn't matter a great deal. If a program needs to run another program, then the startup time involved in the creation of the process is probably an order of magnitude (often several) less than the time that the new program takes to do its work. So, you only typically care about process creation overhead when you're creating a very large number of parallel children.

Unfortunately, I work in the world of system monitoring, and in that world, creating a few hundred or thousand programs a second during peak times (subordinate monitoring tools) is not a rarity, and even when the load is much lighter, large amounts of process creation overhead isn't always ignorable. For example, if your system is doing a lot of IO, then large memory operations during process creation might reduce the amount of caching the system can do.

All of these factors lead me to test the subprocess module against os for a simple case: I want to run a process under a shell with standard output being captured. With the os module, my test looked like this:

 import os
 for i in range(10000):
    f = os.popen("exit 0")
    f.close()

With the subprocess module, my test looked like this:

 import subprocess
 for i in range(10000):
    f = subprocess.Popen("exit 0", shell=True)
    f.wait()

I timed the two scripts and came to a surprising conclusion: subprocess has about a 40% process creation overhead over os.popen! That's an awful lot of increase, so what could be going on? My next step was to use strace to determine what could be taking up that extra time. Here's a partial strace fo the subprocess example under Linux:

pipe([3, 4])                            = 0
fcntl(4, F_GETFD)                       = 0
fcntl(4, F_SETFD, FD_CLOEXEC)           = 0
clone(...) = 3306
close(4)                                = 0
mmap(NULL, 1052672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7f358acc6000
read(3, "", 1048576)                    = 0
mremap(0x7f358acc6000, 1052672, 4096, MREMAP_MAYMOVE) = 0x7f358acc6000
close(3)                                = 0
munmap(0x7f358acc6000, 4096)            = 0
wait4(3306, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, NULL) = 3306
The os.popen example had one less fcntl and an extra fstat, both of which are fairly light weight. The real culprit here are the mmap, mremap and munmap calls that subprocess is doing. Why are those there, I wondered. In looking at the subprocess code, it seems that these operations may be the result of thread creation which subprocess uses to manage reads and writes on subprocess inputs and outputs, but I'm not sure. What is clear is that the subprocess module is about 1,300 lines long while popen was a builtin supplied by the interpreter.

Conclusion

subprocess is a valiant attempt to make a complex snarl of library calls into a uniform tool. The problem is that process creation is one of the most fundamental operations that a language performs, and when a simple task like running an child process and reading its output becomes too heavy, a language suffers for it. Perhaps subprocess should be simplified and its convenience routines re-written as low-level operations that are optimized per-platform. Or perhaps os.popen should be undeprecated. After all, I'm willing to bet that managing fork, pipe and exec operations from Python will never be as low-impact as calling the C library popen(3) function.

Sunday, January 9, 2011

Evaluating the "readability" of programming languages

Being a Perl programmer (among many other languages), I've often entered into the debate over what "readability" is. People who don't program in Perl and many who do, find it difficult to read. The problem is that those who know the language typically don't mean the same thing at all when they say, "Perl is unreadable," as those who don't know the language. Perl isn't unique in having readability concerns directed at it. C, C++, Lisp, FORTRAN, Haskell, Ruby, PHP, Bourne Shell and friends, AWK, and dozens of others have had the same complaints leveled against them, and in many cases, the same dichotomy exists between those who know the language and those who don't.

So, if you don't program in a language, what makes it "readable?" Presumably, it "reads" as if it were a language you know. For example, if many of your operators are English words, and you seek to minimize the amount of punctuation in a language, then people who don't know the language might feel more comfortable with it than if you re-purpose ever bit of punctuation that one particular extended keyboard symbol set provides (yes, I'm looking at you, APL). Now, once you learn that symbol set, of course, you have an entirely different problem on your hands: can you understand the code that you can now map to a sequence of operations in your mind?

It turns out that this latter question is where many users of Perl over the years have staked out a more sophisticated set of concerns. For example, the language provides no fundamental object model, just the tools with which to build one. So, when you see code that uses objects, you have to wonder, "what is this doing?" When you have to ask that, then I think it's fair to criticize code as "unreadable."

On the other hand, I've had a debate with someone recently about this snippit of Perl 6:

 1, 1, *+* ... *

This is the infinite Fibonacci sequence in Perl 6, and while it's an obscure looking thing to those who don't work with Perl 6 on a regular basis, to someone who knows the language, there is no ambiguity at all. You never find yourself wondering what the hidden mechanics are, here, because there are none. It's simply a series of numbers, generated by adding the previous two values together to get the next value.

Perl 6 has its problems with readability, to be sure, but until we have a very large base of programmers using it on a regular basis, I don't know that we'll have a clear handle on what those are. I think the adverbial tagging of expressions will either end up being a boon or a substantial hindrance to readability, for example. The ability to build mini-languages ala Lisp is extremely powerful, but to say that it's open to levels of abuse that could make the International Obfuscated C Programming Contest look like a poetry slam is an understatement of epic proportions. Will this be a problem in reality? Maybe.

Moving past Perl, however, I'd love to have a generic set of metrics to apply to any language so that we can understand what it is that we're talking about. I see there being 5:

First, a definition: a "symbol" is any sequence of 1 or more characters that have a defined meaning in the language. In Perl, for example, "my", "+" and "die" are all symbols. The name of a function is not a symbol because its meaning is not part of the definition of the language.

Granularity - This is a measure of code abstraction. High granularity means that it takes many symbols to represent a given concept while low granularity means that a smaller number of symbols is required. One measure of granularity is the number of symbols in an average piece of code vs. cyclomatic complexity, but I don't believe that this is sufficient in modern programming languages.

Density - A simple average of the number of characters (in whatever character set you like) per symbol the language uses.

Vocabulary - The number of symbols in the language.

Textiness - The ratio of symbols which are comprised of "letters" (which may vary in meaning by character set) to those comprised of non-letters or a mix.

Morphability - The ease with which a programmer may change the behavior of the language or invoke behavior which is ambiguous at compile-time. This includes everything from textual macros to operator overloading to Lisp-style macros to simple polymorphism. Every way in which a piece of code may have multiple valid meanings or the meaning might change depending on the behavior of previously evaluated code. It's important to understand that "if a then b else c" has two behaviors that depend on the value of a, but if we could expect this code to compute pi because of something I've done ahead of time, then it would demonstrate a high morphability.

These five metrics need more specific definitions. We need to understand what their domains are and how different languages map into those domains. I'll tackle that in a latter article...

Monday, December 6, 2010

Wikileaks, Assange and Stross

Charles Stross has a great bit on his blog about Julian Assange. He is trending toward irrational exuberance, here, but I have to agree on the core point: anyone willing to publish these documents (Google for "Wikileaks" if you're not aware of the situation) is a kind of hero. It used to be that you went to the papers, but the papers are gasping for air and can't jeopardize their relationship with the government. Is there anything all that awful in the leaks? Not really, but it does make it harder, as Stross points out, for any group larger than 5-10 people to do anything they're going to be shocked or embarrassed by when it hits the Interwebs. This, to my thinking, is a good thing.

The counter-claim I've heard is that this makes it harder for foreign governments to trust us with their secrets (one example being given relates to Yemen cooperating over Al Qaeda raids, a reasonable concern). However, this is a reason to keep such secrets ... well, secret. Keeping them on file in a massive government bureaucracy isn't that, and the fact that someone then leaks that information is on the head of the person who does so, IMHO. Interestingly, we're not even talking about that person. We're just talking about the guy who did the actual publishing, which seems odd. He's just one guy with a Web site and no access to internal U.S. government information. Why not just prevent him from getting the information in the first place?

On a less wholesome note, I'll point out that Stross is a bit off on his analysis of the rape charge. While this definitely looks like a case of politically-motivated and kind of whacked-out revenge rather than a real rape claim, the charge being leveled against Assange isn't that he slept with another woman after the claimant, but that he had unprotected sex. This, apparently, under Swedish law can constitute rape even if the sex is consensual. I'm a bit shocked by this, but none the less, this is the claim I read on Wikipedia which is citing a Sydney Morning Herald piece about the rape charges. The part of the case that seems weird, however, is that the Sweedish authorities responded to Assange's willingness to meet with them at the Sweedish embassy or Scotland Yard with a request to Interpol and the EU for extradition. That's going way over the top, it would seem, given that their request was for interrogation, not trial.

Anyway, the wonderful thing about the Internet is: someone's going to pop up and offer the same service, even if Assange is buried for his role in this. It's awful to see him go through what even the women in question admit are rape charges over consensual sex, but in the end, I think his fears that he'll be handed over to the U.S. and harmed are unfounded... at least, I hope that will be the case. I want to think we haven't sunk that far...

Wednesday, October 27, 2010

The I Can't Take It Anymore Diet

Several years ago I was diagnosed with severe sleep apnea with a minimum blood oxygen saturation of around 70% which is to say, "you're dying because your brain is starving." I was placed on a CPAP machine and was immediately relieved of almost all symptoms. Problem solved, right? Well, no.

Life on a CPAP machine isn't fun. You have a bulky thing that you have to carry around anywhere you intend to sleep and anyone who intends to sleep with you has to deal with your having medical equipment strapped to your face and making airflow noises all night long. Mind you, it's much quieter than most apnea sufferers themselves, and far less distressing to hear than a pause in breathing followed by a loud snoring, choking intake of breath, but it's not exactly sexy.

Those were just cosmetic/lifestyle issues though, and I could have lived with that. The real problem was that I found the thing painful to wear and while it gave me a better night's sleep, it also caused me to get less of it, due to the difficulty of falling asleep and plenty of events where I would wake up with it pulled partially off.

So, I decided to give it up. That's not a small decision. I essentially decided to risk injury to my brain and/or death, so I needed a plan. Most apnea sufferers, myself included, have their breathing problems as a result of their weight. Over a certain, individual-specific weight threshold, your airway just doesn't work the way it was designed to. In my case, I was 265 lbs. and the threshold was about 230 lbs.

This made a diet a fairly obvious solution. (click the title for the rest of the story)

Wednesday, October 13, 2010

Something new every day: Bourne Shell variables

I've worked with the Unix Operating System and its variants since the late 1980s. I've worked with the Bourne Again Shell (bash) since the early 1990s. And yet today I learned something new about variable expansion. In the startup scripts for a source code indexing system called OpenGrok, I found this gem:
somecommand ${PROG:+-c} ${PROG}
Now, I know that ${FOO-bar} will be replaced with the value of $FOO if it is currently set or "bar" if it's not. That much I learned many years ago, but this usage of "+" was new to me. After some testing, I found that "+" substitutes the following text if and only if the variable is set, otherwise it substitutes nothing. Thus if $PROG were set to "foo", the above text would execute:
somecommand -c foo
But if $PROG were not set, then somecommand would be run with no arguments at all. Very slick!

How I managed to go over 20 years without learning that, I'm unsure (then again, perhaps I've learned and forgotten it...)

Tuesday, October 12, 2010

A Google App Engine failure

Long ago, I wrote a Perl script that generates random names for my roleplaying games. It's a simple thing, but it can take input lists from any language and spit out similar-sounding made-up names. It's a powerful, but simple tool, and it seemed a natural fit for my first exploration of Google App Engine. Sadly, it didn't work out that way, and I thought it might serve as a useful caution to others who might plan the same sort of work.

The fundamental problem is that my app is IO-hungry. It reads in the entire source list every time someone asks for a made-up word, crunches it down into first-parts, mid-parts and end-parts (2-3 letter segments which are rooted at the beginning or end of the word or neither). We then sort the lists of parts according to frequency of occurrence and perform a weighted, random pick of a first part, then each subsequent part is chosen in the same way, but from a subset of all of the parts, which overlaps the previous segment. The combination of weighted choice and overlapping leads to words which tend to be pronounceable in the source language of the input list.

This process of reading and processing all of the words every time wasn't something I was going to be able to do in Google App Engine, however, since costs are associated with resources consumption. So, I set out to store the pre-digested versions of the input lists as sorted word-segments in the Google App Engine datastore. This is where my problems began. While it's entirely possible to store the data this way, what I found was that my need to access so many records from the database as I performed my random walk down the lists of word-parts left GAA gasping for breath. In practical terms, I'd created the world's slowest tool for producing babble. Of this, I'm sure my mother feels proud.

Frankly, I'm not sure what I can do about this. GAA just doesn't seem to have been designed for this sort of thing. A shame, really. Of course, I could pre-compute a queue of results for each source namelist and keep re-populating them with a periodic job, but that really seems like a cheesy way to solve a problem that takes a few seconds for my original Perl script.

How not to open source a roleplaying game

After the phenomenal success of open source software development in the 1990s, someone at Wizards of the Coast decided to try to follow this model for their roleplaying game, Dungeons & Dragons, purchased in 1997 when they acquired TSR. Specifically, they decided to publish a cut-down set of rules called the SRD, which represented the core of what it takes to publish a Dungeons & Dragons-compatible game. The idea at the time, which worked well, was to encourage what marketing people call an "ecosystem" of publishers who made everything from full roleplaying game systems to source books for D&D.

All of this made sense, but Wizards had opened the genie's bottle, and there was no putting it back. As long as they continued to encourage their new ecosystem, they were assured of their place as king of the mountain. However, in 2008, they announced that they would end support for D&D version 3.5 and begin publishing version 4.0. This new version would not be available to third party publishers for creating their own games, at least not at first (they have since published a new, more restrictive set of rules called the GSL).

At the same time, a small publisher named Paizo had been licensed the rights to publish Dragon and Dungeon magazines. These two magazines were in decline at the time, but Paizo manged gain the enthusiasm of the roleplaying community by dipping into the well of nostalgia that many players had for the game. The printed updates to classic adventures in Dungeon and breathed new life into old stapes of Dragon such as the "Ecology of the..." series. They also fanned the flame of the original Dungeons & Dragons setting: Greyhawk. Their most impressive accomplishment was the enthusiasm generated by their "Adventure Paths." These 12-issue serial adventures were published in Dungeon magazine and the first was then collected as a hardcover book. Overall Paizo did quite a bit of cheerleading for Wizards and in 2007 they received their reward: notice that their licenses to all Dungeons & Dragons products were being revoked so that Wizards could develop a stand-alone Web site to replace the print magazines in coordination with the launch of the 4.0 edition.

Of course, the consensus at the time was that Paizo was doomed, but Wizards' Open Gaming License for the 3.5 edition was their way out. They immediately converted all existing Dungeon and Dragon subscribers over to new products based on the Open Gaming License. In less time than most publishers take to decide to take on a project, Paizo had their own campaign setting for D&D 3.5, published under the OGL along with a new adventure path, Rise of the Runelords. However, the OGL also allowed for full 3.5-compatible systems, and that was Paizo's next step. Over the course of the next two adventure paths that they published, Paizo continued to work on their variant system: the Pathfinder RPG.

Today, with several adventure paths published under their new system and a constant stream of supplement books published for their world of Golarion, one has to wonder if Wizards of the Coast is feeling burned. The popularity of the Pathfinder RPG hasn't reached the level of D&D, but it continues to build a dedicated base of adherents and is frequently referred to as "D&D 3.75."

So should Wizards not have created their "ecosystem?" Of course they should, it was a shot in the arm to what many declared a dead product. What they should not have done is assume that because they wanted to move on with a new edition that the industry would either follow or contentedly watch their businesses crumble. Wizards should have built support for 4.0 among their publishers and then eased into it without yanking the rug out from under Paizo. This would have resulted, at the very least, in reducing the number of long-term players that walked away from their version of the game.