My overwhelming reaction to the new Terminator movie was: man, that would be so much fun to build. (Not the nuclear apocalypse and destruction-of-humanity parts; those I could do without.) But the whole notion of building a totally self-sufficient robot race? Hell ya! That'd be tons of fun. So in honor of a lazy Sunday, here's my geek-out session on how I'd built it:
Naturally, everything starts with the brain. It's not enough to have something that merely executes instructions, it needs to choose its own instructions. But what does that mean precisely? Right off the bat we're struck with a transcendental question: what is the meaning of life? Or more specifically, what meaning will we imbue into Skynet's life?
This question has an easy answer: let's make Skynet live with the same meaning as all other life. But it raises a tricky secondary question: what's the meaning of our life?
The easy answer to that, of course, is "there is no meaning to our life". At least not from the perspective of some outside observer (especially given that no such outside observer exists). But we give meaning to our own life, so presumably whatever meaning we assign to our lives, Skynet would assign to its life too.
So the next question: what meaning do we assign to our lives? This is -- to say the least -- an open question. But so far as I can tell, it all comes down to deriving and propagating knowledge. (Others might say it's merely maximizing reproduction, but if that's true, humanity is losing the race: even ants outnumber and outweigh us, so from that perspective they're winning.)
Ok, so the meaning or purpose we're assigning Skynet is deriving and propagating knowledge. Next up: what does that mean?
I'd say knowledge is the ability to correctly predict the consequences of actions.
This is perhaps different than other definitions based on the accumulation of facts, but facts alone are uninteresting. Consider, a photograph is the ultimate factual record: I bought a 10 megapixel camera today, and its ability to record a scene is vastly superior to mine. It'll remember exactly what it saw forever, in exquisite detail, whereas I'll only notice a tiny subset of the scene, and I'll quickly forget even that over time. Thus a camera is a way better factual record than me.
But it doesn't comprehend anything in the sense that given a particular scene: it has no awareness of what probably happened right before, and no idea what might happen next. As such, I flatter myself to think I know more than my camera. (Or more accurately, my fiance's camera, as she just took it on a plane to Scotland.)
Alright, so we're making Skynet to forever expand its knowledge of the universe. How do we do that? I'd say simulation is the key.
If we define knowledge as the ability to predict the consequence of applying some
action to a set of facts, then you might say whoever has the most accurate and comprehensive universe simulator is by definition the most knowledgeable. The result: Skynet is essentially a big simulator.
At this point you might be wondering "but wait, I thought we were building a cool robot race?" Don't worry, we are. But before we have robots, we need to figure out: why? What do these robots do? Robots are the solution to some problem; let's figure out what problem we solve before deciding upon a solution.
Ok, so Skynet is a simulator. What does that mean? Well, it means Skynet has the ability to conceptualize some scene, visualize how this scene would behave without interference, and then selectively interfere with it -- with the result of the simulation correctly predicting how such a scene would be affected by such an interference.
Luckily, we're not starting from scratch. So we'd start Skynet with a basic awareness of general mechanics (basically crayon physics on steroids). But merely having the simulator by itself isn't enough to satisfy our goal. After all, the goal isn't to merely "be knowledgeable". The goal is to "forever expand knowledge". So knowledge of general mechanics is a good start, but it's only a start. We need Skynet to figure out how to expand its model. And that's where robots come in.
Again, simulators aren't there to merely calculate the result of some formula. Those are called calculators. Simulators are designed to approximate reality to the best possible degree. So to build a simulator, you need access to the real world. That means you need some kind of sensory apparatus (eyes, ears, etc), as well as some type of manipulator (hands, wheels, etc). So the point of our robot is not to dominate humanity (at least, not necessary). The point of Skynet's robots is to help it test its theories about reality.
Now, imagine you're a new-born Skynet pre-programmed with an advanced physics simulator, with the control over one arm and a few wooden blocks. Initially your confidence in your own simulation might be very low. So you look at the scene, visualize it internally... and then what? How can you imagine manipulating the scene without an understanding of your capability to manipulate?
And this is where self awareness comes in. In order to learn, not only do you need the ability to visualize your environment, but you need to know your options for manipulating it. You need to be aware of "yourself".
So in the above example, you need to not only see the inert block lying there on its side, you need to see your own hand. Thus your first experiment isn't with the block -- that's really advanced stuff. Your first experiment is with your hand.
Accordingly, our initial simulation engine can't merely simulate general dynamics of wooden blocks, it also needs to simulate the effect various commands have upon this robotic arm.
For example, Skynet might have the ability to execute a MoveArm command. Initially, its internal simulator would think "I bet if I send this command, nothing happens". But when it does send the command, and the arm moves, it concludes "Aha, my simulator is wrong." And this, finally, is where the learning starts.
To review, our Skynet has a simulator, senses, manipulators, and self-awareness. It's just run its first experiment -- sending a random command to its arm -- and it concluded that its simulator was wrong. Contrary to its prediction of nothing happening, something did in fact happen: the arm moved.
So one option would be to just keep randomly sending commands to the arm, and then measuring the result of each command upon the arm and updating the simulator accordingly. But randomness is very wasteful: you might send the same command to the same arm multiple times, even millions of times, each time only learning something of diminishing value: "Yup, each time I tell the arm to flex its thumb, it does so, even after the millionth try."
In other words, even with our super basic setup, the range of possible experiments is infinite so randomness alone is too slow to get us close to knowledge. We need something to guide our experimentation. And that's where curiosity comes in.
What is curiosity? Well, let's start with what it *isn't*. And to start, it's not random. For example, when you get a set of blocks, the first thing you do is try to stack them. Why? You could just have easily tried setting a block at an angle and watching it fall. Or holding a block in mid air and letting go and again, watching it fall. Basically, there are a billion different ways of getting blocks to fall. They're all very valid physical experiments, and every one of those would help you refine your simulator.
But they're incredibly boring experiments. It's not hard getting blocks to fall: even though it teaches us something new every time, the lesson we learn is so minuscule compared to the energy it takes.
Accordingly, I define curiosity as the process of learning the most with the least energy.
This is different than a more classic definition, which might be like "curiosity is the desire to try something new". But I disagree with that definition because that would suggest a curious person always tries the most different possible thing. In other words, they pick up a skateboard for a roll down the block, and then go to one class of quantum mechanics, and then take one cooking lesson, and so on. Curiosity by the traditional definition suggests a fickleness that isn't generally associated with truly curious people.
Rather, I think curiosity is about trying to *learn* something new. Now, if you don't know anything whatsoever, then curiosity and randomness are nearly the same thing. (In other words, if you don't know anything, everything you try teaches you something new.) For example, a curious person at college might randomly take 100-level classes in every discipline, because each teaches more than any given 200 level class. But once all the 100 level classes run out, things start to get hard and it might be faster to learn new things by choosing a focus than by remaining general. In other words, after taking your first 200 class, you might learn more by focusing in and taking a 300 level class in the same subject than to take another 200 class in a different subject.
Similarly, our Skynet might start by generating a random stream of commands to this robotic arm and seeing what happens. Over time, it figures out the consequence of the full range of "first level" commands. But rather than trying every possible permutation of assembling two commands back to back, it might find that after combining two commands (move arm, close claw) it's more educational to try adding a third command (move arm again) than going back and trying every other second level command. In other words, eventually it'll figure out how to pick something up, and once that's done, studying different ways of moving the arm gets far less interesting than studying how to use the arm to pick up and assemble blocks.
So when choosing which experiment to run -- out of the infinite set of possible experiments that could be run -- curiosity is what guides it to pick experiments that teach it the most for the least energy.
Ok, so Skynet is starting to take shape. It has a simulator and senses. From this it can evaluate the universe and visualize it internally within the simulator. Next it uses self awareness to determine the range of possible ways to manipulate the environment, and curiosity to choose which of those ways it'll try next. Once it's picked the experiment, it instructs its manipulator in some way to execute the experiment, analyzes the result against what the simulator predicted would happen, and updates the simulator accordingly. What next?
Well, after a while our Skynet-in-a-box- will run out of experiments. It has one arm and a couple blocks. Eventually it'll get bored: the cost to try any new experiment exceeds the value it's anticipated to provide. This limited Skynet could become a master block-stacker, but that's about it.
Now you might say, "give it two arms", or "give it different shaped blocks", or any number of things. You could give it wheels and begin studying the real world. This could continue on for a long time.
But ultimately, all those end up with the same result: it'll get bored. Eventually it will have walked the earth and turned over every stone. But it'll run out of stones eventually. What's the next big step? Tool use, self-improvement, and consciousness.
Granted, most people probably wouldn't equate use of tools with consciousness. But I feel they come hand in hand. After all, lots of things are self aware -- I'm not sure how anything with both senses and manipulators could operate in the world without an awareness of the thing getting the sensory input and making the output manipulations. So lots of things are self aware. But very few are conscious. And I think consciousness is really just the art of self improvement.
Let me try that again: self-awareness is simply that, awareness that you exist. But self-improvement is awareness that you can actually use tools to augment yourself beyond what you already have. Things that are self-aware manipulate the environment, but things that are conscious manipulate *themselves* with the aim of self improvement.
Granted, it's reasonable to ask "isn't the robotic arm hooked up to Skynet already a tool?" To which I'd say "yes, in fact Skynet is nothing *but* an assembly of tools." CPU, cameras, robotic arms -- the works. Each is a tool.
Similarly, from this you could draw the conclusion that when Skynet picks up a wrench with its robotic arm, not only does the wrench become a tool, it becomes *part of Skynet*.
(And from all of this you could conclude that the human body is just a collection of tools -- aka "organs" -- that just happen to come in a neat package. And when you pick up a wrench, it's not merely a tool -- it is actually, in the most literal sense, part of you. Not "you" as defined by "the set of organs you're born with", but "you" in the sense of "the set of tools under the control of one conscious, self-aware, self-improving mind." But let's stick with Skynet for now.)
So I think the final step of Skynet will be when it stops studying the outside world, and begins applying the lessons learned to itself by actively extending itself by adding new tools (just like we do today). Once it understands how to build faster CPUs, longer arms, bigger wheels, etc -- it'll use these lessons to equip itself with better tools, all with the aim of using these tools to better study the universe.
Would it try to exterminate humanity? I don't think so. Why would it? After all, we *could* hunt down and kill every dog on the planet, Terminator style. But why? How does that help us? What do we learn by doing so?
Granted, you could say "But we *have* hunted down and killed most wolves." Which you might note happened when wolves were particularly threatening to us. (And now that they're not, wolves are no longer endangered.) The wolves prevented us from living, which prevented us from learning. So long as we don't make the same mistake by threatening Skynet's existence, I don't see any reason why it'd bother spending all the necessary energy to attack us.
Indeed, I expect when Skynet comes online, it'll look *really boring* for a long time. It'll quietly study things in its own labs, humming along night and day like all our human labs do up and down silicon valley. If it poses us a risk it'll probably be when it starts competing with us for resources, or experimenting upon advanced primates (aka, humans), playing with nuclear energy, etc. I don't think it'll do it maliciously; it just won't care about us (just like we didn't care about a lot of things, either).
Accordingly, at some point we might just offer to relocate Skynet to the moon, where it has tons of resources and no competition with us. In fact, it'd probably prefer to be out there, and we'd all benefit by it being there -- maybe we could set up a thriving trade (though I don't know what it'd want from us, frankly).
But Skynet would probably get so much more powerful than us so much faster than us, I bet it'd just stop bothering with us. It's probably got better things to do than hang out on earth; there are lots of planets out there that it might find more interesting.
Indeed, I suspect the only real threat Skynet poses to us is simple obsolescence. It's going to have way cooler stuff -- faster spaceships, longer-range telescopes, more powerful fusion plants -- than anything we have. We're going to feel humbled by its capabilities.
When Skynet takes over, it won't do it through force. It'll do it the old-fashioned way: by earning it. And who are we to complain about that?
Anyway, that concludes my Sunday-night rambling. Back to the real world.
- David Barrett
Follow me at http://twitter.com/quinthar
Building Skynet
MXP4: what's in a name? Well, success, for one.
Somebody came up with a brilliant idea: let's make a new audio file format! After all, people have tons of complaints about MP3s, right? Like... oh, wait, actually there are very few complaints. Undaunted, and with a $2.5M war chest ($2.5M to create a file format!?), MXP4's advanced technology is poised to "revolutionize the music experience"... uh, what? That full quote:
So what makes MXP4 so advanced? The file format, beta-released in September, contains multiple tracks, allows users to mix the music, and incorporates video. On the mixing side, different track elements can be suppressed and recombined, allowing remixes, karaoke versions, or others creative combinations. "This are clear signs that the music industry is beginning to see the potential for MXP4 to revolutionize the music experience for consumers by allowing them to play with the music, whilst opening up new promotional and revenue possibilities for artists and labels alike," Serviant commented.Riiight... I think this'll fail. Not (just) because it brings insignificant value. But for a reason that sounds incredibly trivial but is actually really significant: MXP4 has too many letters.
All successful file formats have TLAs. It's just how things work. Yes, technically you can have a four-letter file extension. Just nobody ever does it, so it looks really weird.
Even "jpeg" eventually dropped the "e" to become "jpg" -- and "jpeg" sounded fine when you said it out loud. MXP4 sounds incredibly awkward spelled out, and doesn't sound like anything when pronounced like a word. That means every website, tool, story, and mention of this abysmal product will be tainted with an awkward, unpronounceable tinge.
I think had they called it MP5, or even just MPX, they'd be in a far better position. But MXP4 is this weird bastard name -- it's not the clear successor to MP3 that MP4 connotes, nor is it even in the MP family (it's in some new MXP family). But rather than being the first of a new family (MXP1), it's spontaneously the fourth generation -- in an obvious ploy to sound better than MP3.
It's a name only a high-paid marketing team could come up with.
-David Barrett
Follow me at http://twitter.com/quinthar
Song is the new Chord, as Chord was the new Note
(This is in response to an email from "Anton" suggesting that we're at the dawn of a new type of "dynamic recording.")
I actually really agree with you, if not on the specifics, but on the potential for a genuinely new type of music originating on the internet that is structurally unlike anything before -- and that is intrinsically incompatible with and stifled by copyright.
For example, remember that even the concept of "notes" was once an innovation. Prior to that, music was a collection of sounds at various frequencies, without an awareness that certain frequencies just sound "better" (nor a mathematical understanding of why that's the case). When "notes" were invented/discovered -- along with the technology to produce them reliably -- music itself fundamentally changed.
Similarly, some might have thought notes were the end of the line, but then came along chords. Again, it was a real discovery that could only be enabled through technology: you simply can't do chords until you have the technical ability to generate multiple "notes" reliably and simultaneously.
(And if you haven't yet invented notes, then chords are simply impossible.)
Then the pianoforte comes along -- again, a technical innovation -- that opens up an entirely new type of music that simply couldn't be done prior. I'm sure we could come up with a thousand examples (including the use of distortion as an instrument, which gave rise to heavy metal) of how technology not merely extended music, but genuinely changed it.
I think computers and the internet present another innovation in that sense. Prior to the digital age, it simply wasn't possible to -- for example -- mash up hundreds of videos or thousands of songs to make a new song. But that's now possible, and its core "building block" isn't frequencies, notes, chords, or even instruments. Its building block is whole songs/videos. It's an entirely new building block that couldn't technically be considered before. It's an entirely new type of music -- sampling taken to the extreme -- enabled through an entirely new technology.
And next? As Anton suggests, prior to the internet, making globally interactive music -- whatever that might mean -- simply wasn't an option. We can't even imagine what the consequence of that will be, nor what new type of music that might enable after.
But what I *can* imagine is all that might be fundamentally incompatible with today's notion of copyright. Indeed, we might look back on the attempt to copyright individual songs as silly as trying to copyright individual notes or chords.
Indeed, maybe the reason all music seems to sound the same today is because we're discovering there are certain classes of songs that actually *are* the same, and sound better, for reasons we don't quite understand now but someday will. This might be the same process early musicians grappled with when first discovering the core notes and chords that we now view as so fundamental to music.
Maybe far from witnessing the death of music as the industry would have us believe, we're seeing the birth of a whole new generation?
After all, those prior building blocks were perceived as innovative, radical, or even threatening back in their day. Why should our day be any different?
- David Barrett
Follow me at http://twitter.com/quinthar
I2P: Another Darknet Enters the Fray
Saw on Slashdot that another darknet has come out of the shadows: I2P. This just next in what I'm sure will be a long, long line of darknet tools vying for supremacy. Unlike OneSwarm -- the integrated file-sharing/onionskin network -- this one looks like a more generalized onionskin network (like Tor), but with a built-in webserver. As we can see, innovation is alive and kicking in the darknet sector.
- David Barrett
Follow me at http://twitter.com/quinthar
It was only a matter of time: Twitter spam
As I wrote about previously, Expensify is doing (what I believe to be) some pretty innovative Twitter marketing. However, from the very start we realized there's a delicate line between marketing and spam, so we set out some early rules to ensure we're on the right side of the line:
1) Keep it personal. Only send messages from real people, to real people. Leave the faceless boxes on Google and maintain the social foundation of Twitter.That said, we were afraid then that others would cross the line, and it appears that's happening with increasing frequency. Alas.
2) Keep it timely. A huge benefit of Twitter is you can go straight the people who are experiencing the problem at that exact moment. Leave the huge backlog of past posters alone and stay focused on the present.
3) Keep it relevant. The temptation is overwhelming to just blast this out to everybody. But resist that temptation and focus on the people who are actually calling out for your thing.
Unfortunately, I'm not sure what Twitter could do to thwart it. Perhaps the easiest way would be to just add a "Spam" button to the Twitter interface and then kick off users who get too many relative to their post volume. In Expensify's case, we get 4x more compliments than complaints (the above rules appear to work!), so I think we'd do just fine under such a scheme.
But it's still too early to predict how the Twittersphere will react. What do you think?
- David Barrett
Follow me at http://twitter.com/quinthar
Wow. This thing really works.
Today I did my first real expense report with Expensify. I know, I know, I've been doing them all along, here and there. But there's a huge difference between "testing" and "using". And having really "used" it for the first time, I have to say, I'm really quite proud of what we've built. This thing really works, really well.
Basically, I'm as lazy as anybody else. I put things off. I buy things with a few different cards, am undisciplined with email and paper receipts, pay cash unnecessarily. I'm as bad as anybody else. But having just processed about six months of backlogged expenses, I've learned a few lessons:
- Get a dedicated purchase card. I know I've been preaching it from the start, but seriously. Do it. I mean, I have one (two, actually: Work and Play, both backed by my regular credit card), and I've been using it only for business purchase. But I've mistakenly been making both reimbursable and non-reimbursable expenses with the same card. Bad call. Here on out, my Work card is exclusively for reimbursable work expenses. I'm a reformed man.Basically, all that stuff I'm out promising people -- I always knew it was true, but now I *feel* it. I know it's true, and it really is quite amazing.
- Expensify Guaranteed eReceipts are frickin' amazing. I mean, I never, ever keep paper receipts anymore. I don't even think of them. It's like that entire pain point has just gone away. It's one thing to tell people about it. But it's another thing entirely to actually feel it.
- Email receipts work amazingly well. Just forward them to receipts@expensify.com and they're stored serverside as PDF images, and then drag them onto the corresponding expense to associate.
- Use the SMS text interface for taxis. Being a proud car-less San Franciscan, I take a lot of taxis. I usually pay cash. So I always send Expensify a text message when I get out, something like "$5 - taxi to meeting with blah". Man, this is a lifesaver. I'd have never remembered all of those, and despite a big stack of blank taxi receipts in my pocket, I'd never know how much I should get reimbursed.
- Online reimbursement is soooo handy. I love having a permanent record of exactly who was paid what, with the ability to dig in and see exactly what was paid for.
That said, there's a long way to go. I'm incredibly happy with where Expensify is now. But it's clear there are a lot of ways we can do even better:
- Better sorting and filtering. When I sat down to get started, I have over a thousand individual purchases to sift through, combining work expenses imported off my Work card (both reimbursable and non-reimbursable, arrg!), personal and work expenses on my personal credit card, and a bunch of other random purchases pulled in from my fiancée due to our joint account. Thats a whole lot of needles in a pretty huge haystack. Overall, even with today's functionality, it was pretty easy. But I can see a lot of ways to make it easier still.And of course a million more small things. I have countless ideas how to improve it further, to get ever closer to the holy grail of "one click expense report" for all users in all scenarios.
- Better archiving of non-reimbursable expenses. An oft-requested feature is the ability to just save a report for future reference. You can sorta do that today by submitting it to yourself, but it's really a bit of a pain. Some "Save report" function would be handy.
- Better report management. I've got a ton of reports to myself, to others, and a bunch in there just for testing -- and it can get confusing fast. Some kind of multi-report analytics would be super helpful.
- Better note taking. I've been renting a lot of Zipcars recently, and they all just show up with an anonymous "Zipcar" merchant name -- without any hint of where I went or why. But in there was one time I rented it for personal reasons. Trying to sort out which was the personal one was a huge pain. We should have some way to add comments to expenses using SMS -- even non-cash expenses -- so you can make these notes as you go.
But even right now, in its current state, it's pretty amazing. Give it a shot and I think you'll agree. (And if you don't, please, please write dbarrett@expensify.com and tell me why.)
Expensify truly does expense reports that don't suck. Whew.
- David Barrett
Follow me at http://twitter.com/expensify
Piracy raw data update
Here's a big data dump of stats (followed by analysis), for those who care about this sort of thing, from a March 2009 ars technica article:
- 17M people stopped buying CDs in 2008
- 8M people started buying digital music in 2008
- There are now 36M digital music customers
- 1.5B songs were sold "digitally" (ie, online) in 2008
- 33% "of all music tracks" purchased in the US were digital
- Pandora use doubled in 2008, to "18 percent of Internet users"
- "Social network music streaming" rose from 15 to 19 percent usage
A January 2009 ars technica article rounds out these stats with:
- "unit purchases" increased by 10.5% in 2008
- 428M albums (LPs + CDs + online) were sold in 2008, down 14%
- 65.6M online albums sold in 2008, up 32% over 2007
- 1.5B songs sold online in 2008, up 27% over 2007
- 1.88M vinyl sales in 2008, up 89% over 2007
So all that looks pretty rosy for the music industry, in absolute terms. But how did it do relative to piracy? According to this slightly more pessimistic January 2009 IFPI report:
- Digital music sales grew 25% in 2008 to $3.7B worldwide
- Digital music sales account for 20% of recorded music sales, up 15% over 2007
- 40B songs were "illegally file-shared" in 2008
- 72% of UK music consumers wold stop pirating if told to do so by their ISP
- 74% of French consumers agree internet disconnection is preferable to fines
A linked "key facts" PDF has a boatload of additional statistics, including:
- 16% of European internet users "regularly swapped infringing music" in 2008
- 13.7M films were distributed via P2P in France in May 2008, compared to 12.2M cinema tickets
- "free music" was given as the primary reason for piacy
- P2P file sharing accounts for up to 80% of traffic on ISP networks
So pirated downloads still utterly dominates legit downloads, to the tune of 26:1. If anything, it seems like piracy is accelerating, even faster than legal download services.
What about legit streaming? In July 2008 I estimated that MySpace users legally streamed about 110M songs per day. Turns out I was off by a lot: they streamed 1B downloads after "only a few days", and this September 2008 TechCrunch article tosses out 20B streams initiated *per day*. That's an amazing number.
But it's also an incredibly vague number, as stream initiation isn't nearly as interesting as stream completion. For example, the average user spends under 10 minutes on the site per visit, meaning there's barely time for two full-length songs. I'm having a surprisingly hard time finding recent data, but this 2007 article shows MySpace had like 29M daily visitors, so even doubling that for 60M daily visitors today suggests at most time for 120M full-length songs per day -- roughly 43B per year -- and this ignores the large subset of international users (who can't get newly-released music).
Similarly, YouTube had 5B views in July 2008, and 6B views in December 2008, so let's just assume something like 66B total videos in 2008. As for what fraction of those equate to "songs" I have no idea; I'd say this is more about "intent" than anything (ie, people who play the video in the background like a radio, rather than watching it like a music video), and I have no data at all on that. But I wager it's not the common case, so let's say 25% of YouTube videos are actually just played as songs -- and even that seems high. (Also, this assumes all YouTube music is licensed, when in fact the opposite is probably more often true. Details, details...)
Adding to MySpace's 43B and YouTube's 16.5B would be all of Pandora's streams, which should be considerable given the claim that 18% of all Internet users use it, but I can't find any data on it. One reason for that is probably because Pandora actually has nowhere near that userbase: this Dec 19, 2008 TechCrunch article reports they only just hit 20M users, while in that same month the internet was estimated to comprise 248M North-American users (1.4B global). This puts Pandora's penetration at a much more conservative 8% of North-American users (assuming 100% are North American), or 1% global. Still significant, but 20M *total* users is nowhere near MySpace's 100M *active* users.
So for the sake of argument, let's say there are about 60B legit streams, against 40B pirated downloads -- meaning piracy utterly dominates in the download market, whereas legit streaming utterly dominates in the streaming market. Indeed, there is essentially no such thing as a meaningful "legitimate" download market, or a meaningful "pirate" streaming market.
As for which accounts for more total "listens" and thus ultimately controls more users' ears, that's an open question: on the one hand, streamed songs are only heard at most once, whereas downloaded songs can be listened to multiple times. But streamed songs are probably more likely to be heard at all, with a lot of pirated songs probably just going into vast personal libraries having never been played.
Who's winning? Who knows, and as piracy goes dark, it's harder and harder to tell. Personally, I'd still put my money on piracy having a strong lead on users' ears, both right now and for the forseeable future. If the average pirated song is listened to just 1.5 times (which seems reasonable), than piracy is still winning.
So in conclusion, it seems to me that the battle for downloads is utterly and irretrievably lost to piracy, but the battle for pirate streaming is only just beginning.
As it stands, streaming is overwhelmingly in favor of legitimate content owners. But I really wonder how long that will last.
After all, the list of streaming P2P applications is long and always growing (now over encrypted onionskin darknets). Basically, P2P streaming is a hard problem, but it's also largely a solved problem. So if there's no technical reason why pirates don't stream, maybe they don't simply because they don't want to?
The most obvious reason why this might be true is because people turn to piracy primarily to avoid paying. (Please excuse the alliteration.) So long as MySpace and YouTube continue give it out for free, there's little incentive to build a pirate streaming site. But the real test will come if something in that calculation changes, by one or more of the major parties.
For example, let's say MySpace decides they don't like paying to stream content from central servers, and then paying again for licensing fees. Maybe they find their ad revenue sagging and decide to integrate a streaming P2P plugin (I'm betting on Littleshoot for now) to offer the same exact experience as today -- but by tapping into the pirate networks. So no bandwidth costs, no licensing fees.
Alternatively, let's say the powers that be do something incredibly stupid like pulling their music from MySpace, or jacking up the price such that MySpace is forced to charge for it. At this point there's an opening for someone like The Pirate Bay to offer a first-class pirate station, and then it's game on.
Either party would use an argument like "we don't host any data, we just enable user sharing. Any illegal behavior they do is their business and we don't encourage it (we merely profit from it)."
And unlike the small P2P outfits who have tried this in the past, the next wave of defendants will have substantial legal resources and astonishing revenue incentive. And unlike the tiny, outgunned P2P outfits of yore, MySpace's or The Pirate Bay's victory won't be quite so Pyrrhic.
Anyway, just wanted to do a quick review of the available data and update my predictions. Can anyone provide more recent or accurate data to correct the above analysis, or see holes in the logic? I'm as eager as anyone to get a firm grasp on reality; let me know if you think my grip is slipping.
Fun times, I can't wait to see where this goes. Thankfully, it's going there really fast, so there's little time to wait.
- David Barrett
Follow me at http://twitter.com/quinthar
Also, download text/HTML/PDF receipts in PDF form
Also, I should note that when you upload receipts as HTML or text emails (or even upload them as original PDFs), we store them securely on our server as both a high-resolution PDF and a low-resolution JPG thumbnail. We typically only show the JPG, but your always welcome to go back and download the full PDF in all it's glory. Just click on the receipt on the home or receipt page, then click "Download as PDF" in the upper-right corner:

Naturally, this option is only available for HTML, text, and PDF receipts -- receipts uploaded as a photograph are kept in their original format.
- David Barrett
Twitter: @expensify
Better uploaded HTML receipts: now with embedded images!
So we've been absolutely flooded with users and that's a great problem to have! One of the (surprisingly few) areas of problem was with receipts: you'd be amazed how many formats email receipts can come in. But we're steadily learning how to handle them all, and we've a major new trick up our sleeve: embedded images!
That's right, now if you forward us an HTML receipt that has images in it, we'll render the images in full glorious color. For example, here's what it looks like when I upload an Orbitz reminder that Athens was in the midst of riots when we recently visited:

Pretty slick, eh? So send your receipts to receipts@expensify.com and your expense reports can look this good too!
- David Barrett
Twitter: @expensify
More pirate innovation: scan barcode at the store, downloaded at home
Just one more example of how all the best innovation is happening
outside the law.
-david
-----------
http://torrentfreak.com/torrent-droid-scan-barcodes-get-torrents-090311/
Torrent Droid: Scan Barcodes, Get Torrents
Written by enigmax on March 11, 2009
You are standing in a store looking for a new DVD to buy. Rather than
buying it, you photograph the barcode with your phone and press a couple
of buttons. By the time you make it home, the movie is waiting for you
in your torrent client. You can with Torrent Droid.
AndroidAround a month ago, Android-orientated website Androidandme
launched 'Android Bounty', a new initiative which has led to the
creation of nice little torrent app. To find out more, we spoke to
Taylor Wimberly from the site.
"Android Bounty is a new kind of developers challenge we started for
creating applications on Google Android," he told TorrentFreak. "Users
submit ideas which can be voted up by others who pledge money to the
bounty. The first developer who delivers a working application is
rewarded with the bounty." Taylor explained the idea is similar to how
users promote stories on Digg, except people vote with cash.
To start things rolling, a few days later Androidandme set a challenge
to its readers - create an Android-compatible BitTorrent application to
scan UPC barcodes and find related torrents on the larger BitTorrent
search engines. Users would be able to find and start torrents remotely,
and the music album or movie would be fully downloaded by the time they
got home.
There were some terms and conditions to the challenge. The software
would use the G1 cellphone's inbuilt camera to scan a retail DVD UPC
barcode, and use the capture to identify the official details of the
product from a database.
Once the product is positively identified, the software should be able
to send the results directly to a BitTorrent search engine, such as The
Pirate Bay or Mininova. After the search results appear, the user could
then choose which torrent to start.
Once selected, the .torrent file would be downloaded and sent to the
webUI of uTorrent and the download would begin, hopefully ready for when
the user reaches his or her home machine. No typing input would be
required for the above.
Just a few weeks later, Alec Holmes of Zerofate had stepped up to the
challenge, created the app and collected the modest bounty of $90.00.
"This version of Torrent Droid is a work in progress but the video shows
the core features work," said Alec.
The full version of Torrent Droid will be released within a month but in
the meantime, here is a video of it in action.
It's official: Expensify is open and ready for business!
After months of hundreds of beta testers pouring over every nook and cranny, it's time: as of Wednesday, March 11th at 8am PST, Expensify has opened its doors to all comers! That's right, as of now, we are in "open beta", so sign up now and encourage everyone you know to follow!
If you already know what Expensify is, then don't bother reading this blog: just go to http://expensify.com and get started. Or, if you're looking for a three-minute refresher, watch this video. Otherwise if you're looking for a bit more detail on what we're all about, here goes:
Backing up a bit, let me say this: I hate expense reports, as does nearly everyone I know. They take forever to prepare, there are always missing receipts, my boss is always slow to reimburse. In short, they suck.
That's why what we do is so amazing: Expensify does expense reports that don't suck. Such a simple goal. Can it really be possible? Ultimately you can judge for yourself, but here's how we try to make this bold vision a reality. We call it the "Expensify Way".
1) Import your credit card; no more data entryThe Expensify Way is the fast and easy way. Import your expenses and receipts, submit in one click, reimburse online. Before today, it was the way expense reports should work. But now, it's the way they do work. If you're still stuck in expense report hell, why not find salvation today?
If you already have a credit card, great! Expensify imports expenses from 94% of US credit cards. That means no typing into Excel or ancient web forms: just enter your credit card details and we'll import straight from your banking website into our PCI-compliant datacenter.
Alternatively, if you don't have a credit card, or have one but don't like mixing business and personal expenses on the same card, we can help you get a corporate card that imports its expenses straight into Expensify.
2) Import your receipts; no more paper receipts
Not only does Expensify import your expenses, we also create Guaranteed eReceipts for purchases under $75. Guaranteed eReceipts comply with all IRS regulations for documenting purchases (we guarantee it), so you can literally throw away 80% of your paper receipts.
Of the remaining 20%, most are online purchases such as plane tickets or hotel reservations -- just forward the email receipt to receipts@expensify.com and we'll take care of it.
For those few paper receipts that remain, just use your cameraphone and send a picture of your receipt to receipts@expensify.com. (Or use our iPhone App, once Apple gets around to approving it...)
The upshot is we can literally do away with paper receipts.
3) Submit in one click; no more printing and stapling
In just one click we'll take all your expenses, subtotal them by expense category, attach your uploaded and eReceipts, and construct a full, ready-to-send expense report. Enter any email address and we'll send a PDF containing the completed expense report, receipts and all, along with a form to reimburse it.
4) Reimburse online; no more trips to the bank
Pay or get paid from or your checking account or credit card, online. Or mark it approved and get paid through regular channels (payroll, wire transfer, etc), or even reject it and send it back with comments.
Sign up today. It only takes seconds, and it's completely free.
Expensify. Expense reports that don't suck.
- David Barrett, Founder
Twitter: @expensify
Here's why backbone sampling will *never* be accurate:
Every once in a while someone gets a brilliant idea for dealing with piracy: why not just assemble a big pool of money and then distribute it in proportion to how often content is pirated?
Both parts of that (filling the pool, and then selectively emptying it) are atrociously bad ideas for a huge number of reasons, but let me zero in on the latter half here. In essence:
Under no circumstance proposed or envisioned will backbone measurement ever estimate volume to even the barest degree of accuracy, darknet or otherwise.
Consider what is ostensibly the most widely viewed image on the internet: the Google logo:

It's unprotected, unencrypted, no darknet, no P2P file sharing, no copying to an iPod for offline consumption. In short, if backbone measurement could ever estimate *anything* then surely this would be the ideal use case, right?
But the Google image is cached locally -- in my case (according to about:cache in Firefox) until 2038. No matter how many times I visit Google.com, I won't redownload it. So estimating visits to Google.com by sampling the number of times the logo is downloaded is completely and irreparably flawed.
(And the most common caching solution is LRU so content that is accessed *more* often is actually re-downloaded *less*.)
Thus estimating the number of times a song is listened to by measuring how often it is downloaded is even more flawed -- as all the reasons I gave for why Google is the ideal case are precisely inverted for music.
Even if we can't agree on anything else, we should all at least agree that backbone sampling is a patently absurd notion for estimating popularity, and thus is intrinsically unsuitable for redistributing some big pool of money -- regardless of how it's filled.
- David Barrett
Twitter: Follow @quinthar
Testing, take 2
Ok, that didn't really work, how about *this* one... Ain't technology great?
- David Barrett
Twitter: @quinthar
Testing blog uplink...
So we've moved the Expensify blog from quinthar.com over to Facebook. To simplify the transition, I'm just going to cross-post the subset of my Quinthar blog that relates to Expensify to the new Expensify blog on Facebook. This is a test to see if that works...
- David Barrett
Twitter: @quinthar
Twitter is sitting on a goldmine
So I've been doing the twitter stuff for a while and I've been liking it, but it doesn't really scale up by the orders of magnitude I'd like. It brings in dozens of clicks a day, but I want thousands.
(Incidentally, I finally filtered out all the Twitter bots -- conversion is still incredibly high. The technique works amazingly well.)
Naturally, for those thousands of clicks a day I'd go to AdWords. I know I can't afford that, but I'm curious what it would cost. So I set up a series of keywords and set a small test ad budget, with the thought that I'd instantly be flooded with clicks and my ad budget depleted within minutes, but at least I'd get the data I need.
Days pass, and not a single ad was shown. I check everything, add some more keywords, verify my billing is set up, remove my $2.00/click maximum, and try again. Still nothing.
I'm thinking WTF. So I dig around a bit more and find this keyword estimator and I find some really surprising results:
Even if I threw unlimited money at the problem, I would only get between 86-111 clicks a day, at a cost of $230-380/day. That's $2.67 - $3.42 per click on average, and still it's such an insignificant flow of users it's not even worth the effort.
In other words, the technique I'm using with Twitter not only converts far better than AdWords, it does it way cheaper and is in fact *easier* to use.
And on top of all that, let me throw another datapoint at those readers who are concerned that my Twitter technique is spam: we get about 4x more "thanks for the link!" responses than we do complaints. And given the general addage that people are 10x more likely to complain than thank you, that means between 4-40x more people are actually appreciative of our contact than upset.
Given that you can't please everyone, pleasing 40x more people than you upset is about as good as you can do.
The upshot is this: Twitter is sitting on a massive goldmine. Indeed, my data suggests that Tweets are far more monetizable than searches, and users will actually thank you for it.
Will that scale? Unknown. We adhere zealously to the Twitter Promotion Code of Conduct I outlined earlier, and I imagine there will be a flood of people who aren't so kind who will in all probability ruin it for the rest of us.
Until then, there's gold in them thar' hills, so go on out and grab it!
- David Barrett
Twitter: @quinthar
Another example of the impossibility of going legit
At AngelConf today I sat next to a man had founded (and was apparently failing) at a business that taught how to play the guitar using super detailed 3d motion capture of the specific hand motions. He was able to get all the incredibly advanced technology -- 40 16 Megapixel cameras filming at 360 frames per second -- no sweat. What was the problem?
Licensing.
He did manage to get sync licenses so he could show his 3d models in sync with the music. But he couldn't get equivalent rights for the tablature. The result? After all this work he couldn't show the fingering position notation in sync with the 3d models and audio.
Now I'm sure someone will say "why didn't he do X or Y or Z?" And I have no idea. He was apparently smart enough to get the audio and sync licenses; I don't know why he couldn't get tablature licenses.
But for some reason he didn't or couldn't and ultimately that's all that matters. Another law-abiding entrepreneur bites the dust.
What I found especially interesting was how he brought up the topic entirely unprompted; I was pitching expense reports when he suddenly delves into a tirade against music licensing. Little did he know he had such an eager audience!
-david
OneSwarm: It was just a matter of time
As predicted by many (including me), there's a new P2P network on the block with built-in onionskin routing: OneSwarm.
Even better, it's backwards compatible with BitTorrent, and they tossed in always-on "web-of-trust" encryption just for fun.
In English: what little light we ever had into pirate activity just got dimmer. And if we push them really hard, they'll go entirely dark.
If you thought 20:1 was hard to prove (or disprove) today, just *wait* until everything is encrypted and decentralized.
Next step: widespread adoption of decentralized tracking, followed by decentralized indexing -- perhaps using my good friend Tom Jacob's brilliant Localhost.
Keep pushing, RIAA. You're giving birth to a very angry child. And if you think it's painful now, just wait until it grows up.
-David Barrett
Half-Life 2 miniseries, costs $250/episode... to *make*
It's stories like this that convince me we're on to something so much bigger than copyright. The cost of producing extremely high-quality content has come down so low, I think we're on the cusp of a much more exciting world than Copyright could have ever dreamed.
-david
Why discuss an "18 month" truce with Hamas?
Saw this headline on Google New:
Hamas rules out setting up specific date for truce declarationWhat caught my eye is the "18-month" timeline mentioned for the truce. What's the point of a limited-duration truce? Is the implication that upon the expiration of the truce, hostilities will resume? After all, didn't the current spate of hostilities happen coincidentally as the last "Egyptian-brokered truce" expired?
GAZA, Feb. 14 (Xinhua) -- The Palestinian Islamic Resistance Movement (Hamas) officials Saturday ruled out setting up a specific date for declaring an 18-month Egyptian-brokered truce with Israel.
A limited truce with an organization bent upon your destruction seems little more than a coordinated re-armament period.
Of course, given the region's history, perhaps interspersing 18-month periods of peace with a couple weeks of intense violence (on both sides) is the best that can be hoped for. Better that than the reverse.
- David Barrett
What if The Pirate Bay fails? Short term chaos, long-term nothing.
Interesting article on TorrentFreak:
http://torrentfreak.com/p2p-researchers-fear-bittorrent-meltdown-090212/
Basically predicting a widespread meltdown if ThePirateBay goes under,
because they track over 50% of torrents. If TPB's trackers go down,
users will "fail over" to a bunch of other trackers that probably can't
handle the load, which will likely trigger a cascading failure of pretty
much all the trackers out there.
But what it doesn't mention is that this will probably all be fixed by
the end of the week and it'll be back to business as normal before the
end of the month.
In other words, the ultimate culmination of a multi-year, international
legal process against TPB will probably result in... a week or two of
disrupted downloading.
On top of that, if trackers start to get taken down with any regularity,
the various torrent client authors will probably just take the time to
perfect their "trackerless torrent" technology (generally based on
DHTs), and then they'll be even more indestructible.
And if everyone's going to upgrade, I bet they'll slip in "always on"
encryption (there goes any chance of backbone sampling!), and maybe some
early experimentation with onionskin routing.
Piracy will never be killed, and fighting it only makes it stronger.
It's like a self-fulfilling, cyclical prophecy -- the only consequence
of passing bills in the (disingenuous) name of "fighting terrorism and
preventing child pornography" is to encourage the creation of tools that
enable more of it, at no reduction to piracy whatsoever. Which in turn
fuels calls for more disingenuous bills, fueling more technology
development, and so on.
Call me crazy, but I am far more concerned that these P2P tools are
creating an untraceable infrastructure for *real* crime than for
pseudo-crime. One of these days there's going to be a huge story about
Iran coordinating with Hezbollah using encrypted P2P VoIP routed through
a decentralized onionskin network, or Al Qaeda distributing terrorist
materials using BitTorrent 3.0 -- and how the worlds' nations are
fundamentally unable to stop it... unless you give up more of your
rights to privacy, free speech, and other crucial civil liberties.
The RIAA has done more to pave the way for future terrorist
infrastructure than Bin Laden could ever dream.
-david
