Showing posts with label Data Portability (11 posts). Show all posts

July 15, 2011

Google's Data Liberation Front Frees Your +1s

Google's Data Liberation Front Frees Your +1s

In a new and innovative way to leverage the company's new Hangout feature as part of Google+, Google's Data Liberation Front held a small press briefing with a handful of tech journalists today, walking them through why the company is focused on leveraging open standards and helping users get their data out. Alongside the discussion, engineering manager Brian Fitzpatrick said the company has extended its exports to include the +1s you've made for Web sites - a small bump obviously, but one that demonstrates their seriousness about making getting your data out of Google as easy (if not easier) than it is to get it in.

Without speculating on other company's practices (namely Facebook), Fitzpatrick asked those of us participating if we would recommend a restaurant which locked its doors to prevent us from leaving after we sat down for a meal, or if we would recommend people rent an apartment that demanded it keep your furniture and family photos once you moved. The obvious answer is of course not... and Fitzpatrick said the same should be true for your online content. He recounted how when the Data Liberation Front first started with Blogger, there were some internal concerns at Google that users would leave the platform en masse for WordPress or other solutions, but in fact, they instead regularly downloaded content but kept posting - using the exports as local backup. (This is what I do as well)

Google Takeout Liberates Your Content from Multiple Services

Fitzpatrick pointed to Google Search as an example of not locking in one's data, as users, with many choices, will use the engine and can go to any other when they like. But with many online services, content goes in and doesn't come out. As I've demonstrated with my own personal backup of my Facebook wall and photos, you can get your data out of the social network, but it's possibly not configured very simply to move to another platform altogether. This is a main focus for the Data Liberation Front team, said Fitzpatrick, who said the effort is made to point to XML, Activity Streams and Microformats wherever possible, letting your data be interchangeable within services.

Downloading My +1s is a Mere 25KB.

Now that My +1s are Downloaded, I Can Move them Elsewhere

Lost in much of the coverage of Google+ a few weeks ago, the Data Liberation Front introduced its Google Takeout policy on the same day, making it easy to download data from multiple services and take them elsewhere. This list includes your +1s now, as well as Google Buzz, your Circles and Streams from Google+, and Picasa Web albums. As they mentioned at the time of that launch, if they make it simple for you to get your data out of the company, they have to work harder to keep you in.

June 4, 2011

A Scorched Data Policy Is Bad for Web, Bad for History

A Scorched Data Policy Is Bad for Web, Bad for History

In a world where the cost of storage is practically zero, and the incentive to delete data declines, I'm befuddled by the lack of prioritization for some companies to focus on a complete search history, and in parallel, an intentional erasure of information by more prominent content producers, who seem to arbitrarily decide that the long tail of the Web, and the interest of future readers is less important than being seen as participating in the latest hip wave.

Even as I've changed technologies, blog providers and structures, I make extra effort to not lose historical archives and keep comment streams intact, not for ego purposes or for SEO, but because it's the right thing to do.

Steve Rubel, an executive vice president of global strategy for Edelman, a longtime blogger who was among the first to espouse the benefits of blogging, social media and was often at the leading edge of tech only a few years ago, has more recently swayed to and fro based on the hot startup of the day, leaving thousands of broken links in the process. After running out of time to update his blog regularly, back in 2009 he ditched the blog to run a lifestream, based on Posterous. (Google cache). Last month, he pivoted again (another hot thing to do in the Valley these days for those with unsuccessful ideas) and now is the proud owner of a Tumblr-powered blog.

Pivot Number One Saw Steve Go to Posterous

Pivot Number Two Saw Steve Go to Tumblr, and Delete Everything

But instead of leaving the older sites open, he deleted everything, proudly stating:
"With just two clicks of a mouse I rid the web of literally thousands of blog posts, some of which I am proud of - others less so - and redirected the URLs to the new site."
While no doubt many of his older posts (like mine) have limited or no value to today's readers or those in the future, they are an insightful historical record of one of the more visible bloggers of a specific period, one who is now in the process of erasing his tracks.

With more than five years of blogging myself here, the blog archives now take on their own role as a personal reference desk. When Kevin Rose left Digg, I was able to go back to my comments on Digg in 2006 to see my thoughts at the time. When TweetDeck was sold to Twitter, I could go back to 2008 and see my first thoughts on the service. With PostRank selling to Google yesterday, I found a post I made two years ago that had comments on it from people who now work at Google. If I didn't manage to keep the posts alive and the comment threads as well, this would not be possible.

Steve's Community Is Very Unhappy With the Deletions, Calling it "Nutty"

Others Say Erasing "A Bit Much", Want Their Comments Out

For those of you who don't consume my blog exclusively by RSS, you might have noticed some recent changes to the look and feel (Go ahead and look). It's not a major change, but an upgrade nonetheless. Part of the reason for my slow migration was the criticality for me to not lose the existing posts, structure, external links and attached discussions. I know some comments from a few years ago are still out of reach, but I'm hoping to bring them back.

Of all the Web content I have produced or managed in the last fifteen years, one of my biggest regrets is the complete void from my time in college, both from my own personal home page, and from the student newspaper where I was the online editor and also wrote hundreds of posts, most on the front page. Current Valleywag editor Ryan Tate, who picked up the online editor job at the paper after I had left, struggled with a series of malicious hacks, and all our collective work was gone, erased from the Web like a bad memory. Whenever we trade emails or talk in person, we both lament this loss.

Steve's Original Micro Persuasion Blog, Pre-Deletion

This data obliteration is something that really is avoidable now, and yet, we let it happen on a near-constant basis. Most newspaper stories from the terrorist attacks in 2001 come up as 404s, beyond the reaches of the Archive.org project, or Google's search engine cache.

What I would like to see is a proposal from Google, or some other well-intended Web entity, such as Amazon, to offer a solution, embedded in today's modern browsers, as an option, that solves for intentional or unintentional content deletion. All those links that I provided back to Steve Rubel's MicroPersuasion blog from 2006 to 2009 should automatically be detected as dead, and then presented, to the best of the tech's ability, as they originally were, using Google cache or S3, Archive.org or something. And yes, I'd love it if somebody like Google or Microsoft would also give Twitter or Facebook a helping hand to get their own search archives into something useful.

Steve Thinks The World Won't Care About His Old Posts.

The issues I have with Steve's approach to pouring gasoline on his past and then lighting it on fire is not one of a choice of platforms. While seeing him join Tumblr is about as hip as your dad trying to snowboard with the cool kids, the more important part is that it eliminates the choice for readers, present and future, to ever get that data, and fill in the blanks. It's not his call to decide what has value for others, even if he sent me a tweet saying "we're foolish if we think the world really cares."

The world should care about walking through the historical record - be it on Steve's blog, or Dave's blog, or Robert's blog, or Mike's blog, or Penelope's, or any of the people who have been chronicling the world they see around them. I wish we had full archives from newspapers for years and decades backward, or the personal journals of people famous and ordinary from centuries past. What they might have found mundane is intriguing to others of us - maybe not massive populations, but to one person, they could contain serious insight. The Web is supposed to cater to the long tail, and the history of what we've produced should be there when they come looking for it.

May 19, 2011

Google Music Move: Days of File Transfers Ahead

Google Music Move: Days of File Transfers Ahead

Around noon on Wednesday, I fired up my old laptop, downloaded Google's Music Manager, and started the process of uploading my vast, but long-neglected, music collection to the cloud. Assuming the migration is complete, I'll have a Web-based repository for all the music I've purchased, as a companion to all the on-demand music I can find using Spotify - a match I like a lot. More than 30 hours later, thousands and thousands of my songs are on Google Music in beta. Yet, there are thousands more to go - highlighting a process that may be a challenge for users who may consider making a similar transition. And once they are uploaded, it's not clear if there's a way to get them down from the system if I want to move again.

My blogging colleague MG Siegler, he of ParisLemon and TechCrunch, argued yesterday that Google's move to open up a music locker (as Amazon did) without first gaining expensive deals with the music labels, has essentially locked up the market for Apple, who is quite unwilling to lose their leadership position in the market. He argued that the complexity of moving one's library to the cloud, instead of a simpler, faster way of matching your library to a known database and providing you access (as LaLa once did) would be a major stumbling block for the curious consumer. Putting MG's pro-Apple bias aside, there is some truth to that, if a known brand has a simpler solution and an official relationship with approved music downloads, the value to the cloud locker from a competitor is not obvious.

Downloading 1,945 high quality iTunes Plus tracks

As you know, migrating one's library out of iTunes and to a secondary location, be it Google Music or anywhere else, cannot happen for DRM-locked content. To strip one's iTunes library of the DRM, one must first pay Apple upwards of $.30 per track to upgrade to iTunes Plus. For a library like mine with thousands of tracks, this upgrade can (and did) cost hundreds of dollars - a steep pair of handcuffs, considering one doesn't get any new music in the deal. Converting to iTunes Plus doesn't actually upgrade the songs directly, but downloads brand new tracks in their place, and you have the option to delete the old files.

With a goal of completing this process, come Hell or high water, I gave Apple more than $340 yesterday to upgrade 1,945 tracks - at the same time as I had already told Google's Music Manager to upload my non-DRM tracks to the cloud. After about an hour, Apple sent me an email to say my tracks were ready, and I could begin the downloading process. So I set my old Mac aside, and put it to work, downloading the nearly 2,000 tracks at the same time as Google was uploading thousands more. Good thing I have unlimited bandwidth...

Google's Music Manager Has Uploaded 3,000+ Songs So Far

With my iTunes library now fully Plusified, Google Music has full go-ahead to get my music and that's happening. At this pace, it's likely the first pass of my full library will complete mid-day or late tomorrow, about 48 hours after I started. This process would take longer if it were on my primary machine, doing other tasks, or constantly interrupting the move by putting the machine through normal on/off paces.

Assuming the process completes, I am excited to have all my music in the cloud, searchable and playable from any computer where I am logged in. But as one person asked me, what if Google Music goes the way of Google Wave? I don't think it will, of course, but what if I want to take my more than 50 gigabytes of music down from the cloud, back to a desktop, or move from this service to another?

My Top Songs from iTunes Are Making it to Google Music

One admirable passion Google promises is the support of Data Portability. Its Data Portability Project is focused on letting users move data in services and out of services. You can see the conflicts over data ownership flare up especially in the battle with Facebook, who let users pull in Gmail contacts, but didn't let Google do the reverse. So I am a little surprised to see there isn't yet a way to bring data out of Google Music. Yes, it's brand new and early, but there's no talk, unless I missed it, about getting your music downloaded - be it one track or many. You can delete songs from the service. You can edit them and all their metadata, but you can't move them once they're uploaded. This means if you don't fully believe that Google Music will be around forever, you might still want to keep a second copy somewhere.

There is no doubt a considerable back story to the labels' fights with Google, Apple, Amazon and Spotify. I've heard people say there are hundreds of millions of dollars being promised up front and the biggest payer wins. Google said at Google IO that the expectations were prohibitive and they couldn't agree to such arduous terms. If Apple offers a comparable service on the cloud that knows one's iTunes purchasing history, and can eliminate the need to upload and store songs, but get the benefits of cloud, it will have a leg up in usability. I expect in coming months we will hear more from the Google Music team about improving migration and just how we can be sure we can always own our music.

November 19, 2010

Closed to Open: Downloading Your Data from Facebook

Closed to Open: Downloading Your Data from Facebook

Six weeks ago, Facebook announced they would enable users to download the entirety of personal profile on the network - a big move for the company often seen as being "closed" when other sites have attached themselves to phrases like "open" and "portable". The promise was that content you put into the system could be easily retrieved, from your own personal profile (or Wall) to your photos, videos and even the friend list - while that famously does not contain e-mails, and has been at the heart of a data tug of war between the company and its occasional nemesis, Google.

The data comes as a ZIP file which expands into simple HTML pages, with each major section of one's Facebook experience hosted as browsable directories. In fact, the data is easily expandable and movable that you could choose to replicate the Facebook profile experience on your own site, simply by uploading the content to a server under your control.

My Facebook Data Weighed In at 31 Megabytes

Eager to try this, earlier this week, I set out to get a local copy of the data I put into Facebook. Not a tremendously active Facebook user, I didn't imagine it would be a huge mass of data, but it did weigh in at a compressed 31 megabytes, complete with a pair of short videos, and 13 photo albums, including "Wall Photos", "Mobile Uploads" and "Profile Pictures".

The first thing to note upon downloading and expanding is that when Facebook says they let you download all your info, they aren't messing around. The index.html file shows the data you list on your wall, from your real birthday and spouse info to religion, political party, employer, and interests (such as all those Likes you have racked up). This you probably anticipated. But it also backs up all your messages, so you now run the risk of downloading entire e-mail message threads from friends and posting them live for all to see, making the private public.

Facebook Said It Would Take A While

I Had to Confirm It Was Me Before Downloading

If you do choose to upload this content somewhere outside of Facebook, the major heft of the content is on the "Wall", which is not paginated by length at all - so you can expect it to weigh in at several megabytes.

My Downloaded Facebook Profile

Mine hits more than 7, but you can embarrassingly scroll down to the bottom of my Facebook history, when I joined late July 20th, 2007, updating my first status to read "wondering why he's looking at this facebook stuff and considering getting out before it's too late...", following on the next day, adding, "Any activity I have here will be met with serious resistance on my part. I see Facebook more like the MySpace and Geocities of the future than a serious Web platform, especially as it's password-protected and outside the range of Google."

Looks like I was wrong on that one. Point Zuckerberg.

As one's Wall can be impressively long, it actually can be fun to search for keywords posted over the lifetime of one's Facebook use. You can also scan quickly and see what applications you may have used and abandoned.

In addition to Profile, Wall and Messages, there are top level directories for Photos, Videos, Friends, Notes and Events. Photos and Videos are self-explanatory, providing the option to back up your rich media, while friends just lists your friends in alphabetical order by first name. Your Notes are included and Events shows those Events you said you participated in.

In the interest of sharing, I downloaded my Facebook details and posted them on my own site. You can find my full experience at http://www.louisgray.com/facebook if you are curious. I removed Messages for obvious reasons, but the rest is there.

Of note, according to an interview with the LA Times, the project was the last one worked on by FriendFeed cofounder Paul Buchheit before he left the company to join Y! Combinator - and not the revamp of Messaging many people thought would be under his umbrella.

Getting one's data out of Facebook is a good thing. It doesn't mean they don't still have it, unless you were to close your account, but the idea that one's posted content and personal data is available to the author is the right concept.

July 4, 2010

A Day to Call for Data Independence

A Day to Call for Data Independence

I Want My Data Independent of Companies and Services

We don't live in a data democracy. Every day, we give over more of our data to people we don't know, whose motives we may not fully understand, and the possibility of our getting it back is very slim. Unlike in a democracy, we don't get to vote in the leaders of companies who build the programs and sites that harness and manipulate our data. We don't get to set term limits on CEOs or throw the bums out of corporate offices if our data is used in ways that are unsavory. And just like our tax dollars, we keep creating more and don't always know where our data is going to end up.

In today's world, as much of the data we create transitions from offline to online, and from our personal computer hard drives to the cloud, we are being forced to make choices in terms of what companies we trust, what services we believe will have a future, and what applications or Web sites can do what with our data. If we choose wrongly, we stand in danger of losing data, losing access to that data, losing historical data or metadata. Our data could be compromised by ill-meaning people, or simply cannot be moved, written once and made to lack true portability.

As I said in January, I think the time has come to deliver personal clouds with OS and application-neutral data. I don't want to be forced to accurately guess the right programs and providers, and I want anytime instant access to my own personal data, including contacts, relationships, rich media, e-mail, documents and more from any device.

We have not yet scratched the surface of data interoperability and portability between services and devices. While work on standards continues, there remain significant challenges, and often, as they do in government today, politics play a big role, between and often inside companies. Is what is in your best interest also in the best interest of those storing your content? And should they be one and the same?

I am calling for:
  • The ability to export one's social profile from network to network.
  • The ability to export one's shares and postings from network to network.
  • The ability to change blog software back-ends quickly, without losing content.
  • The ability to change blog comment providers quickly, without losing data.
  • The ability to change e-mail providers without long and tricky transfers.
  • The ability to purchase applications once even if you change operating systems on mobile or desktop.
  • The ability to access one's purchased media from any device if you provide your identity.
  • The ability to access one's bookmarks and history from any browser if you provide your identity.
  • The ability to access one's contacts and messages from any browser or e-mail client.
  • The ability to export one's social graph and reconnect it in a new place.
  • The ability to purchase media once even if you change media sources.
  • The ability to switch carriers without penalty.
  • The ability to collaborate with people across geographies, OS and Web browsers.
  • The ability to add actions to entries (such as likes or comments) and have them flow to all entry points.
  • The ability to migrate from one RSS reader or shared link blog and not lose subscribers.
  • The ability to download all videos and images in one's network rapidly.
These are just a few ideas off the top of my head, and the list could no doubt grow dozens long with your input and more hours spent trying to learn just where our data goes when we hit publish or submit. I am trusting an increasing amount of my data to the cloud with every post I write, every status update, every photo I upload, or every video I add to YouTube. With more than a decade and a half of activity on the Web, the legacy of my choices often impacts what I can do in the future, or the speed at which I can make change. I have had to walk away from created data before, and I have had to make compromises on choice or accept things as less than ideal much more often.

There are people out there working on very real standards to let our data move from place to place without hiccups - such as those hammering away on the DataPortability Project. Some of these issues are being solved in front of our eyes, while others are getting much worse. On a day when one nation is celebrating with big words like Freedom and Independence, it's worth knowing just who owns our data now, and striving to one day see a time when it can truly be moveable and independent - not beholden to any company, service, software or environment.

That would be a day worth celebrating.

March 13, 2010

Activity Streams Aim to Be DNA of the Future Web

Activity Streams Aim to Be DNA of the Future Web


When the well-respected open source advocate Chris Messina announced he was joining Google in January, many folks were concerned that his being absorbed in to the big company Borg would mean a cessation or redirection of some of his projects targeting the next generation Web, possibly in exchange of proprietary efforts to promote the company's products. Today, at the South by Southwest Interactive event in Austin, Texas, he spoke on how he and others in the community, both at Google and outside of it, are working to bring more meaning to our social networks, activity, and feeds, through extending today's data portability standards to include more information and more relevance. Messina walked through a history of the Web's publishing, from static portals of a decade ago, to today's RSS and Atom-powered sites, and suggested a future with even more information, based on streams, that tells a story.

Messina, after expressing his excitement about working with a team he believed was leading the industry in things he cares about, including the Data Liberation Front, letting data move from one site to another, said he was focused on what he called "generative structures" that were the underpinnings and DNA of how information is shared, updated and transmitted.

As he recounted, in 1999, portals ruled the Web, and people, myself included, would put their data in sites like My Netscape and My Excite, customizing these sites with headlines from third party services, primarily tapping RSS, which offered the headline of a story, a link and its description. In an era when publishers wanted to not give away their data, it was "the best we could do", he said. By 2005, a new extension of RSS was promoted, called Atom, which was still the essential concept of syndicating data from one Web site to another, but also adding an author and an identifier for the atomic bit of information.

Now in 2010, little has changed. Most news feeds of today, be they on Facebook, on customized portals, or the headline and link model that dominates Twitter, are fairly simple, and they don't indicate intent. As Messina said, "It's not all that different from the last 10 years, and that gets kind of depressing".

But what has changed is the increase of sharing rich media in these places, on platforms designed for "dead tree media". He said we should be able to show what we did, who with and why we were doing it, and that needs to happen through new richer formats for the social Web.

Activity Streams, an extension to the Atom Feed format, is looking to accomplish this by extending Atom and RSS with new aspects, including a verb and an object type. The world of FriendFeed, which supported a unified feed of 58 different services, where people could have one single stream that represented their identity online helped guide much of ActivityStreams' framework. As entries flowed to FriendFeed, they largely represented actions of posts, shares, bookmarks, reviews, from different sites. But ActivityStreams is aiming to do more than just syndicate data from one home page to another, as RSS and Atom have done for a decade, but also display intent and meaning.

"If your goal is to help people produce meaning, knowledge and culture, you have the basics for a pretty compelling social application and can motivate people to act," Messina said.

But with more streaming of information from many different sites, it can exacerbate the assumed problem of information overload - and tools need to be further developed to help us consume the data.

"We snack on information. It may feel like overload, but the tools haven't caught up," Messina said. "The solution to data overload is more metadata and we are at that point where can start generating that. We take the basic construct from 1999 and weave in some additional information - data about data."

As I outlined in my summary of DeWitt Clinton's talk on Google Buzz at the beginning of the month, Activity Streams are playing a big role in this new network, and these streams are intended to be open, not just for a company like Google, but other social networks that are tracking individual's activity and intent. The goal is to make discovery of intent data ubiquitous and transitive between sites, in the same way that RSS and Atom focused on publishing of data from one site to another.

Messina called the work on Activity Streams as iterative "baby steps", but ones that focus on getting today's rich media activity a home with a rich experience, and to make this process easy for service providers.

"If you have a Web site that has people doing things on it, and they have a feed they are taking with them, it is fairly trivial to add ActivityStreams information," he said. "Essentially we have the verbs and object types represented."

You can find out more on the continued development of ActivityStreams at http://activitystrea.ms.

March 4, 2010

Designing Buzz for a Google-Free World

Designing Buzz for a Google-Free World

If you haven't seen a lot of applications built in the last few weeks that leverage the Google Buzz API, it's because there aren't any. In fact, Google hasn't yet rolled out any API for Buzz. According to the company, this isn't due to any backroom dealings where they plan to introduce proprietary code and hooks that tie activity to their platform, but instead, because they wanted to be sure they could first build a product that in fullness leveraged open Web standards, and start with that foundation to deliver an interoperable system that could continue to function even if Google were to "disappear off the face of the earth".

In a presentation to the Silicon Valley Google Technology Users Group last night, held at the Google campus, DeWitt Clinton, a software engineer for the company, talked to developers and other tech enthusiasts about the company's API strategy and approach to Buzz, and explained that Buzz is designed not to increase lock-in to Google, but instead, to leverage open technologies that will let data flow to and from sites without central ownership. While a Buzz API will eventually be released, it will leverage the same open standards that power it today.

"The first principle of Buzz is that we can build this on protocols that are open and free, but not centralized," DeWitt said. "Can Google disappear off the face of the earth and Buzz still works? We need to make this data federated and distributed."

On the day Buzz launched, I referenced much of the foundation for Buzz in a quick article about the open tools and APIs that "make Buzz hum". But last night, DeWitt expanded that story to include 9 major open APIs, briefly outlined below.

1. Atom

DeWitt called Atom "the lingua franca of the programmable Web today", explaining that Atom contains entries that are "well structured", and include source entry, GUIDs that enable deduplication, and specification of the content type. He said, "You are able to pass rich data in that Atom feed in a way that is more specific than other feed types."

2. AtomPub

DeWitt said AtomPub "has become the most popular paradigm for restful APIs on the Web." AtomPub expanded the original Atom format to include the ability to both create and update feeds, not just passively read.

3. ActivityStreams

ActivityStreams essentially watch users' activity and can specify rich verbs and actions within those feeds. This enables feeds for all comments posted on Buzz, all likes, or even alerts that one person following you on Buzz also follows you on another network. DeWitt's examples hint at future developments for the platform, as these specific feeds are not yet clearly visible.

4. Pubsubhubbub

Much discussed here on the blog, Pubsubhubbub reduces the need for sites to poll for updates, and powers real-time updates between services. DeWitt reiterated "the hub is decided on by the publisher" and "there is nothing Google-specific about that.", saying that the infrastructure and plumbing for Buzz has been laid for the last few years. Pubsubhubbub has been pioneered by Brad Fitzpatrick and Brett Slatkin, both Google employees.

5. MediaRSS

Developed by Flickr, MediaRSS syndicates rich media through both RSS and atom feeds, creating a structured namespace inside RSS for content and a thumbnail. Buzz leverages MediaRSS, letting you pull rich content, like Flickr photos, into the platform. Of course, PicasaWeb, a Google property, also supports MediaRSS.

6. OAuth

The product of engineers from all corners, including Twitter, OAuth was engineered "to solve a vexing problem in the industry," Dewitt said, explaining OAuth prevents the need to ask users for their name and passwords on third party sites, acting as a delegated authorization protocol that gives permission to the application. Google Buzz, like Twitter, leverages OAuth to provide authenticated access to your data.

7. WebFinger

A new-age version of the old command-line prompted, text responding Finger protocol, WebFinger aims to be a way to get public information tied to an individual, through their identity, assigned to an e-mail address. "We want people to identify themselves, and we want people to discover people," DeWitt said.

WebFinger is similar to the strategy of OpenID, but OpenID hasn't had massive adoption by end-users who have found it unwieldy. WebFinger, aiming to be less arcane, enables the independent nature of Buzz, helping to federate the data and distribute it by domain, owned by the end user. DeWitt said, "The profile lookup and notification mechanism can be in the hands of the user being addressed."

8. Salmon

Still in earliest stages of development, Salmon is an extension or replacement for the old PingBack model that had blogs informing the other about references or links. This "flawed" model only provided minimal data, and could not be verified, letting me send PingBacks anywhere I wish if I chose. Salmon's goal is to leverage what's being called "Magic Signatures", signed with a public key to prove and verify linkage.

The first approach for Salmon will be to migrate comments from aggregators to originating posts, as covered a few times on this blog. But DeWitt said that "Likes" are similar activities that could flow back with Salmon, or be used to notify users of "following" or other activity. DeWitt forecast that sites like Blogger and StatusNet would rapidly adopt and federate Salmon to transmit data updates.

9. Portable Contacts

Simply described, Portable Contacts show your information and that of the friends who you follow, providing users a secure way to get access to address books and friends lists without having to request credentials or scrape the data.

DeWitt also noted XFN, the XHTML Friend Network, and FOAF (Friend Of a Friend) as being key contributors to the Buzz technology stack today, adding that he was "glad smart people were working on this ten years ago because we are all benefiting from it now."

DeWitt, on his Buzz feed, has been talking a lot about open standards and their importance to the Google team at large. See @Jesse Stay A few points of clarification to your most recent post [1], because I believe getting the details right matters. and "The thing I find most attractive about Google Buzz is its stated commitment to open standards.", as well as his first post from February 21st, which thanked the standard developers: Standing on the shoulders of giants—a look at the people behind the protocols behind Google Buzz:

Given Google's size, there is a good amount of distrust on the Web from people who think they own too much of your data, know too much about you, or have goals that run contrary to your own ideals on privacy, communication and sharing. Not even DeWitt's detailed presentations and explanations and promises of openness and data portability will convince everyone that they are on the right path. But I personally believe the frankness and detail that is being shown here is not just promising a strong future for this individual product (Buzz), but also in extending the groundwork done for the entire Web, for products and services we haven't even seen yet.

DeWitt adds: "All of these protocols are open. They are literally also all free. They are intended to be used by everybody, with or without Google being involved. You don't have to ask us if you can use Salmon or Pubsubhububb. We have a liberal and permissive patent license."

Is Google going away? Not today, and not this year. Is Buzz perfect? No. Of course not. Can it do all the things I can do on other sites, like FriendFeed? No. Not yet. But it seems that the Buzz team has opted to make tradeoffs that favor fast shipping and openness over completeness and individual features. And if you don't trust Google, it sounds like you can do something about it.

"We are pretty adamant about not building this on proprietary technology," DeWitt said last night. "If any of you feel that it is not going in the right direction, you have the power to change its direction and Google will not stop you."

You can find me on Buzz here and can follow DeWitt Clinton on Buzz here.

November 18, 2009

Open Web Foundation Speeds Protocols' Legal Contracts

Open Web Foundation Speeds Protocols' Legal Contracts



On Tuesday, the Open Web Foundation released an agreement aimed to speed new specifications' ability to be adopted by downstream users, with the intent of spreading open tools throughout the Web. Though occupying the always-complicated intersection of both the legal world and the tech world, the agreement is very interesting. The non-profit organization, featuring leading geeks from many of Silicon Valley's best known and most-respected companies, is hoping to promote data portability and open Web standards, no matter their source. Tuesday's agreement makes it easier for others to implement specifications without requiring lengthy bureaucratic legalities, and already features 10 major protocols and services as having signed up.

Among the services that have committed to using the new agreement include Yahoo!'s Media RSS standard, OAuth, Microsoft's WebSlice, and my often mentioned personal favorites, the PubSubHubbub and Salmon Protocols, being promoted by employees from Google.

As explained on the Yahoo! blog, on Facebook's Developers' blog and at Standards Law, services such as OpenID and OpenSocial were both forced to spend a great deal of effort working on legalities, taking their sharp engineering resources away from doing what they do best - code. The hope is that by setting a standard for approvals and access, much of these headaches can be eliminated.

The agreement itself is lightweight, compared to many legal tomes, and essentially mirrors standards set by Apache and Creative Commons, both of which have much history in the Web community. It covers how to handle attribution, that users can be trusted to leverage the work without fear of patent lawsuits, and that downstream users will not lay claims to others' efforts.

It could be yet another important step in making sure the Web is open, and that users can expect similar behavior and access capabilities from site to site and service to service. See also:
The Blurry Picture of Open APIs, Standards, Data Ownership
from October 29th.

October 29, 2009

The Blurry Picture of Open APIs, Standards, Data Ownership

The Blurry Picture of Open APIs, Standards, Data Ownership

Look beyond "real-time" and "social", and you'll easily find another pair of tech buzzwords that everybody wants attached to their product or service - "open" and "standards". Companies are practically falling over one another to show they have embraced developers or users, letting data stream in and out of their products, while avoiding words like "proprietary" and "closed", which are PR death. But as you might imagine, the very definition of "open" can vary depending on who you talk to, what the service's goals are, and how they may leverage existing standards on the Web. Following the much-discussed news of Facebook debuting its "Open Graph API" on Wednesday, I traded a few e-mails with a few respected tech-minded developers, and found, unsurprisingly, that not everyone believes Facebook is fully "open". In fact, it's believed some companies are playing fast and loose with terms that should be better understood.

To quickly summarize the discussion, there are essentially three major ways to bucket "open" APIs, agreed those I contacted.
  • The first, "open access", means that anybody can use the API, but all the data in or out of the services is owned or controlled by the company whose service you are using. The Facebook Open Graph API "is open insofar as you do not violate their ToS", one developer wrote. "Here, 'open' is superfluous -- no (question) you're giving people open access to it, how else would they use it?"
  • The second type is that of an API that leverages open standards, including those such as XML, HTTP, and others. But that doesn't mean APIs that leverage those standards are open by definition. For example, Twitter's API is proprietary, even though it is built on open standards. The developer adds, "Here 'open' is just saying they've tried to incorporate best practices from other engineers -- it would be stupid if they didn't."
  • The third type is the most "open", including open standard APIs like OpenSocial, OpenID, PubSubHubbub, AtomPub and others. These APIs have a clear definition that can be utilized by multiple providers in a way that is interoperable, decoupling providers and consumers.
In short, you have "open but we control the process", "standing on the backs of open" and "truly open", if this opinion is accepted. The developer adds, "In short, the first two mean nothing, the last one actually fits the dictionary definition. The Web is built on open standard APIs and protocols."

Chris Saad, VP of Product and Community Strategy at JS-Kit, well known for his efforts in the data portability space, concurred, writing over e-mail:
"Facebook in particular has made a concerted effort to dilute the word open and use it in reference to a human/cultural thing when talking about the platform and their products."

He added, "In reality there is a VERY big difference between having an 'Open API', an 'Open Standards API' and an 'API'. An API is just a thing you poke and you get data back. When you get FaceBookPropietaryXMLData using FacebookPropietaryAuthMethod and you can only cache the data for 24 hours - that is NOT an open API - it is an API."
So who cares? Historically, services like Facebook and AOL have been characterized as walled gardens, meaning their information is sealed within, beyond the reach of the standard Web. Other services are known as "data roach motels", where data gets in, but never gets out. As the first developer said, the Web is built on open standard APIs and protocols, so sites can work well with each other, and activities operate in a similar manner, regardless of service.

Jesse Stay, a friend of mine, fellow blogger, and well-versed developer for both the Facebook and Twitter platforms, agreed that there is a tremendous amount of confusion around the definition of "open". In fact, just last month he wrote a post on his site, "The Open Web – Is it Really What We Think it is?"

Today he said Facebook's move gave full access to "users' walls, comments, likes and social graph... accessible from any Web site, desktop application or mobile application, using open API access protocols." Meanwhile, Facebook users can now opt into letting their status updates indexed by search engines, and the company is open sourcing architecture like the Tornado Web server (acquired as part of the FriendFeed buy) so other developers can make new platforms.

Jesse is more optimistic about Facebook's goals than was Chris. He said that the site lets users decide how open they want to be with their data, and that they are "working to give users full power" in that regard. But he also states frustration with the company's restricted access to search, and a lack of access to the entire network in aggregate, with the exception of their fan page directory. And he didn't address the core issue with Facebook in terms of them owning your data bidirectionally, and yes, them having the option to block your access if they felt you had violated the terms of service. (Remember this one? Scobleizer: Facebook Disabled My Account)

Web standards are very well known and we usually recognize them by their acronyms. JSON. HTTP. XML. POP3. Atom. Open means that developers can tap into the standard and use it as they wish, both procuring data and pushing it elsewhere. When we start to blur the lines about open and associate them with specific companies, like Twitter, Facebook, Yahoo! or others, you can usually guess that the solution is slightly less open. Somebody has the option to change their proprietary code and block you from having full access.

As stated more than a few times here, I have chosen to trust companies with my data. I put a lot of data into the Web and move it around. I expect standards to work the same way across sites, and I hope that those services that I use treat developers as well as they do their users. I recognize I am not as technical as folks like the developers I pinged today, and thus I need to trust their comments at times once my expertise is surpassed. But we need to be more knowledgeable about what is "open" and what is "sorta', kinda' open". Maybe Facebook can help us all understand their level of openness as time progresses.

February 6, 2009

Gnip Says To Make Money, Make Sure Your Customers Have Money

Gnip Says To Make Money, Make Sure Your Customers Have Money

Following on to yesterday's visit at Lijit, I knew my two-day trip to Boulder would not be complete without making time to visit Gnip, the interesting company started by former MyBlogLog and IGN founder Eric Marcoullier. So Micah Baldwin, my trusty sidekick and part-time chauffeur for the trip, and I caught breakfast with Eric and the team this morning, and visited their cozy headquarters. While we didn't get hours to sit with the executive team, as we did with Lijit, we learned that Gnip is growing and hiring talented developers, and has made an important discovery in its business model - target companies that have money and are willing to pay for your product.

In July of 2008, when Gnip first launched (See: Gnip CEO's Goal: Make Twitter's Data Flow Suck Less), the company made headlines for finding ways to move data around more quickly and without as much overhead, acting as an arbiter between different Web services. But as I was told today, making Web 2.0 services the primary client was pretty much a guarantee for low revenue. After all, if your customers aren't making money, how could you expect them to pay you?

As a result, Gnip, which has grown to 11 full-time employees, all of whom Eric says are required to be smarter than him in order to get hired, has gotten more activity with more traditional companies, including one market research firm that is analyzing as many as 100,000 different Twitter accounts and checks for user sentiment.

Gnip's office in Boulder has the industrious start-up feel to it. Desks are pushed together in a small space that reminds me of my freshman year dorm room I had to share with two other guys. But while the company is growing, it has a small quandary, as Boulder commercial real estate works great for small spaces and large spaces, I was told, but there just aren't enough options for medium-sized companies looking for an "in between" solution.

In fact, there is so little open space at Gnip's office that Eric, Shane Pearson and I talked on the front porch. Unfortunately, the front porch bench that adorned the office had been stolen overnight. The main suspect? "Stinking hippies", Eric tweeted.

Such growing pains are good, of course, because that means the company is growing, period, and it sounds like they are focused on continuing to improve the product, hire smart developers, and they even finally managed to grab the ever-elusive gnip.com domain, after using the gnipcentral.com site since launch.

To learn more about Gnip's unique view on the data world, check out their blog. Product news will no doubt be coming soon. Of the 11 employees, 8 are full-time engineers.

December 5, 2008

Getting Started With Google Friend Connect

Getting Started With Google Friend Connect

By Mike Fruchter of MichaelFruchter.com (Twitter/FriendFeed)


Back in May, Google announced their plans to launch Friend Connect. Friend Connect allows website owners to incorporate social aspects and elements onto their sites. It's powered by OpenSocial, which was developed by Google along with MySpace and a number of other social networks. In a nutshell, OpenSocial is basically a set of common APIs used for building and distributing social applications and their data across many websites. This data includes profile information, friends information, activities and so forth. You can find more about the OpenSocial platform here.

This post is a brief tutorial on the features and the implementation of Friend Connect.

Getting started:


Installation is very easy, and takes less then 5 minutes. First head over to the Google's Friend Connect website. Click the "Set up a new site" link to begin.

Give Google some basic info:

The next step is pretty much self explanatory. Input your site name, and the url of your website.


FTP two files to your server:

This step involves downloading two html files to your desktop. FTP those two files to your web server, and make sure they are placed in the root folder of your domain. This is important if you are going to use Friend Connect on a sub directory of your site, I.E /blog.

Test your setup
:

This will test to see if you uploaded the html files correctly. Now that you are almost finished, it's time to add and customize your gadgets.

Now for the social part, adding gadgets to your site:

The code you generate for the gadgets are html. It's a simple copy and paste process. There are two types of widgets available, social and members.

Members gadget: This allows visitors to join your site, sign in and out, see other members, and invite other people to join your site. A screen shot of the members gadget is shown below. You can also see and interact with the gadget live on the sidebar of my blog.


The other option is a sign in gadget pictured below. This is a lighter version of the members gadget without images. It's meant for small spaces and allows visitors to join your site, sign in and out, and invite others to become members.


Clicking a member's profile image on the member gadget will display their Google service profile. You can also add that person as a friend, and once they accept the request, you will able to see each other across all Friend Connect sites. You also have the option of blocking users.

Clicking the invite link brings up the email sharing aspect of the gadget. You can share a website with your Gmail contacts by clicking the Google tab. It's not limited to Gmail contacts, any email address will work. You can also add a quick, personalized note with the share.

The share tab takes it a step further up the social ladder by allowing you to post the website to Facebook, Myspace or Del.icio.us. You can also submit the site to Digg, Fark, Mixx, Live Spaces and StumbleUpon. In its current state it needs more social networks and services to plug into. Integrating other services, should be a top priority for Google. There are currently a dozen or so other free social sharing type services on the market that offer more of a selection than Friend Connect in its current state.

Social gadgets: This area is kind of sparse at the moment. There are only two worthy gadgets that are of any real social use and benefit. One is a wall gadget and the other is a review/rate gadget.Bwana, from bwana.org has all of these gadgets integrated live on his blog. Head over to his site to see them in action. The screen shots below were captured from his blog.

Wall gadget:This allows people to post comments, and also links to (YouTube) videos on your site. It's customizable and makes for a great shout out box.



Review/Rate gadget: This gadget can be used in a variety of ways. It lets people rate your content via a simple starring system. You can use this gadget to rate any type of content, videos, photos, etc. Comments can also be posted on it. This can also be used numerous times on a single page, or on separate pages of your site. I believe this will become widely used by bloggers, once they catch onto it. This is something I will be experimenting with on future blog postings.


Games? No just lame!

Being that the service just officially launched, it would of been nice for them to include at least a variety of already developed widgets other than the two mentioned above. All we get for now is the Lame Game. Hey, they named it that, not me, but what else could you really call it? Its only purpose is to allow your friends to go click crazy. This is done by clicking as often as they can which leads to trying to increase their score vs. their friends and other site members.



Lets touch on the admin area:

There is not a lot of back end features to report. I expected to see some type of url referral reporting function, similar to that offered by MyBlogLog. Where members came from, what they viewed, what they clicked, this is what I want to see. Google are you listening?

Clicking the reports link will show you two simple graphs. The first displays total membership over time, and the second displays new members over time.


Clicking the moderate posts will allow you to delete any obscenities. You can also choose to have all posts be approved before they are posted. Wall posts are immediately visible to anyone by default.

Clicking the manage members link allows you to do the normal friending functions. View members Google service profiles, add members, block members, and you can also view their friends as well. You also have the ability to grant administrative privileges to specific members.
Wrapping it all up, and some additional thoughts:

In regards to it being a MyBlogLog killer. I would seriously consider totally removing the MyBlogLog widget off my blog in favor of Friend Connect. Of course that's under two circumstances. The first is to show me the most recent visitors to my site, as does the MyBlogLog widget. The second is to give me url referral reporting and additional stats on where readers came from, what they viewed and so forth. For now, both of these widgets compliment each other fine, and although similar in function, in some aspects are two totally different beasts.

The product is still rough around the edges, as to be expected with any beta. It's a step in the right direction for data portability, and that's a plus in my book. I'm fascinated with the direction and possibility that this could lead to in helping us better connect with members in our social graph and abroad. The fact that Google is behind it, could be a blessing or a curse. Now of course getting the other players in the space, such as Facebook, to play nice, is a whole different ball game.

Will Google's Friend Connect and Facebook Connect initiatives be a progressive step toward evolving social networking or will it be just another passing fad? I believe the former, not the latter.

Read more by Mike Fruchter at MichaelFruchter.com.