Showing posts with label Information. Show all posts
Showing posts with label Information. Show all posts

Saturday, January 17, 2009

Information Overload, a problem or an opportunity

I wrote this one a while back and realized I had posted its sequel already, so maybe it is time to post this one already. If anyone is ready of it, sorry about that.

I look at information overload in two ways. On the one hand there seem to be so much information out there, how can you find the stuff you care about? The other one is how easily you can publish information about yourself on-line, whether you want to or not.

For the first part, does the Internet have too much information? The Internet is has much about democratization of publishing capability as anything else (well maybe not, but it is one of the best side effects). I still read books, but most of the information I use on a daily basis in on-line. As soon as it costs less than $5 for me to print an entire book at home, I am not sure I will ever need to go to a bookstore. Although browsing on-line does not really the same experience, I can make my own hot chocolate while browsing the books, I down need the cafe experience as well. I have been subscribing to o'reily radar for a while where the subject of the future of publishing comes up on a regular basis. They should know, I find their discussions about the Internet fascinating.
I came across this Ted presentation where the speaker discuss the fact that we do not have too much information, we always had that, we only have inappropriate filters.

I suppose I already have my own ideas about what those filters could be. So, it would then seem that it is not so much there is too much out there, it is just that we need better ways to get to it.

As for the second part, information about yourself being published. I goggled myself once, and I found a reference that seemed to match. I went to that web site where I could find my phone number and address. I can Google my phone number and find my name and address. I got more involved and tried a few web sites. I am now in FlickrFacebook, LinkedIn, I share articles I like in Google Reader and even started twittering (is this a verb in the dictionary yet?). It took me a while to get to twitter because I could not see myself telling others small things on a regular basis ('I am eating', 'I am thinking about going to bed', etc.). Still it seems like a nice way to learn about people I know. I got to learn a few things here are there, it was nice. I might get to publish some of my own at some point. After watching the Ted video, I went and paid attention to the security settings. Apparently it is one of the best for a social site. I was hoping to at least have a differentiation between friends and family, like Flickr does. I was surprise to discover it was not there. So what do you when your parents subscribe to your page. I don't have that problem, but I will definitely register to my son's when the time comes :). I don't really like people knowing about my birthday, I am still trying to find a way to hide that. Maybe I should change it to a false one and stop worrying about it. The other concern is that I need 3 profiles at least, a work one, a family one and last one for friends. Until this is possible, the lowest common denominator will probably win. Is it just me or it should be work for all working individual?
So, my issue now is how to keep all those sites up to date? Updating my status in Facebook does not update Twitter. That's a nice application to write for Facebook, I should look into that. I am spending so much time thinking about blog, pictures, keeping in touch... Finding a way to communicate efficiently is becoming difficult. Can't my Flickr pictures be automatically referenced in Facebook and notify twitter about it?
The next worry is of course privacy. Because of all the information I post to all the sites, there is so much more about myself on-line than there had been before two years ago, I have looked at Mint and PageOnce as well, they look great. Unfortunately I have not found the courage to give that much information about myself in one site. You are never sure where the company will go in the future and they could be. I suppose that's if you are lucky, the company is still there and your information still exists. I am a bit scared of having one place which would have the list of all my web sites and a way to aggregate all that information about me. Just too scary to think about it at this point. All I need is for them to expose my information using the Google social API and everyone knows.

So what is the conclusion of all that? I need to subscribe very carefully to a few feeds and keep an eye for new ways to filter information so I do not spend 3 hours everyday catching up with the world while still learning about things I did not know I wanted to know more about. For now I need to find a way to keep my private information locally and synchronized will all my computers and all. I have started writing my own contact management application because nothing I find satisfy me and I am curious about what could be done. Nothing like doing it yourself to find out. I will wait for this magic web application which will let me control my information and not give it to anyone else. How long will the wait be?

I suppose I already have my own ideas about what those filters could be. So, it would then seem that it is not so much there is too much out there, it is just that we need better ways to get to it.

As for the second part, information about yourself being published. I goggled myself once, and I found a reference that seemed to match. I went to that web site where I could find my phone number and address. I can Google my phone number and find my name and address. I got more involved and tried a few web sites. I am now in FlickrFacebook, LinkedIn, I share articles I like in Google Reader and even started twittering (is this a verb in the dictionary yet?). It took me a while to get to twitter because I could not see myself telling others small things on a regular basis ('I am eating', 'I am thinking about going to bed', etc.). Still it seems like a nice way to learn about people I know. I got to learn a few things here are there, it was nice. I might get to publish some of my own at some point. After watching the Ted video, I went and paid attention to the security settings. Apparently it is one of the best for a social site. I was hoping to at least have a differentiation between friends and family, like Flickr does. I was surprise to discover it was not there. So what do you when your parents subscribe to your page. I don't have that problem, but I will definitely register to my son's when the time comes :). I don't really like people knowing about my birthday, I am still trying to find a way to hide that. Maybe I should change it to a false one and stop worrying about it. The other concern is that I need 3 profiles at least, a work one, a family one and last one for friends. Until this is possible, the lowest common denominator will probably win. Is it just me or it should be work for all working individual?
So, my issue now is how to keep all those sites up to date? Updating my status in Facebook does not update Twitter. That's a nice application to write for Facebook, I should look into that. I am spending so much time thinking about blog, pictures, keeping in touch... Finding a way to communicate efficiently is becoming difficult. Can't my Flickr pictures be automatically referenced in Facebook and notify twitter about it?
The next worry is of course privacy. Because of all the information I post to all the sites, there is so much more about myself on-line than there had been before two years ago, I have looked at Mint and PageOnce as well, they look great. Unfortunately I have not found the courage to give that much information about myself in one site. You are never sure where the company will go in the future and they could be. I suppose that's if you are lucky, the company is still there and your information still exists. I am a bit scared of having one place which would have the list of all my web sites and a way to aggregate all that information about me. Just too scary to think about it at this point. All I need is for them to expose my information using the Google social API and everyone knows.

So what is the conclusion of all that? I need to subscribe very carefully to a few feeds and keep an eye for new ways to filter information so I do not spend 3 hours everyday catching up with the world while still learning about things I did not know I wanted to know more about. For now I need to find a way to keep my private information locally and synchronized will all my computers and all. I have started writing my own contact management application because nothing I find satisfy me and I am curious about what could be done. Nothing like doing it yourself to find out. I will wait for this magic web application which will let me control my information and not give it to anyone else. How long will the wait be?

Can too much information be bad?

I was watching this great video about from Kevin Kelly who gave a great interview at the Web 2.0 Summit. I liked most of the interview, I could not help react to his final vision of the future. Us being connected to all of the data in the Internet cloud, our perceptions being expanded beyond our current "shell". I find the concept scary. I find it scary for 2 reasons. First with all this data available, even I who is a bit wary of doing it too much, I spend most of my time trying to catch up with information. The information overload, I do not need to start again. Second, the notion of always being connected seems a bit disturbing.

On the one hand, I understand the need to have better information. Better connected information with access to the multi-dimensional aspect of information available. A bit like what is available is mashups . A map of my travel pictures combined with pictures taken by other people and articles from wikipedia. But somehow this gets blended in an easy way, not sure I have a vision of what that is at this point. The Internet has so much to be positive about, there is no corner of the planet, or soon anyway, that does not have access to free information, unfiltered. Most people can express opinions and expose anyone (ok, this one could be bad, but it is also pretty cool).

I believe we can run away from ourselves. By trying to learn all the time, we do not really try to understand ourselves. There is a lot that can be done here, but it can only be done while the mind is quiet. Are we going to measure ourselves by the amount of information we have? Or are we going to take the time to understand our nature by taking the time to listen to it.

This kind of thoughts always brings me to how ideas are created. We often refer to the stupidity of crowds and how it is difficult to come up with something brilliant while debating with the rest of the team, the best ideas are created in the shower. I think group meetings are important, this is where you level the plain field, this is where you learn about things that are relevant to what you are trying to achieve, providing you are not too set one way already. But then individuals will take the matter at hand and do something with it. If the group tries to dictate the final answer, you will end up with a common denominator and nothing brilliant (well in most cases anyway). But you can use the group as a way to gather the necessary information to make the idea possible.

So, on the other hand, if the information is available so easily, maybe we do not have to pursue knowledge to that extend, maybe we can relax a bit knowing that once information is needed, it will be easily found. Maybe we will find new grounds to compete in, because we will always compete. Maybe we will excel in synthesizing knowledge and create new intelligence that we will freely give for everyone else to make more knowledge with. That is a bit of an optimistic vision, but I kind of like it. I think it will take a while before people feel comfortable giving away information like this. We might have to change the way we make money or even live, so we do not need money. One can always dream...

Saturday, October 4, 2008

Contact and Document Management

I have been involved in creating a search application for a while. The application was hinging on also becoming a document management system while it enabled users to collaborate on the documents in order to create intelligence based on the raw data contained in the documents. The application did not really make it that far, but after I had taken a step back and thought about what I had been doing for a while. I came to the conclusion that at the basis of all information gathering and collaboration is good contact management.
  • We needed some security of course, that is the easiest one of all, but you need to make sure every one is well identified. So that is the mostly easy part. Of course combine this with groups and roles and the like and you are covered. It is of course not as easy as that, but everyone can do at least that much, so it is not that hard.
  • What comes next is that not everyone creates data with the same skill and quality of content as everyone else. You must identify who you experts are and what their field of expertise is. Of course one of the issue it that this will change over time. If you have just recruited a new assistant and he is creating documents reflecting his training in a given field, you would want to give priority to the more senior and experienced resources when searching data on that field. You could even go further and privilege any efforts made when several experienced resources collaborate together.
  • Of course this seem to apply only at a company level where that information is well known and understood. It seems a bit more difficult to achieve this at the Internet level. We have lots of information about one single document, and potentially about a site as well. I believe we could also collect information about the documents a single person has written and what the subjects are. I believe we collect a list of keywords for each document and who references those documents (for Google anyway). Those keywords could be cross referenced (or not) with fields of expertise.
  • The next stage would be to combine that with user feedback. Present them with a top list of expertise they declare themselves to be cognizant about and then have them rate articles, and by extension the authors of those articles. The last stage, or at least the last one I can think of, would be to include social information. I suppose I mean for you to be able to generate a graph of trusted sources for specific subjects. It would be great to be able to be able to use that information for future searches and being able to visualize it as well as customize it. Based on previous document feedback, we could already build such a graph already, but being able to customize it would increase its efficiency. I could define an author I have heard of, without knowing whether there are any articles written by that person. But we could also recommend other authors who have recommended him and are recommended by her.
  • The other thing that would be great, would be to be automatically notified when a preferred author has published, but only if the web site is actually related to my interests. If I publish a new picture in Flickr or if I publish a new entry in my blog about the latest coding practice, they do not speak to the same interest and to be honest my photographic skills and not really worth keeping track of.

This is probably too complex to ever be understood by everyone or not many people will actually bother setting all those filters. There must be a compromise here somewhere. 

Sunday, September 28, 2008

Just a quick one on search

I was reading this article and I have to say my first reaction was why can't you just ask the user. Of course few paragraphs later, there was the line indicating they use trained individual who evaluate the results returned.

But how can you train someone to think like me and look for the things I am looking for? Of course I am not that special and probably fall right in the middle of the bell curve of regular person. But I'd still like to think I am not exactly learning what everyone else is, or even not learning it the same way everyone else is.

One easy way would be to check which position within the results I selected to look at and which ones I actually reported as being relevant to my search. This could be implemented already without much changes to the current Google results, then again maybe they are already doing that, well the first part anyway.
Is there a way for me to disqualify pages/sites from search results already?  From this question comes the obvious next point. It would be nice to be able to exclude the places I have already been to. Or maybe not put them on top of my search results. Been there, done that type of approach, probably don’t want to do it again for a while. It is a bit counter productive for me to have to scroll through all those places I have already been to. We could at least flag those pages as recently visited, that would help me with the issue of, I think I have already read this, why could I not remember this based on the obscure hyperlink I clicked on.

As soon as I started thinking about how to collect what I think relevant is, I quickly started wondering about what relevant really means. What is a relevant result? Is it something that gives me exactly the information I was looking for or is it something that also gives me the ability to find something I did not think about but opens my understanding of the searched topic? Every time my mind wonders in that direction, I think about ‘people who bought this book also bought…’. Which of course is not as relevant as people who loved this book also read those ones when looking for a similar topic.

The other thought that came to my mind was, when I am googling, I am trying to learn about something 80% of the time and then I am probably trying to entertain myself. Can the search engine differentiate between one mode and the other? Should I be fed similar results depending on that search mode? Personally I am more interested in hard core articles (with lots of texts) full of information when learning and I need lots of images and moving pictures when entertaining myself.  You can tell I just like text by looking at this blog, I might need to add more picture if I ever want someone else to look at it and get something out of it. I almost want different modes, the casual lots of pictures showing in the results mode and the hard core facts one. And I should be able to choose every time based on the type of search I am currently conducting.Then again, not everyone learns the same way, it is a bit of a conundrum. Some people like to learn by looking at pictures. Is there ever the hope a search engine could understand the difference?
Here comes the basic point again: it will never happen unless we collect information from the user and usage of the system. And of course, I would not agree to it unless I could review the information collected. I would even go as far as saying I would turn the collection off is there is no way for me to be able to control the information itself. Anyway, I have some crazy probably never going to happen kind of ideas about what searching should really be. It cannot remain a cookie cutter approach for much longer, it will need to include personalization and preferences. It is only a matter of time.

Thursday, September 4, 2008

My wish list for Google Search

My current belief is that for us to truly gather information, we need to understand the source of the information. The easiest thing we could do today is to record the author of articles, web pages, blog entries...
We have lots of ways of doing that, tracking the keywords, owner of the blog, signature at the end of the page. Yet what surprises me is that the information is not displayed in a search result page.

It would be nice if we could see the author, have a way to review other articles from that author ordered by relevance to your current search. If you could use something similar to the Social Graph API to access other resources by that author. Obviously you probably want to limit those to the ones the author is comfortable exposing as public web sites (if there is such a thing as a private one still). I suppose you would also want to have references to web site the author contributes to without necessarily owning the sites.

Of course we all know how evil Google is, it is tracking so much information about us without telling us about it. That data is used to know who you are, what you are interested in and use that information to narrow your search result items to what might be more relevant to you. I believe the problem lies in the fact that this information is hidden from you and that there is little you can do to influence how the data is being used. I know you can view the history and you can remove items but there is a difference between removing a history item and deciding not to use it for future searches. If users could truly own their data and influence its content, I believe the privacy issue would be less of a problem. Nobody wants a system that gathers data about you, to help you of course, but does not expose how that data is used and barely exposes what the data is. When you make a search, I want to have an explain option where I will be able to see how my previous searches/wanderings have influenced the results for this page. It would be especially important then to have the option of saying don't use that one, this is not me really. Then I might be ok to have the information used to also deliver some targeted material. I need to know what the benefits to me are before I feel comfortable it being used by someone else. This does not mean Google needs to expose the algorithm used to deliver the results, only how my information is influencing the results presented to me.

Let's say the user now has a way to set those preferences, we could take the approach a step further. I would want to use something similar to Stumble Upon where I can rate authors and articles in the context of my original search. This would allow dynamically adding more intelligence about the document itself but also customize my search results based on previously approved authors and articles. On the top of that you can then cross reference results and ranking from other users 'similar to you' or from other previous 'similar' searches and maybe give more relevant results. Probably something not too far of the Stumble Upon toolbar which can rate the results based on your preferences.

For the more advanced user, you can start adding a social aspect to this and add the ability to register some of your friends and when you look for something your friends already found, you could have some additional information equivalent to your friends' recommendations. But by then that is probably a little complicated and too cumbersome to really be useful. I might have people I trust in my circles, but I probably do not trust them in all aspects of my interests. If we have to think that much before deciding how to influence my searches, no-one will go through the effort.

So to summarize:

  • Start displaying the authors of the information and give access to their work

  • Show the user how the data gathered influences the search results

  • Allow the user to enter preferences for sources of information

  • Ignore everything else I added