The personal and the institutional

Twittering and microblogging not permitted
Image by cameronneylon via Flickr

A number of things recently have lead me to reflect on the nature of interactions between social media, research organisations and the wider community. There has been an awful lot written about the effective use of social media by organisations, the risks involved in trusting staff and members of an organisation to engage productively and positively with a wider audience. Above all there seems a real focus on the potential for people to embarrass the organisation. Relatively little focus is applied to the ability of the organisation to embarrass its staff but that is perhaps a subject for another post.

In the area of academic research this takes on a whole new hue due to the presence of a strong principle and community expectation of free speech, the principle of “academic freedom”. No-one really knows what academic freedom is. It’s one of those things that people can’t define but will be very clear about when it has been taken away. In general terms it is the expectation that a tenured academic has earnt the right to be able to speak their opinion, regardless of how controversial. We can accept there are some bounds on this, of ethics, taste, and legality – racism would generally be regarded as unacceptable – while noting that the boundary between what is socially unacceptable and what is a validly held and supported academic opinion is both elastic and almost impossible to define. Try expressing the opinion, for example, that their might be a biological basis to the difference between men and women on average scores on a specific maths test. These grey areas, looking at how the academy ( or academies) censor themselves are interesting but aren’t directly relevant to this post. Here I am more interested in how institutions censor their staff.

Organisations always seek to control the messages they release to the wider community. The first priority of any organisation or institution is its own survival. This is not necessarily a bad thing – presumably the institution exists because it is  (or at least was) the most effective way of delivering a specific mission. If it ceases to exist, that mission can’t be delivered. Controlling the message is a means of controlling others reactions and hence the future. Research institutions have always struggled with this – the corporate centre sending once message of clear vision, high standards, continuous positive development, while the academics privately mutter in the privacy of their own coffee room about creeping beauracracy, lack of resources, and falling standards.

There is fault on both sides here. Research administration and support only very rarely puts the needs and resources of academics at its centre. Time and time again the layers of beauracracy mean that what may or may not have been a good idea gets buried in a new set of unconnected paperwork, that more administration is required taking resources away from frontline activities, and that target setting results in target meeting but at the cost of what was important in the first place. There is usually a fundamental lack of understanding of what researchers do and what motivates them.

On the other side academics are arrogant and self absorbed, rarely interested in contributing to the solution of larger problems. They fail to understand, or take any interest in the corporate obligations of the organisations that support them and will only rarely cooperate and compromise to find solutions to problems. Worse than this, academics build social and reward structures that encourage this kind of behaviour, promoting individual achievement rather than that of teams, penalising people for accepting compromises, and rarely rewarding the key positive contribution of effective communication and problem solving between the academic side and administration.

What the first decade of the social web has taught us is that organisations that effectively harness the goodwill of their staff or members using social media tools do well. Organisations that effectively use Twitter or Facebook enable and encourage their staff to take the shared organisational values out to the wider public. Enable your staff to take responsibility and respond rapidly to issues, make it easy to identify the right person to engage with a specific issue, and admit (and fix) mistakes early and often, is the advice you can get from any social media consultant. Bring the right expert attention to bear on a problem and solve it collaboratively, whether its internal or with a customer. This is simply another variation on Michael Nielsen’s writing on markets in expert attention – the organisations that build effective internal markets and apply the added value to improving their offering will win.

This approach is antithetical to traditional command and control management structures. It implies a fluidity and a lack of direct control over people’s time. It is also requires that there be slack in the system, something that doesn’t sit well with efficiency drives. In its extreme form it removes the need for the organisation to formally exist, allowing a fluid interaction of free agents to interact in a market for their time. What it does do though is map very well onto a rather traditional view of how the academy is “managed”. Academics provide a limited resource, their time, and apply it to a large extent in a way determined by what they think is important. Management structures are in practice fairly flat (and used to be much more so) and interactions are driven more by interests and personal whim than by widely accepted corporate objectives. Research organisations, and perhaps by extension those commercial interests that interact most directly with them, should be ideally suited to harness the power of the social web to first solve their internal problems and secondly interact more effectively with their customers and stakeholders.

Why doesn’t this happen? A variety of reasons, some of them the usual suspects, a lack of adoption of new tools by academics, appalling IT procurement procedures and poor standards of software development, and a simple lack of time to develop new approaches, and a real lack of appreciation of the value that diversity of contributions can bring to a successful department and organisation. The biggest one though I suspect is a lack of good will between administrations and academics. Academics will not adopt any tools en masse across a department, let alone an organisation because they are naturally suspicious of the agenda and competence of those choosing the tools. And the diversity of tools they choose on their own means that none have critical mass within the organisation – few academic institutions had a useful global calendar system until very recently. Administration don’t trust the herd of cats that make up their academic staff to engage productively with the problems they have and see the need to have a technical solution that has critical mass of users, and therefore involves a central decision.

The problems of both diversity and lack of critical mass are a solid indication that the social web has some way to mature – these conversations should occur effectively across different tools and frameworks – and the uptake at research institutions should (although it may seem paradoxical) be expected to much slower than in more top down, managed organisation, or at least organisations with a shared focus. But it strikes me that the institutions that get this right, and they won’t be the traditional top institutions, will very rapidly accrue a serious advantage, both in terms of freeing up staff time to focus on core activities and releasing real monetary resource to support those activities. If the social side works, then the resource will also go to the right place. Watch for academic institutions trying to bring in strong social media experience into senior management. It will be a very interesting story to follow.

Reblog this post [with Zemanta]

“Friendfeeds for Science” pt II – Design ideas for a research focussed aggregator

Who likes me on friendfeed?
Image by cameronneylon via Flickr

This post, while only 48 hours old is somewhat outdated by these two Friendfeed discussions. This was written independently of those discussions so it seemed worth putting out in its original form rather than spending too much time rewriting.

I wrote recently about Sciencefeed, a Friendfeed like system aimed at scientists and was fairly critical. I also promised to write about what I thought a “Friendfeed for Researchers” should look like. To look at this we need to think about what Friendfeed, and other services including Twitter, Facebook, and Posterous are used for and what else they could do.

Friendfeed is an aggregator that enables, as I have written before, an “object-centric” means of interacting around those objects. As Alan Cann has pointed out this is not the only thing it does, also enabling the person-centric interactions that I see as more typical of Facebook and Twitter. Enabling both is important, as is the realization that all of these systems need to interoperate effectively with each other, something which is still evolving. But core to the development of something that works for researchers is that standard research objects and particularly papers, need to be first class objects. Author lists, one click to full text, one click to bookmark to my library.

Functionality 1: Treat research objects as first class citizens with special attention, start with journal papers and support for Citeulike/Zotero/Mendeley etc.

On top of this Friendfeed is a community, or rather several interlinked communities that have their own traditions, standards, and expectations, that are supported to a greater or lesser extent by the functionality of rooms, search, hiding, and administration found within Friendfeed. Any new service needs to understand and support these expectations.

Friendfeed also doesn’t so some things. It is not terribly effective as a bookmark tool, nor very good as tool for identifying and mining for objects or information that is more than a few days old although paradoxically it has served quite well as a means of archiving tweets and exposing them to search engines. The idea of a tool that surfaces objects to Google is an interesting one, and one we could take advantage of.  Granularity of sharing is also limited, what if I want slidesets to be public but tweets to be a private feed? Or to collect different feeds under different headings for different communities, public, domain-specific, and only for the interested specialist?

Finally Friendfeed doesn’t have a very sophisticated karma system.  While likes and comments will keep bringing specific objects (and by extension the people who have brought them in) into your attention stream there is none of the filtering power enabled by tools like StackOverflow. Whether or not such a thing is something we would want is an interesting question but it has the potential to enable much more sophisticated filtering and curation of content. StackOverflow itself has an interesting limitation as well; there is only one rank order of answers, I can’t choose to privelege the upmods of one specific curator rather than another. I certainly can’t choose to order my stream based on a persons upmods but not their downmods.

A user on Friendfeed plays three distinct roles, content author, content curator, and content consumer. Different people will emphasise different roles, from the pure broadcaster, to the pure reader who doesn’t ever interact. The real added value comes from the curation role and in particular enabling granular filtering based on your choice of curators. Curation comes in the form of choosing to push content to Friendfeed from outside servces, from “likes”, and from commenting. Commenting is both curation and authoring, providing context as well as providing new information or opinion. But supporting and validating this activity will be important. Whatever choice is made around “liking” or StackOverflow style up and down-modding needs to apply to comments as well as objects.

Functionality addition 2: Enable rating of comments and by extension, the people making them

If reputation gathering is to be useful in driving filtering functionality as I have suggested we will need good ways of separating content authoring from curation. One thing that really annoys me is seeing an interesting title and a friendly avatar on Friendfeed and clicking through to find something written by someone else. Not because I don’t want to read something written by someone else, but because my decision to click through was based on assumptions about who the author was.  We need to support a strong culture of citation and attribution in research. A Friendfeed for research will need to clearly mark the distinction between who has brought an object into the service, who has curated it, and who authored it. Both should be valued but the roles should be measured separately.

Functionality addition 3: Clearly designate authors and curators of objects brought into the stream. Possibly enable these activities to be rated separately?

If we recognize a role of author, outside that of the user’s curation activity we can also enable the rating of people and objects that don’t belong to users. This would allow researchers who are not users to build up reputation within the system. This has the potential to solve the “ghost town” phenomonen that plagues most science social networking sites. A new user could be able to claim the author role for objects that were originally brought  in by someone else. This would immediately connect them with other people who have commented on their work, and provide them with a reputation that can be further built upon through taking on curation activities.

This is a sensitive area, holding information on people without their knowledge, but it is something done already across indexing services, aggregation services, and chat rooms. The use of karma in this context would need to be very carefully thought out., and whether it would be made available either within or outside the system would be an important question to tackle.

Functionality addition 4: Collect reputation and comment information for authors who are not users to enable them to rapidly connect with relevant content if they choose to join.

Finally there is the question of interacting with this content and filtering it through the rating systems that have been created. The UI issues for this are formidable but there is a need to enable different views. A streaming view, and more static views of content a user has collected over long periods, as well as search. There is probably enough for another whole post in those issues.

Summary: Overall for me the key to building a service that takes inspiration from Friendfeed but delivers more functionality for researchers, while not alienating a wider potential user base is to build a tool that enables and supports curation rating and granular filtering of content. Authorship is key, as is quantitative measures of value and personal relevance that will enable users to build their own view of the content they are interested in, to collect it for themselves and to continue to curate it for themselves, either on their own or in collaboraton with others.

Reblog this post [with Zemanta]

The Panton Principles: Finding agreement on the public domain for published scientific data

Drafters of the Panton principlesI had the great pleasure and privilege of announcing the launch of the Panton Principles at the Science Commons Symposium – Pacific Northwest on Saturday. The launch of the Panton Principles, many months after they were first suggested is really largely down to the work of Jonathan Gray. This was one of several projects that I haven’t been able to follow through properly on and I want to acknowledge the effort that Jonathan has put into making that happen. I thought it might be helpful to describe where they came from, what they are intended to do and perhaps just as importantly what they don’t.

The Panton Principles aim to articulate a view of what best practice should be with respect to data publication for science. They arose out of an ongoing conversation between myself Peter Murray-Rust and Rufus Pollock. Rufus founded the Open Knowledge Foundation, an organisation that seeks to promote and support open culture, open source, and open science, with the emphasis on the open. The OKF position on licences has always been that share-alike provisions are an acceptable limitation to complete freedom to re-use content. I have always taken the Science Commons position that share-alike provisions, particularly on data have the potential to make it difficult or impossible to get multiple datasets or systems to interoperate. In another post I will explore this disagreement which really amounts to a different perspective on the balance of the risks and consequences of theft vs things not being used or useful. Peter in turn is particularly concerned about the practicalities – really wanting a straightforward set of rules to be baked right into publication mechanisms.

The Principles came out of a discussion in the Panton Arms a pub near to the Chemistry Department of Cambridge University, after I had given a talk in the Unilever Centre for Molecular Informatics. We were having our usual argument trying to win the others over when we actually turned to what we could agree on. What sort of statement could we make that would capture the best parts of both positions with a focus on science and data. We focussed further by trying to draw out one specific issue. Not the issue or when people should share results, or the details of how, but the mechanisms that should be used for re-use. The principles are intended to focus on what happens when a decision has been made to publish data and where we assume that the wish is for that data to be effectively re-used.

Where we found agreement was that for science, and for scientific data, and particularly science funded by public investment, that the public domain was the best approach and that we would all recommend it. We brought John Wilbanks in both to bring the views of Creative Commons and to help craft the words. It also made a good excuse to return to the pub. We couldn’t agree on everything – we will never agree on everything – but the form of words chosen – that placing data explicitly, irrevocably, and legally in the public domain satisfies both the Open Knowledge Definition and the Science Commons Principles for Open Data was something that we could all personally sign up to.

The end result is something that I have no doubt is imperfect. We have borrowed inspiration from the Budapest Declaration, but there are three B’s. Perhaps it will take three P’s to capture all the aspects that we need. I’m certainly up for some meetings in Pisa or Portland, Pittsburgh or Prague (less convinced about Perth but if it works for anyone else it would make my mother happy). For me it captures something that we agree on – a way forwards towards making the best possible practice a common and practical reality. It is something I can sign up to and I hope you will consider doing so as well.

Above all, it is a start.

Reblog this post [with Zemanta]

Friendfeed for Research? First impressions of ScienceFeed

Image representing FriendFeed as depicted in C...
Image via CrunchBase

I have been saying for quite some time that I think Friendfeed offers a unique combination of functionality that seems to work well for scientists, researchers, and the people they want to (or should want to) have conversations with. For me the core of this functionality lies in two places: first that the system explicitly supports conversations that centre around objects. This is different to Twitter which supports conversations but doesn’t centre them around the object – it is actually not trivial to find all the tweets about a given paper for instance. Facebook now has similar functionality but it is much more often used to have pure conversation. Facebook is a tool mainly used for person to person interactions, it is user- or person-centric. Friendfeed, at least as it is used in my space is object-centric, and this is the key aspect in which “social networks for science” need to differ from the consumer offerings in my opinion. This idea can trace a fairly direct lineage via Deepak Singh to the Jeff Jonas/Jon Udell concatenation of soundbites:

“Data finds data…then people find people”

The second key aspect about Friendfeed is that it gives the user a great deal of control over what they present to represent themselves. If we accept the idea that researchers want to interact with other researchers around research objects then it follows that the objects that you choose to represent yourself is crucial to creating your online persona. I choose not to push Twitter into Friendfeed mainly because my tweets are directed at a somewhat different audience. I do choose to bring in video, slides, blog posts, papers, and other aspects of my work life. Others might choose to include Flickr but not YouTube. Flexibility is key because you are building an online presence. Most of the frustration I see with online social tools and their use by researchers centres around a lack of control in which content goes where and when.

So as an advocate of Friendfeed as a template for tools for scientists it is very interesting to see how that template might be applied to tools built with researchers in mind. ScienceFeed launched yesterday by Ijad Madisch, the person behind ResearchGate. The first thing to say is that this is an out and out clone of Friendfeed, from the position of the buttons to the overall layout. It seems not to be built on the Tornado server that was open sourced by the Friendfeed team so questions may hang over scalability and architecture but that remains to be tested. The main UI difference with Friendfeed is that the influence of another 18 months of development of social infrastructure is evident in the use of OAuth to rapidly leverage existing networks and information on Friendfeed, Twitter, and Facebook. Although it still requires some profile setup, this is good to see. It falls short of the kind of true federation which we might hope to see in the future but then so does everything else.

In terms of specific functionality for scientists the main additions is a specialised tool for adding content via a search of literature databases. This seems to be adapted from the ResearchGate tool for populating a profile’s publication list. A welcome addition and certainly real tools for researchers must treat publications as first class objects. But not groundbreaking.

The real limitation of ScienceFeed is that it seems to miss the point of what Friendfeed is about. There is currently no mechanism for bringing in and aggregating diverse streams of content automatically. It is nice to be able to manually share items in my citeulike library but this needs to happen automatically. My blog posts need to come in as do my slideshows on slideshare, my preprints on Nature Precedings or Arxiv. Most of this information is accessible via RSS feeds so import via RSS/Atom (and in the future real time protocols like XMPP) is an absolute requirement. Without this functionality, ScienceFeed is just a souped up microblogging service. And as was pointed out yesterday in one friendfeed thread we have a twitter-like service for scientists. It’s called Twitter. With the functionality of automatic feed aggregation Friendfeed can become a presentation of yourself as a researcher on the web. An automated publication list that is always up to date and always contains your latest (public) thoughts, ideas, and content. In short your web-native business card and CV all rolled into one.

Finally there is the problem of the name. I was very careful at the top of this post to be inclusive in the scope of people who I think can benefit from Friendfeed. One of the great strengths of Friendfeed is that it has promoted conversations across boundaries that are traditionally very hard to bridge. The ongoing collision between the library and scientific communities on Friendfeed may rank one day as its most important achievement, at least in the research space. I wonder whether the conversations that have sparked there would have happened at all without the open scope that allowed communities to form without prejudice as to where they came from and then to find each other and mingle. There is nothing in ScienceFeed that precludes anyone from joining as far as I can see, but the name is potentially exclusionary, and I think unfortunate.

Overall I think ScienceFeed is a good discussion point, a foil to critical thinking, and potentially a valuable fall back position if Friendfeed does go under. It is a place where the wider research community could have a stronger voice about development direction and an opportunity to argue more effectively for business models that can provide confidence in a long term future. I think it currently falls far short of being a useful tool but there is the potential to use it as a spur to build something better. That might be ScienceFeed v2 or it might be an entirely different service. In a follow-up post I will make some suggestions about what such a service might look like but for now I’d be interested in what other people think.

Other Friendfeed threads are here and here and Techcrunch has also written up the launch.

Reblog this post [with Zemanta]

Science Commons Symposium – Redmond 20th February

Science Commons
Image by dullhunk via Flickr

One of the great things about being invited to speak that people don’t often emphasise is that it gives you space and time to hear other people speak. And sometimes someone puts together a programme that means you just have to shift the rest of the world around to make sure you can get there. Lisa Green and Hope Leman have put together the biggest concentration of speakers in the Open Science space that I think I have ever seen for the Science Commons Symposium – Pacific Northwest to be held on the Microsoft Campus in Redmond on 20 February. If you are in the Seattle area and have an interest in the future of science, whether pro- or anti- the “open” movement, or just want to hear some great talks you should be there. If you can’t be there then watch out for the video stream.

Along with me you’ll get Jean-Claude Bradley, Antony Williams, Peter Murray-Rust, Heather Joseph, Stephen Friend, Peter Binfield, and John Wilbanks. Everything from policy to publication, software development to bench work, and from capturing the work of a single researcher to the challenges of placing several hundred millions dollars worth of drug discovery data into the public domain. All with a focus on how we make more science available and generate more and innovative. Not to be missed, in person or online – and if that sounds too much like self promotion then feel free to miss the first talk… ;-)

Reblog this post [with Zemanta]

Peer review: What is it good for?

Peer Review Monster
Image by Gideon Burton via Flickr

It hasn’t been a real good week for peer review. In the same week that the Lancet fully retract the original Wakefield MMR article (while keeping the retraction behind a login screen – way to go there on public understanding of science), the main stream media went to town on the report of 14 stem cell scientists writing an open letter making the claim that peer review in that area was being dominated by a small group of people blocking the publication of innovative work. I don’t have the information to actually comment on the substance of either issue but I do want to reflect on what this tells us about the state of peer review.

There remains much reverence of the traditional process of peer review. I may be over interpreting the tenor of Andrew Morrison’s editorial in BioEssays but it seems to me that he is saying, as many others have over the years “if we could just have the rigour of traditional peer review with the ease of publication of the web then all our problems would be solved”.  Scientists worship at the altar of peer review, and I use that metaphor deliberately because it is rarely if ever questioned. Somehow the process of peer review is supposed to sprinkle some sort of magical dust over a text which makes it “scientific” or “worthy”, yet while we quibble over details of managing the process, or complain that we don’t get paid for it, rarely is the fundamental basis on which we decide whether science is formally published examined in detail.

There is a good reason for this. THE EMPEROR HAS NO CLOTHES! [sorry, had to get that off my chest]. The evidence that peer review as traditionally practiced is of any value at all is equivocal at best (Science 214, 881; 1981, J Clinical Epidemiology 50, 1189; 1998, Brain 123, 1954; 2000, Learned Publishing 22, 117; 2009). It’s not even really negative. That would at least be useful. There are a few studies that suggest peer review is somewhat better than throwing a dice and a bunch that say it is much the same. It is at its best at dealing with narrow technical questions, and at its worst at determining “importance” is perhaps the best we might say. Which for anyone who has tried to get published in a top journal or written a grant proposal ought to be deeply troubling. Professional editorial decisions may in fact be more reliable, something that Philip Campbell hints at in his response to questions about the open letter [BBC article]:

Our editors […] have always used their own judgement in what we publish. We have not infrequently overruled two or even three sceptical referees and published a paper.

But there is perhaps an even more important procedural issue around peer review. Whatever value it might have we largely throw away. Few journals make referee’s reports available, virtually none track the changes made in response to referee’s comments enabling a reader to make their own judgement as to whether a paper was improved or made worse. Referees get no public credit for good work, and no public opprobrium for poor or even malicious work. And in most cases a paper rejected from one journal starts completely afresh when submitted to a new journal, the work of the previous referees simply thrown out of the window.

Much of the commentary around the open letter has suggested that the peer review process should be made public. But only for published papers. This goes nowhere near far enough. One of the key points where we lose value is in the transfer from one journal to another. The authors lose out because they’ve lost their priority date (in the worse case giving the malicious referees the chance to get their paper in first). The referees miss out because their work is rendered worthless. Even the journals are losing an opportunity to demonstrate the high standards they apply in terms of quality and rigor – and indeed the high expectations they have of their referees.

We never ask what the cost of not publishing a paper is or what the cost of delaying publication could be. Eric Weinstein provides the most sophisticated view of this that I have come across and I recommend watching his talk at Science in the 21st Century from a few years back. There is a direct cost to rejecting papers, both in the time of referees and the time of editors, as well as the time required for authors to reformat and resubmit. But the bigger problem is the opportunity cost – how much that might have been useful, or even important, is never published? And how much is research held back by delays in publication? How many follow up studies not done, how many leads not followed up, and perhaps most importantly how many projects not refunded, or only funded once the carefully built up expertise in the form of research workers is lost?

Rejecting a paper is like gambling in a game where you can only win. There are no real downside risks for either editors or referees in rejecting papers. There are downsides, as described above, and those carry real costs, but those are never borne by the people who make or contribute to the decision. Its as though it were a futures market where you can only lose if you go long, never if you go short on a stock. In Eric’s terminology those costs need to be carried, we need to require that referees and editors who “go short” on a paper or grant are required to unwind their position if they get it wrong. This is the only way we can price in the downside risks into the process. If we want open peer review, indeed if we want peer review in its traditional form, along with the caveats, costs and problems, then the most important advance would be to have it for unpublished papers.

Journals need to acknowledge the papers they’ve rejected, along with dates of submission. Ideally all referees reports should be made public, or at least re-usable by the authors. If full publication, of either the submitted form of the paper or the referees report is not acceptable then journals could publish a hash of the submitted document and reports against a local key enabling the authors to demonstrate submission date and the provenance of referees reports as they take them to another journal.

In my view referees need to be held accountable for the quality of their work. If we value this work we should also value and publicly laud good examples. And conversely poor work should be criticised. Any scientist has received reviews that are, if not malicious, then incompetent. And even if we struggle to admit it to others we can usually tell the difference between critical, but constructive (if sometimes brutal), and nonsense. Most of us would even admit that we don’t always do as good a job as we would like. After all, why should we work hard at it? No credit, no consequences, why would you bother? It might be argued that if you put poor work in you can’t expect good work back out when your own papers and grants get refereed. This again may be true, but only in the long run, and only if there are active and public pressures to raise quality. None of which I have seen.

Traditional peer review is hideously expensive. And currently there is little or no pressure on its contributors or managers to provide good value for money. It is also unsustainable at its current level. My solution to this is to radically cut the number of peer reviewed papers probably by 90-95% leaving the rest to be published as either pure data or pre-prints. But the whole industry is addicted to traditional peer reviewed publications, from the funders who can’t quite figure out how else to measure research outputs, to the researchers and their institutions who need them for promotion, to the publishers (both OA and toll access) and metrics providers who both feed the addiction and feed off it.

So that leaves those who hold the purse strings, the funders, with a responsibility to pursue a value for money agenda. A good place to start would be a serious critical analysis of the costs and benefits of peer review.

Addition after the fact: Pointed out in the comments that there are other posts/papers I should have referred to where people have raised similar ideas and issues. In particular Martin Fenner’s post at Nature Network. The comments are particularly good as an expert analysis of the usefulness of the kind of “value for money” critique I have made. Also a paper in the Arxiv from Stefano Allesina. Feel free to mention others and I will add them here.

Reblog this post [with Zemanta]

Everything I know about software design I learned from Greg Wilson – and so should your students

Visualization of the "history tree" ...
Image via Wikipedia

Which is not to say that I am any good at software engineering, good practice, or writing decent code. And you shouldn’t take Greg to task for some of the dodgy demos I’ve done over the past few months either. What he does need to take the credit for is enabling me to go from someone who knew nothing at all about software design, the management of software development or testing to being able to talk about these things, ask some of the right questions, and even begin to make some of my own judgements about code quality in an amazingly short period of time. From someone who didn’t know how to execute a python script to someone who feels uncomfortable working with services where I can’t use a testing framework before deploying software.

This was possible through the online component of the training programme, called Software Carpentry, that Greg has been building, delivering and developing over the past decade. This isn’t a course in software engineering and it isn’t built for computer science undergraduates. It is a course focussed on taking scientists who have done a little bit of tinkering or scripting and giving them the tools, the literacy, and the knowledge to apply the best of knowledge base of software engineering to building useful high quality code that solves their problems.

Code and computational quality has never been a priority in science and there is a strong argument that we are currently paying, and will continue to pay a heavy price for that unless we sort out the fundamentals of computational literacy and practices as these tools become ubiquitous across the whole spread of scientific disciplines. We teach people how to write up an experiment; but we don’t teach them how to document code. We teach people the importance of significant figures but many computational scientists have never even heard of version control. And we teach the importance of proper experimental controls but never provide the basic training in testing and validating software.

Greg is seeking support to enable him to update Software Carpentry to provide an online resource for the effective training of scientists in basic computational literacy. It won’t cost very much money; we’re talking a few hundred thousand dollars here. And the impact is potentially both important and large. If you care about the training of computational scientists; not computer scientists, but the people who need, or could benefit from, some coding, data managements, or processing in their day to day scientific work, and you have money then I encourage you to contribute. If you know people or organizations with money please encourage them to contribute. Like everything important, especially anything to do with education and preparing for the future, these things are tough to fund.

You can find Greg at his blog: http://pyre.third-bit.com

His description of what wants to do and what he needs to do it is at: http://pyre.third-bit.com/blog/archives/3400.html

Reblog this post [with Zemanta]

Why I am disappointed with Nature Communications

Towards the end of last year I wrote up some initial reactions to the announcement of Nature Communications and the communications team at NPG were kind enough to do a Q&A to look at some of the issues and concerns I raised. Specifically I was concerned about two things. The licence that would be used for the “Open Access” option and the way that journal would be positioned in terms of “quality”, particularly as it related to the other NPG journals and the approach to peer review.

Unfortunately I have to say that I feel these have been fudged, and this is unfortunate because there was a real opportunity here to do something different and quite exciting.  I get the impression that that may even have been the original intention. But from my perspective what has resulted is a poor compromise between my hopes and commercial concerns.

At the centre of my problem is the use of a Creative Commons Attribution Non-commercial licence for the “Open Access” option. This doesn’t qualify under the BBB declarations on Open Access publication and it doesn’t qualify for the SPARC seal for Open Access. But does this really matter or is it just a side issue for a bunch of hard core zealots? After all if people can see it that’s a good start isn’t it? Well yes, it is a good start but non-commercial terms raise serious problems. Putting aside the fact that there is an argument that universities are commercial entities and therefore can’t legitimately use content with non-commercial licences the problem is that NC terms limit the ability of people to create new business models that re-use content and are capable of scaling.

We need these business models because the current model of scholarly publication is simply unaffordable. The argument is often made that if you are unsure whether you are allowed to use content then you can just ask, but this simply doesn’t scale. And lets be clear about some of the things that NC means you’re not licensed for: using a paper for commercially funded research even within a university, using the content of paper to support a grant application, using the paper to judge a patent application, using a paper to assess the viability of a business idea…the list goes on and on. Yes you can ask if you’re not sure, but asking each and every time does not scale. This is the central point of the BBB declarations. For scientific communication to scale it must allow the free movement and re-use of content.

Now if this were coming from any old toll access publisher I would just roll my eyes and move on, but NPG sets itself up to be judged by a higher standard. NPG is a privately held company, not beholden to share holders. It is a company that states that it is committed to advancing scientific communication not simply traditional publication. Non-commercial licences do not do this. From the Q&A:

Q: Would you accept that a CC-BY-NC(ND) licence does not qualify as Open Access under the terms of the Budapest and Bethesda Declarations because it limits the fields and types of re-use?

A: Yes, we do accept that. But we believe that we are offering authors and their funders the choices they require.Our licensing terms enable authors to comply with, or exceed, the public access mandates of all major funders.

NPG is offering the minimum that allows compliance. Not what will most effectively advance scientific communication. Again, I would expect this of a shareholder-controlled profit-driven toll access dead tree publisher but I am holding NPG to a higher standard. Even so there is a legitimate argument to be made that non-commercial licences are needed to make sure that NPG can continue to support these and other activities. This is why I asked in the Q&A whether NPG made significant money off re-licensing of content for commercial purposes. This is a discussion we could have on the substance – the balance between a commercial entity providing a valuable service and the necessary limitations we might accept as the price of ensuring the continued provision of that service. It is a value for money judgement. But not one we can make without a clear view of the costs and benefits.

So I’m calling NPG on this one. Make a case for why non-commercial licences are necessary or even beneficial, not why they are acceptable. They damage scientific communication, they create unnecessary confusion about rights, and more importantly they damage the development of new business models to support scientific communication. Explain why it is commercially necessary for the development of these new activities, or roll it back, and take a lead on driving the development of science communication forward. Don’t take the kind of small steps we expect from other, more traditional, publishers. Above all, lets have that discussion. What is the price we would have to pay to change the license terms?

Because I think it goes deeper. I think that NPG are actually limiting their potential income by focussing on the protection of their income from legacy forms of commercial re-use. They could make more money off this content by growing the pie than by protecting their piece of a specific income stream. It goes to the heart of a misunderstanding about how to effectively exploit content on the web. There is money to be made through re-packaging content for new purposes. The content is obviously key but the real value offering is the Nature brand. Which is much better protected as a trademark than through licensing. Others could re-package and sell on the content but they can never put the Nature brand on it.

By making the material available for commercial re-use NPG would help to expand a high value market for re-packaged content which they would be poised to dominate. Sure, if you’re a business you could print off your OA Nature articles and put them on the coffee table, but if you want to present them to investors you want that Nature logo and Nature packaging that you can only get from one place.  And that NPG does damn well. NPG often makes the case that it adds value through selection, presentation, and aggregation. It is the editorial brand that is of value. Let’s see that demonstrated though monetization of the brand, rather than through unnecessarily restricting the re-use of the content, especially where authors are being charged $5000 to cover the editorial costs.

Reblog this post [with Zemanta]

New Year – New me

FireworksApologies for any wierdness in your feed readers. The following is the reason why as I try to get things working properly again.

The past two years on this blog I wrote made some New Year’s resolutions and last year I assessed my performance against the previous year’s aims. This year I will admit to simply being a bit depressed about how much I achieved in real terms and how effective I’ve been at getting ideas out and projects off the ground. This year I want to do more in terms of walking the walk, creating examples, or at least lashups of the things I think are important.

One thing that has been going around in my head for at least 12 months is the question of identity. How I control what I present, who I depend on, and in the world of a semantic web where I am represented by a URL what should actually be there when someone goes to that address. So the positive thing I did over the holiday break, rather than write a new set of resolutions was to start setting up my own presence on the web, to think about what I might want to put there and what it might look like.

This process is not as far along as I would like but its far enough along that this will be the last post at this address. OpenWetWare has been an amazing resource for me over the past several years and we will continue to use the wiki for laboratory information and I hope to work with the team in whatever way I can as the next generation of tools develops. OpenWetWare was also a safe place where I could learn about blogging without worrying about the mechanics, confident in the knowledge that Bill Flanagan was covering the backstops. Bill is the person who has kept things running through the various technical ups and down and I’d particularly like to thank him for all his help.

However I have now learnt enough to be dangerous and want to try some more things out on my own. More than can be conveniently managed on a website that someone else has to look after. I will write a bit more about the idea and choices I’ve made in setting up the site soon but for the moment I just want to point you to the new site and offer you some choices about subscribing to different feeds.

If you are on the feedburner feed for the blog you should be automatically transferred over to the feed on the new site. If you’re reading in a feed reader you can check this by just clicking through to the item on my site. If you end up at a url starting https://cameronneylon.net/ then you are in the right place. If not, just change your reader to point at http://feeds.feedburner.com/ScienceInTheOpen.

This feed will include posts on things like papers and presentations as well as blog posts so if you are already getting that content in another stream and prefer to just get the blog posts via RSS you should point your reader at http://feeds.feedburner.com/ScienceInTheOpen_blog.  I can’t test this until I actually post something so just hold tight if it doesn’t work and I will try to get it working as soon as I can. The comments feed for all seven of you subscribed to it should keep working. All the posts are mirrored on the new site and will continue to be available at OpenWetWare

Once again I’d like to thank all the people at OpenWetWare that got me going in the blogging game and hope to see you over at the new site as I figure out what it means to present yourself as a scientist on the web.

Reblog this post [with Zemanta]

Google Wave: Ripple or Tsunami?

Big Wave Surfing in Tahiti at Teahupoo
Image by thelastminute via Flickr

A talk given at the Edinburgh University IT Futures meeting late in 2009. The talk discusses the strengths and weaknesses of Wave as a tool for research and provides some pointers on how to think about using it in an academic setting. The talk was recorded in a Wave with members of the audience taking notes around images of the slides which I had previously uploaded.

You will only be able to see the wave if you have a Wave preview account and are logged in. If you don’t have an account the slides are included below (or will be as soon as I can get slideshare to talk to me).

[wave id=”googlewave.com!w+-c2g1ggkA”]

Reblog this post [with Zemanta]