Showing posts with label data management. Show all posts
Showing posts with label data management. Show all posts

Friday, October 23, 2009

The Future of Data Policy

The Microsoft External Research Division has launched a book entitled, The Fourth Paradigm: Data-Intensive Scientific Discovery (2009) edited by Tony Hey, Stewart Tansley, and Kristin Tolle. The book was launched on the opening day of the Microsoft eScience Workshop that took place in Pittsburgh, USA from 15-17 October 2009. The book includes a chapter, 'The Future of Data Policy' (pp 201-208), authored by Professor Anne Fitzgerald, Professor Brian Fitzgerald and myself. The book is licensed under a Creative Common Attribution Share Alike 3.0 United States licence, and can be download in its entirety or by chapter at The Fourth Paradigm.

Thursday, April 2, 2009

Copyright protection and data compilations

Over on the Digital Curation Blog, Chris Rushbridge has an interesting post entitled, "Are research data facts and does it matter?"

I have just posted a response, which I am reproducing here:

I would like to take this opportunity to explain some of the research we have undertaken in the OAK Law Project and conclusions we have reached regarding copyright protection of data compilations in Australia. We have two primary publications addressing this area: Building the Infrastructure for Data Access and Reuse in Collaborative Research: An Analysis of the Legal Context and Practical Data Management: A Legal and Policy Guide.

s10(1) of the Australian Copyright Act 1968 defines a literary work to include a "compilation". This is where protection for data compilations under Australian law derives from. Any data that is collected, arranged, organised and presented in a logical fashion will usually be regarded as a compilation.

Chris makes a good point that many data compilations will require a great deal of effort, analysis and creativity. In the US, creativity is a requirement before a data compilation can be protected by copyright. In Australia, creativity is not required. Only that the compilation is a result of the exercise of skill, knowledge or judgment in the arrangement of the data, or the investment of substantial labour or expense in collection the material (Desktop Marketing v Telstra).

It can often be difficult to tell whether a compilation is one that would attract copyright protection. In our work, we have tended to err on the side of caution and assume that most compilations will attract copyright protection. This is because the threshold in Australia is so low. The main case in this area, Desktop Marketing v Telstra, involved the copying of a telephone directory. A telephone directory is merely a compilation of names and numbers listed in alphabetical order. If this is a compilation that attracts copyright, then most other compilations are likely to be protected by copyright under Australian law as well.

Copyright law does not protect mere facts or information. Rather, it protects the expression of facts or information in a material form. This means that generally there would not be a problem with copying some of the basic facts contained in a compilation. For example, if I were to list the names and numbers of a small collection of my colleagues on my website, that would not usually be a problem. I have extracted the data that I need, in a fairly "random" fashion (in that I have not just copied a few pages of names and numbers in alphabetical order directly from the White Pages). I have not copied the way that the data is arranged in the telephone directory (the “expression”).

In regards to Science Commons’ decision to discontinue advocating the application of Creative Commons licences to data compilations, my understanding is that they came to this decision for two reasons:
(1) It was not always clear in the US whether the relevant compilation attracted copyright. If it did not but a person had put a CC licence on the compilation in the mistaken belief that it did, then restrictions would have been imposed on that dataset (e.g. that it could only be used non-commercially) which actually had no legal basis for being imposed; and
(2) CC licences all contain an attribution requirement and Science Commons were concerned about what they call "attribution stacking" - i.e. where a dataset is compiled from data contributed by many different researchers, it would be extremely difficult for a user to attribute all of those researchers.

At OAK Law, we still believe that CC licences can be applied to datasets in Australia because the concerns noted by Science Commons do not arise to the same degree in Australia. Firstly, we have a lower threshold test for copyright protection, meaning that copyright will more readily attach to datasets in Australia and the first problem noted by Science Commons is less likely to occur. Nevertheless, to be sure, we usually advocate that the widest CC licence - the attribution only licence - be applied to datasets. Secondly, unlike in the US, Australian copyright law includes Moral Rights, meaning that creators have to be attributed anyway, regardless of whether a CC licence is applied or not. We think there are various ways of getting around the "attribution stacking" problem - for example, a group of researchers could agree on a common way to be attributed (e.g. we could be attributed as "the OAK Law Project"), or the data could be attributed using a URL, which an interested party can visit and which can list all the contributors (and this list can be added to over time). The advantage of applying CC licences to data, in our view, is that it provides some certainty to users about what they can and cannot do with that data.

If you are interested in this issue, I would also advise reading these posts by Robin Rice and Rufus Pollock.

Friday, February 13, 2009

Access to Victorian fire data

Yesterday afternoon an interesting story appeared on ZDNet Australia: “Vic Govt limited Google’s bushfire map”. I encourage you to read the full post on ZDNet Australia, but in summary, the post documents Google’s trouble in gaining access to Victorian Government data about the movement of bushfires in Victoria.

According to the post, Google has been working with the Commonwealth Fire Authority, which manages fires on private lands, to overlay the Authority’s data onto Google Maps to produce a real-time map of the locations of the fires. The map also uses a colour scheme to convey the seriousness of the fires: green (safe), yellow (controlled), orange (contained) and red (ongoing).

Naturally, this map is immensely beneficial to those in Victoria and elsewhere who are attempting to track the bushfires.

However, Google has run into some problems gaining access to data to plot fires on public lands. This data is owned and controlled by the Victorian Department of Sustainability and Environment, and is covered by Crown copyright. As such, permission is required from the government before the data can be used, and for Google this permission has not been forthcoming. The result is that Google has been unable to plot this data onto their map.

As noted in the ZDNet Australia post, this is not the first time Google has had trouble accessing and using Australian government data. They were expressly denied permission from the Commonwealth Department of Health and Aging to overlay data from the National Public Toilet Map onto a Google Map.

Why is the government so unwilling to share its data? My guess is that there are two possible reasons. The first is that in some cases, the government has a misguided idea that data can be used to build online systems or services (usually these will be geospatial systems or services) which can be used to generate revenue by charging for access. The other is that the government is naturally risk-averse and would prefer to control their data as tightly as possible.

What the government is forgetting is that it is a representative of the people and the government-owned data has been collected using public funds. We, the Australian public, have paid for that data through our taxes and as such, we should have the benefit of that data. Surely it is most beneficial for the public if we can have ready access to that data in the most efficient and convenient way possible. And if that is through a Google Map, then the government should enable this. There can be no argument that in the face of tragedy such as the Victorian bushfires, the government should not hinder our ability to access as much information as possible about that tragedy. This includes the ability to easily track those bushfires via a Google Map.

Arguments have been made that as the access and use issue can be traced back to Crown copyright, then Crown copyright should be removed, as is the case in the United States where government data and publications are held to be in the public domain. I do not believe that this is the answer. Rather than remove Crown copyright completely, the government should be encouraged to release their material where possible under open licences such as the Creative Commons Attribution licence. This should be the default position, unless access to the material must be restricted due to privacy or national security concerns. The government must engage in a “push” model – where it systematically “pushes” its material out to the community – rather than a “pull” model – where members of the public must seek permission or lodge a Freedom Of Information request to access that material. Crown copyright can serve an important purpose, if only through the operation of the requirement of attribution (a requirement imposed through the Creative Commons licence, similar to moral rights), which requires that the author of a material (in this case, the government) to be attributed wherever the material is reproduced. The requirement of attribution for government copyright material can serve a two-fold purpose – (1) it allows the government to retain some control over the material it produces; and (2) it verifies to the public that the material has come from a reliable source.

Our research group at QUT has done some work on this area. See the auPSI website for more information.

Wednesday, October 15, 2008

ARROW Repository Day

On 14 October 2008, I attended the ARROW Repository Day held in Customs House in Brisbane. I presented on the legal issues surrounding management of data for inclusion in a repository. You can access my slides here.

Chris Rusbridge of the Digital Curation Centre in the UK also presented. Some brief notes from his talk are below. Chris was live blogging the day, so if you are interested I suggest you read his notes at the Digital Curation Blog.


Chris Rusbridge (Digital Curation Centre) – Moving the repository upstream

The resistant scholar
  • Uncertainty, risk - about copyright; about Ingelfinger Rule
  • Change
  • Too busy
  • Doesn’t fit into the way they do things now
  • Not well motivated by advantages to others
  • Little in it for them!

Research workflow
  • many different tasks in parallel
  • all different stages
  • teaching (several), research (several), writing up research, writing grant proposals, reviewing papers, administrative tasks etc

On negative clicks

Asked - how many extra clicks are you willing to make to ensure preservation of your record?

Answer - zero

Negative click repository?


Can the repository help rather than hinder?
Towards a Research Repository System? [diagram]

Maybe we could…
  • help with publisher liaison
  • support multiple authoring across several institutions
  • more permissive identity management
  • support multiple versions
  • fine grained access control
  • checkpointing
  • support supplementary data
  • provide basic data management capability
  • provide simple, cross-platform, persistent storage
  • provide some longevity
  • provide additional benefits

Friday, October 3, 2008

ANDS Workshop at eResearch Australasia Conference

On Thursday 2 September, I attended the Australian National Data Service (ANDS) Workshop at the eResearch Australasia Conference 2008. This was a full day workshop, but the ANDS team did a great job of keeping the workshop interesting and highly interactive, and the day went very quickly.

In the morning, there were a few brief presentations – notably from Andrew Treloar of Monash University and the ANDS Establishment Project and Tracey Hinds from CSIRO. I particularly enjoyed Tracey’s presentation, which at a conference that seemed dominated by IT issues, focused on the social issues and the governance issues involved in data management and sharing research data. My notes from Tracey’s talk are below.

The rest of the day was spent in small round-table discussions. The most lively discussion surrounded questions about what institutions and research bodies need to help them in managing and sharing their data, and how ANDS could help. The group found that there was a need for:
  • an openly accessible registry of ontologies for metadata of datasets, so that institutions can start using common and enduring metadata to describe their data;
  • training for researchers, repository managers, research management staff, librarians, archivists and IT staff about data management (including the legal issues surrounding data management), database/repository infrastructure (how to make the database easy to use and sustainable), open access (why should you share your data?) and metadata. It was agreed that the training materials might have a generic introduction component that could be used by all groups, but then there should be different kinds of training materials that provide relevant detail to different groups (e.g. research management staff will have different concerns to IT staff; science researchers may have different concerns humanities researchers);
  • developing conventions for the citation of data, so that researchers can get credit for sharing their data; and
  • proper and comprehensive data management plans (DMP).

There was a consensus that data management plans were particularly important and that it would be useful to develop template DMPs which included specific sections that could be added or deleted as appropriate (for example, a section about compliance with privacy laws might be relevant to medical research but not to astronomy research). It was also thought that ANDS could select a few research projects from different disciplines and assist these projects in formulating a DMP. The resulting DMPs could then be made available online for other projects to use and adapt.

In relation to ANDS selecting particular projects to assist, in a broader way, with their data management and release (“engagement targets”) in the hope that these projects might then appear as “exemplar projects” for other groups, it was considered that appropriate selection criteria might be:
  • broadness of audience and impact;
  • potential for reuse of data and the ongoing reusability/sustainability of the data;
  • the project’s willingness to assist others to develop their data management skills;
  • wide inter-disciplinary appeal;
  • willingness to transfer data around; and
  • projects which will have good exemplary value to attract other communities.


I believe that ANDS will make the notes taken from the workshop available online.


Here are my notes from Tracey’s talk:

Tracey Hind – CSIRO
  • ownership of data should stay with researcher
  • but still need to manage CSIRO’s data at a higher level – maybe provide an “enabling” service for this rather than dictate a “one size fits all” approach
  • As of now, CSIRO still does not formally recognise the idea of data management
  • Real challenges are not technology – it is the human factors – issues of acceptance, understanding, people being prepared to share their data, IP etc
  • High demand for storage, but storage is not management
  • Scientists are not working as well across disciplines as the Flagship vision as hoped, much of this is because “you don’t know what you don’t know” – and it’s hard getting insight into other research disciplines
  • Making data easily discoverable is the key to achieving multi-disciplinary outcomes
  • Lesson is that data is a complex issue – especially when researchers don’t understand the potential benefits – you need exemplar projects to demonstrate the benefits of data management to get buy in.
  • CSIRO’s data management vision (eSIM) – CSIRO scientists will be able to…gather, analyse and share scientific information securely and efficiently, leading to greater scientific outcomes for Australia
  • Four layers – people, processes, technology and governance
  • People challenges = incentives for deposit into a repository;
  • Processes challenges = making sure that the work flows created actually support the technology and make things easy
  • Governance = making sure all of this is properly funded and that data management is a part of the decision making (i.e. make sure researchers have a DMP before they are awarded funding)
  • CSIRO’s exemplar projects = Auscope project; Atlas of Living Australia; Corporate Communications

Tuesday, September 30, 2008

eResearch Australasia Conference 2008 - Tuesday morning (30 September)

John Wilbanks – Uncommon Knowledge and e-Research

Once again, John Wilbanks gave an informative and dynamic presentation. It was geared towards the audience in attendance here at the eResearch Australasia Conference (who are somewhat more IT and science focused than the audience at the OAR conference last week) and so described in detail many aspects of the NeuroCommons Project. If you are interested, I suggest that you see the Neurocommons website. I don’t think any summary that I could provide here would do the project justice. But here are some notes from the beginning of John’s presentation:

Why “eResearch”?

1. eResearch is a requirement imposed on us by the flood of data
  • the web doesn’t give us the same results for science as it does for culture
  • so what can we do?
  • We can…collaborate
  • Eg - Watson and Crick – their success was composed, by building on a series of blocks of knowledge that were available to them from a range of sources
  • But humans can’t build models to scale anymore
  • We need to utilize digital resources
One way to think about eResearch is that it is about:
  • Finding the right collaborator;
  • making big discoveries;
  • getting credit for one’s work
2. We need to convert what we know into digital formats that support model buildings
  • “the web” – no organising topics – hyperlinking allows us to organise things in a dynamic way
  • all the data and all the ides: building blocks
  • open access attempts to solve the legal problems – giving credit where credit is dues; allows humans to read the papers; allows publicly funded research to be accessed by the public
  • but it doesn’t solve the technical problem of paper-based formats that cannot be read by machines
  • we need to develop machine-searchable formats


Kerstin Lehnert, Columbia University – New Science Communities for Cyberinfrastructure: The Example of Geochemistry

Kerstin described eResearch as a vision to provide a genuine infrastructure of highly reliable, widely accessible ICT capabilities to assist researchers in their work – ultimately about people

She discussed the cultural issues involved in sharing data. She identified data citation (what I would call “attribution”) as a big problem. How can all scientists and contributors be cited? Many want to be attributed personally (not just by a project), but there are so many contributors and this quickly becomes a big and messy problem. This observation reflects the problem that we at the OAK Law and Legal Framework to eResearch Projects identified in assessing whether Creative Commons licences could be applied to data compilations. Attribution is an important condition of the CC licence. Researchers and research projects need to decide and identify (before applying a CC licence) how the data compilation is to be attributed, otherwise users could run into all sorts of problems and confusion.


Jane Hunter (UQ) - National Committee for Data in Science (NCDS)

A committee of the Australian Academy of Science – established in February 2008; member of CODATA

Mission – to promote enduring access to Australia’s scientific data assets in order to drive national research and innovation
And to provide a National Data Science voice
Encourage and facilitation cross-fertilisations, between specific science disciplines and other data generation/management disciplines

Future activities include engaging with Chairs of other national committees, including looking at what role they can play within ANDS (Australian National Data Service) to support their goals.

Monday, August 18, 2008

APSR Workshop – The Data Management Plan: Putting Policy into Practice

On Friday 8 August 2008, I attended the Australian Partnership for Sustainable Repositories (APSR) Workshop, “The Data Management Plan: Putting Policy into Practice” at the University of Melbourne.

Professor Anne Fitzgerald, with whom I work at QUT, gave an excellent and very well received presentation on the legal issues surrounding data management. Her slides can be viewed here.


Here are my notes from the workshop (made roughly during the day):


Data management plans: from idea to reality (10:15am – 10:45am)

Dr Markus Buchhorn (ANU) for Karen Visser
  • We need enduring systems that outlive projects and programs
  • Individuals are human – seven deadly fears:
  1. fear of missed “nuggets” in their data – milk it for everything, for ever and veer
  2. fear of missed errors
  3. fear of unknown custodians/stewards
  4. fear inappropriate leaks (privacy/ethics) – can ruin trust relationships with others
  5. fear the cost of effort
  6. fear lack of recognition
  7. fear trusting someone else's data
Plan ahead – help researchers to help themselves as far as possible
Build relationships of trust with researchers – engage with researchers as early as possible


Mark Euston (ANU – Information Literacy Program)
  • tasked with developing a training course, workshop and online, for early to mid career researchers, on Data Management Plans (DMP)
  • Objectives of the course -
  1. what is Data Management (DM)?
  2. benefits and requirements
  3. raising awareness of DM services
  4. DMP
  • Manual based on Guidance on Data Management (UK) and Guide to Social Science Data Preparation and Archiving
  • get researchers in by stressing how they can work with their data more effectively and efficiently

What's happening at... (11:10am – 12:30pm)

Belinda Weaver (UQ)

Issues for the data survey:
  • no 'joined up' services
  • no help
  • inequity – not fair – nothing works etc.
  • costs
  • lack of training (people felt insecure about what they were doing)
  • uncertainty
  • no incentive, no rewards
Recommendations from focus group:
  • standardised DM template for funding applications
  • legal advice centralised and accessible
  • service focused support teams for research projects – specific to the discipline
  • survey of all existing data
  • central data storage system
  • develop a clear UQ data management policy
  • templates
Central management of research data - issues:
  • trust
  • data integrity
  • accidental disclosure
  • control
  • sharing
  • re-use (want to know what use has been made of their data – auditing – and if they give data to a person for a particular purpose, they want to know if the person doesn't end up using the data or not using it for the particular purpose)
  • the long term
Wish-lists:
  • clear policy and guidelines
  • account manager
  • specialists on teams (want to know who to go to for advice)
  • career path?
  • rewards
  • templates for everything
  • funding to do it properly
  • advice and consultancy
  • institutional support
  • tools (but they want to be told only when they want to be told, and be told how they want to be told)
presentations from workshop available at: http://www.library.uq.edu.au/escholarship/orca.html
UQ developing a expert curation advice service


Lyle Winton (Uni of Melbourne)
  • Uni of Melbourne have a research DMP template
  • looking at training for undergrad students
  • looking at how to keep this up to date
  • possible data management registries
  • from 500 charges of research misconduct, 40% could have been avoided by good data management
http://www.esrc.unimelb.edu.ay/dmp/references.html


Suzanne Clarke (Monash)
  • Monash has a Data Management Committee
  • Research Data Management Toolkit for librarians so they know what to talk about to researchers
  • Identified needs: more education required for researchers on statutory requirements for data, IP and the ownership of research data

Gillian Elliot (University of Otago – NZ)
  • As far as she is aware, NZ has no policies surrounding data management
  • so NZ in quite a different position to Australia
  • Survey in 2007 - researchers in NZ had a lot of data and a lot of stuff loosely stuck together that were unpublished and hard to classify – need help with data management
  • data management and copyright concerned researchers – 48% of survey respondents
  • Atlas of Living Australia; Convention on Biological Diversity; Department of Conservation and Land Information New Zealand; Land Care NZ; National Vegetation Survey Databank

Dr Ashley Buckle (TARDIS – Monash University)
  • TARDIS is a multi-institutional collaborative venture that aims to facilitate the arching and sharing of raw X-ray diffraction images
  • Protein Data Bank – growing exponentially – too much data?
  • Benefits to making raw data available – experiment reproducability/validation


Discussion Groups: Group 2 – Processes for Data Management Planning (1:15pm-2:45pm)

How do we make DM part of the usual research practice?

How can we make raw data count as a citation? - for funding etc. - this is very important, if there is greater recognition of the value data in itself as a citable object then researchers will be more willing to manage their data properly.

Ashley Buckle – we need “data journals” - essentially the same as a database but greater recognition

DM needs to give you a reward at the end that is at the same level as rewards from publication

Better tools – build the researchers tools that are so good that they do not actually realise that they are managing their data.


Reporting back to main group and discussion (2:45pm-4:00pm)
  1. Roles, rights and responsibilities
  • Anne Fitzgerald's domains of responsibility
  • Policy plus principles
  • disseminate research data as widely as possible
  • develop practical toolkits
  • risk management for universities
  • simple for universities to completing
  • ongoing legal and policy advice
  • insert data management requirements into research proposals and grants
  • get recognition via NHMRC, ARC and ERA to provide regulatory and reward structure
  • need for national centre for legal policy and advice in regard to the data lifecycle including reuse
  • universities to incorporate data management into risk management strategies
  • provide pragmatic family of licences/responsibility statements (like CC) to identify roles and policies
  • DMPs to be built into research project formulation and management

2. Processes for data management planningbetter tools and incentives: build better workflows
  • allowing data management in their modelling: harness tools onto repositories
  • citation: make sure that citation of datasets happens and is rewarded, as incentive for researchers to create good data
  • persuade ARC to make explicit expression of intent in ERA eventually to credit data citation (at least down the road). This as formal submission from this workshop
  • infrastructure: development of a COHERENT NATIONAL NETWORK of repositories, emphasis on discipline specific repositories (though institutionally supported) as a centre for research activity

3. Making it work
  • know what you don't know
  • each institution needs to:
  1. identify the needs of its researchers (possible role for ANDS here)
  2. map the available services (needs to happen locally)
  3. strategically target the gaps
  4. identify candidate services to drop to fund this
  • Make it easy
  • provide a visible point of contact for the users
  • not necessarily through one channel only
  • not necessarily a one size fits all solution
  • embed regular formal training in how to use services
  • needs to be as easy to use as “MyFlickBook”
  • outreach, marketing, publicity
  • Start small and scale
  1. seed the service and gradual expand it as understanding grows
  2. start with young researchers and use peer group pressure over tie
  3. get good examples going first to generate some quick wins
  4. use growth in tandem with policy
  • Reward innovators in shared services
  1. provide annual performance incentives for going beyond meeting strategic goals
  2. encourage shared services staff to learn new skills
  3. create new job descriptions for new people in management