Skip to main content
  1. Contributors/

Prof. Dr. Michael Hanke

Principal investigator Research data committee member

Forschungszentrum Jülich; Heinrich Heine University Düsseldorf

0000-0001-6398-6370

Michael Hanke

Michael Hanke is a professor at the Heinrich Heine University Düsseldorf, and head of the Psychoinformatics group in the Institute for Neuroscience and Medicine (INM-7) at the Forschungszentrum Jülich. He has co-created several neuroinformatics soft ware projects, among them the NeuroDebian, PyMVPA, and DataLad.

Related content

Projects


Q02: Data management for computational modelling

Data management and training platform. A decentralized data management infrastructure will help focus on developmental and therapeutic longitudinal data, training all participating researchers in the necessary skills for future use. This strategy will lay the foundations for further data-driven computational modelling projects in the next funding period. This is a distributed project, with representatives at all main TRR379 sites. As a key software solution, this project employs DataLad. DataLad is a data management software designed to facilitate the organization, sharing, and reproducibility of scientific datasets. It integrates version control with data handling, allowing researchers to track changes, collaborate efficiently, and ensure the accessibility and integrity of their data. By leveraging tools like Git and Git- annex, DataLad provides a streamlined way to manage large datasets, making it particularly valuable in fields like neuroimaging and bioinformatics.

Publications


Lab in a box: A build-your- own-open-lab software toolkit

Over the past two years, our team has been working on an interoperable software toolstack that is open source, self-hosted, and covers basic relevant needs of a computational neuroscience lab.Notably, a number of software solutions came into existance or were deployed or further developed thanks to interactions of different software communities during RDM workshops or the distribits conference for distributed data management technologies (distribits.live).Our objective is to design an approach that allows storing data of arbitrary size, flexible semantic meta data, and the relations between these data; and to provide ways to query those relations and access the underlying data, as well as exposing selected data for websites, knowledge bases, or data catalogs.The system components are either fully compatible or integrated with the DataLad (Halchenko et al., 2021) ecosystem for data management.At the core of the stack, we have developed the following software components:

DataLad: distributed system for joint management of code, data, and their relationship

DataLad is a Python-based tool for the joint management of code, data, and their relationship, built on top of a versatile system for data logistics (git-annex) and the most popular distributed version control system (Git). It adapts principles of open-source software development and distribution to address the technical challenges of data management, data sharing, and digital provenance collection across the life cycle of digital objects. DataLad aims to make data management as easy as managing code. It streamlines procedures to consume, publish, and update data, for data of any size or type, and to link them as precisely versioned, lightweight dependencies. DataLad helps to make science more reproducible and FAIR (Wilkinson et al., 2016). It can capture complete and actionable process provenance of data transformations to enable automatic re-computation. The DataLad project (datalad.org) delivers a completely open, pioneering platform for flexible decentralized research data management (RDM) (Hanke, Pestilli, et al., 2021). It features a Python and a command-line interface, an extensible architecture, and does not depend on any centralized services but facilitates interoperability with a plurality of existing tools and services. In order to maximize its utility and target audience, DataLad is available for all major operating systems, and can be integrated into established workflows and environments with minimal friction.

Sites


Research Center Jülich (FZJ)

Forschungszentrum Jülich (FZJ) is a German national research institution that pursues interdisciplinary research in the fields of energy, information, and bioeconomy. It operates a broad range of research infrastructures like supercomputers, an atmospheric simulation chamber, electron microscopes, a particle accelerator, cleanrooms for nanotechnology, among other things. As a member of the Helmholtz Association with roughly 6,800 employees in ten institutes and 80 subinstitutes, Jülich is one of the largest research institutions in Europe.