How do I combine data collected by different people or organizations?
Combining datasets is rarely as simple as putting them into the same spreadsheet. Data collected by different people, teams, or organizations can use different methods, formats, variables, taxonomies, spatial and temporal scales, units, metadata, and standards.
Even when two datasets appear to contain the same information, the underlying data can be structured very differently. A date might be recorded as month-day-year, day-month-year, year-month-day, or Julian date. Variables may have different names, formats, units, or definitions. Before the datasets can be compared, those differences need to be understood and reconciled.
Location is one of the most common challenges. Coordinates may be recorded as latitude and longitude in different formats or datums. Locations may be identified by site names that differ between people, or by a pin dropped on a map without enough information to know exactly what that location represents. A location can look precise in a dataset while still being difficult to interpret or compare.
Then there is the question of what was actually observed.
Two datasets might both contain observations of the same species, but one may record individual animals while another records groups. One may use scientific names while another uses common names or an older taxonomy. One may record precise locations while another uses sites or regions. Without understanding these differences, combining the data can create misleading results.
Some variables are particularly difficult to standardize because the same field can mean different things to different people. Depth is a good example. Many ocean datasets contain a field called “depth (m),” but three people could interpret that as the depth of the bottom, the depth of the observer, or the depth of the animal. Those are three very different measurements. If the meaning was never recorded, the number itself may not be useful.
Other variables are so inconsistent or poorly documented that they are routinely excluded from analysis. Valuable information can effectively disappear simply because it cannot be interpreted with enough confidence to use.
The goal is not to make different datasets identical. It is to make them interoperable while preserving the context and meaning of each observation.
Start by understanding the differences
Before data can be compared or combined, you need to understand how they were collected and what each record represents. Methods, sampling effort, variables, definitions, units, classifications, taxonomies, locations, dates, and metadata all affect how the data can be used.
This is where standards matter.
When teams adopt common methods and data structures, their observations become more comparable from the point of collection. A method created in eOceans can also be added to the Methods Catalogue, allowing other teams to adopt it and collect interoperable data from the start.
But existing data will rarely follow exactly the same standards. Historical datasets, independently collected observations, and data from different organizations still need to be brought together.
eOceans makes different data interoperable
Bring existing datasets into eOceans through Magic Uploads. eOceans identifies and maps information into your project's data structure, helping reconcile differences in formats, fields, definitions, classifications, taxonomies, and other data standards.
The original context and metadata remain connected to the data, so you can understand where the information came from, how it was collected, and how it was transformed.
That means a dataset does not have to be redesigned before it can become useful. Existing information can be incorporated while preserving the differences that are important for interpreting it.
New data can then be collected using the same project structure or an adopted method, making future observations easier to combine with existing information.
Combine without losing meaning
Different datasets do not need to become the same dataset.
A survey collected by trained scientists does not need to become indistinguishable from a community observation. A regional dataset does not need to be treated as though it has the same spatial resolution as a precise GPS observation.
The same applies to sampling methods. A BRUV survey, underwater visual census, and gillnet survey measure different things in different ways. They should remain identifiable as different methods, but their results can still be analyzed together or displayed separately depending on the question.
That distinction matters because interoperability is not about removing differences. It is about preserving meaningful differences while making the information usable together.
Once those differences are understood, data can be filtered, standardized, grouped, weighted, or analyzed according to the question being asked and the characteristics of the underlying data.
This makes it possible to bring together information that would otherwise remain isolated and use it for broader spatial, temporal, and multivariate analysis.
And this is where interoperability becomes particularly valuable.
The real world does not fit neatly into data silos. Fisheries, pollution, biodiversity, disease, climate, human activity, and environmental conditions can influence one another. A dataset collected for one purpose can become important for another question when it can be connected to other information without losing the context that makes it meaningful.
Build on what already exists
Once datasets are connected, you don't have to start over each time another organization contributes information or another dataset becomes available.
The same project can continue to incorporate new sources, validate them, analyze them, and update its outputs. As the evidence grows, you can ask new questions without rebuilding the entire data workflow from the beginning.
Turn separate datasets into connected evidence without throwing away the context that makes them useful.
Bring your data together. Keep their meaning. Build more evidence.
Combine existing datasets in eOceans, adopt shared methods where appropriate, and keep adding new data as your project grows.
—> Get powered by eOceans®.