Skip to content

Final omnibus updates#66

Open
iannevans wants to merge 2 commits into
ivoa:mainfrom
iannevans:EndJuneUpdates
Open

Final omnibus updates#66
iannevans wants to merge 2 commits into
ivoa:mainfrom
iannevans:EndJuneUpdates

Conversation

@iannevans

Copy link
Copy Markdown
Collaborator

(1) Numerous small wordsmithing changes.
(2) Split messenger and messenger_pdgid as agreed with Semantics.
(3) Moved the extension table summary table to the end of the section that defines the table entries.
(4) SIgnificant rewrites in section 6.

(1) Numerous small wordsmithing changes.
(2) Split messenger and messenger_pdgid as agreed with Semantics.
(3) Moved the extension table summary table to the end of the section that defines the table entries.
(4) SIgnificant rewrites in section 6.

@bkhelifi bkhelifi left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All good for me, thanks a lot @iannevans

@mservillat

Copy link
Copy Markdown
Collaborator

Thanks for the many modifications and hard work to homogenise the text.

What is now section 6 is quite important to the discussion that will come after endorsement, so it should not be discarded. As it is now, it is truncated and limited to DataLink usages, while the concept of "specific tables" is also important to discuss further data discovery in a more generalised way. I am sure we can find a good formulation to express this solution in a paragraph. I however understand that the content of such a specialized table is not yet settled and requires more science cases. The proposition for a specialized response-function table as it is now could thus simply appear as an appendix. All this will then be discussed at the DM WG level.

About dataproduct_type = “measurements”, it is already planned to avoid its usage, so this paragraph can be removed. It is indeed written in the product-type vocabulary: "Because of its lack of specificity, this term should generally be avoided, and new, more precise terms should be introduced instead." (see https://www.ivoa.net/rdf/product-type/2026-01-15/product-type.html#measurements).

@Bonnarel

Copy link
Copy Markdown
Contributor

Thanks for the many modifications and hard work to homogenise the text.

What is now section 6 is quite important to the discussion that will come after endorsement, so it should not be discarded. As it is now, it is truncated and limited to DataLink usages, while the concept of "specific tables" is also important to discuss further data discovery in a more generalised way. I am sure we can find a good formulation to express this solution in a paragraph. I however understand that the content of such a specialized table is not yet settled and requires more science cases. The proposition for a specialized response-function table as it is now could thus simply appear as an appendix. All this will then be discussed at the DM WG level.

About dataproduct_type = “measurements”, it is already planned to avoid its usage, so this paragraph can be removed. It is indeed written in the product-type vocabulary: "Because of its lack of specificity, this term should generally be avoided, and new, more precise terms should be introduced instead." (see https://www.ivoa.net/rdf/product-type/2026-01-15/product-type.html#measurements).

The "access section" was written by Mireille and myself. I mainly concentrated on DataLink subsection and Mireille on the specific description tables but we are both responsible of the whole due to extensive cross-reading before the pull request. I apologize that we were not able to propose it before interop, but it was admitted several weeks before interop that such a section will come and was needed. During the interop there was several announces that such a section was close to be posted and this didn't seem to be an issue to wait for it.

About the content : there is nothing new there. These ideas have been presented several times in emails, issues, Pull request comment or even in the case of Mireille in previous interop presentation. I remember writing an email in April which had almost the same content as the subsection on DataLink of last Mireille's PR. I also presented some High energy examples in my DAL talk about DataLink explicitly acknowledging the discussions in High Energy Interest Group.
In the general discussion about extensions at last interop, it was clear that the proposal to include response functions and advanced analysis data products in ObsCore was at least questionable and I participated to this discussion.

A compromise such as "various discovery and access methods are possible and are in the hands of data producers" seems to be reachable. But to do this we have to review these methods somewhere in the document. The data variability and the use cases show so different requirements that we have to browse somewhere what can be done. The whole section including the inputs by Ian and the specific table subsection could become an appendix such as "data access implementation note" as suggested by Mathieu .

I still plan to make a couple of comments on Ian's PR.

Updated PR 66 based on feedback from contributors and editors.

@loumir loumir left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Here are some comments about the data product types.
I think it is important that the text reflects the semantics strategy we agreed on, which is
to propose 3 vocabularies and allow some extensions for new terms namely in advanced data products .
These data products are also considered in the radio domain, so identifying the need of a new data product vocabulary for advanced data products in this document is important .
see Advanced data product types for radio observatories

If a data collection consists primarily of individual observations, each with a small set of associated data products such as \glspl{IRF} that are required for further data analysis (this is a very common use case), then either direct access to the {\bf hea-event-bundle}s or access via DataLink is very likely preferable.

However, if the data collection includes data products such as spectra, light-curves, or advanced data products ({\em e.g.\/}, aperture photometry probability density functions) that are extracted from the primary {\bf hea-event-list} and that can be used independently of the primary dataset, then direct access to these products is likely preferable from the user's perspective. Since \gls{HEA} event lists typically simulateously encode spatial, spectral, and temporal information, a very large amount of data may be associated with a single {\bf hea-event-list}. For example, roughly 4,600 individual detectable X-ray sources are located in a single Chandra observation field of view centered on the Galactic center. Using DataLink to the primary {\bf hea-event-list} to access these spectra, light-curves, and advanced data products would be unwieldy, especially when only a small subset of the products are required.
However, if the data collection includes data products such as spectra, light-curves, or advanced data products ({\em e.g.\/}, aperture photometry probability density functions) that are extracted from the primary {\bf hea-event-list} and that can be used independently of the primary dataset, then direct access to these products is likely preferable from the user's perspective. Since \gls{HEA} event lists typically simultaneously encode spatial, spectral, and temporal information, a very large amount of data may be associated with a single {\bf hea-event-list}. For example, roughly 4,600 individual detectable X-ray sources are located in a single deep Chandra observation field of view centered on the Galactic center. Using DataLink to the primary {\bf hea-event-list} to access these spectra, light-curves, and advanced data products would be unwieldy, especially when only a small subset of the products are required.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The last sentence is strange for me . Spectra and light-curves are available directly by Obscore table already as data products from the existing vocabulary. There might be some misunderstanding about the necessity to discover products from the eventlist.
Any column of the Obscore table can be used to discover a spectrum : obs_id , obs_publisher_id, position, energy band, time interval, etc..
we have use-cases with light-curve discovery in appendix A.

\subsubsection{Datalink Access Using a Service Descriptor}
\label{sec:datalinksd}
To avoid preventing a direct access to the {\bf hea-event-list} while keeping the explicit access to the various response functions, it is possible to add a ``DataLink'' service descriptor in the ObsTAP result as can be seen in the example below.
To avoid preventing a direct access to the {\bf hea-event-list} while keeping the explicit access to the various response functions, one can add a ``DataLink'' service descriptor in the ObsTAP result as shown in the example below.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 , thanks

\noindent Note that the referenced ID field is not required to be {\em obs\_publisher\_did\/}. It can be {\em obs\_id\/}, or any free identifier attribute, allowing grouping of several rows together. For example, referencing {\em obs\_id\/} will allow all the datasets belonging to the same observation to be bound to the same \blinks\ endpoint response.

This method is implemented in many archive services such as ESO, CADC, GAVO, HEASARC and even VizieR.
This method is implemented in many archive services such as CADC, ESO, GAVO, HEASARC, and VizieR.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

<GROUP name="inputParams">
<PARAM name="ID" datatype="char" arraysize="*" value=""
ref="referred_ID_FIELD"/></GROUP>
ref="referenced_ID_FIELD"/></GROUP>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

}

In this case the service descriptor provides the root URL to the DataLink \blinks\ endpoint. The full \blinks\ endpoint URL for a specific row (or group of rows) is built by adding the content of the referred ID FIELD this way:
In this case the service descriptor provides the root URL to the DataLink \blinks\ endpoint. The full \blinks\ endpoint URL for a specific row (or group of rows) is built by adding the content of the referenced ID field this way:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1

WHERE (INTERSECTS(s_region, CIRCLE(312.775, 30.683, 1.5)) = 6)
AND (dataproduct_type = `hea-event-list')
AND semantics = `#this')
AND semantics = `#this')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks, +1


Using this approach, {\bf response-function} data products could be described by one or more alternate {\tt ivoa.response$\{$\_xxx$\}$} tables, where {\tt $\{$\_xxx$\}$} is optional and would depend on the type of {\bf response-function} in the case that different sets of queryable attributes are required for different types of {\bf response-function}s. One possible example of such a table is presented in Table~\ref{tab:response_table}.

In order to handle the various possible cardinality relationships between {\bf response-function} and {\bf hea-event-list} datasets, foreign keys must be defined in the {\tt ivoa.responsee$\{$\_xxx$\}$} tables that will allow {\tt JOIN} operations between those tables and the {\tt ivoa.obscore} table.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

a typo : should be "ivoa.response" instead of "ivoa.responsee"


An extensive set of parameter values is often required to construct many types of {\bf response-function} data products, for example depending on the assumed spatial, spectral, and temporal models of the observed astrophysical source as well as multiple instrumental properties. Given the wide range of possibilities, most facilities deliberately choose to limit the set of {\bf response-function} datasets that they make available to the end user, for example by choosing a single spectral model, such as a power-law spectrum, with predefined parameters. Often these {\bf response-function}s will have been used when calibrating the associated {\bf hea-event-list} or will have been created based on some of the assumptions used during the data processing steps ({\em e.g.\/}, based on {\em analysis\_mode\/}, {\em event\_type\/}, and so on). This is evident in the use cases recorded in Appendix~\ref{sec:uc} where typically a single {\bf response-function} of a given type is associated with a single {\bf hea-event-list} and so the selection criteria are simply based on the type of {\bf response-function}, the observation dataset properties, and on the observation properties such as the observing position and field-of-view, the observation date and time, the energy, and time bounds.

More generally, the cardinality of the relationship between a {\bf response-function} dataset and an {\bf hea-event-list} dataset will vary according to the facility and may be one-to-one, one-to-many, many-to-one, or even many-to-many. Compared to the set of queryable attributes included in {\tt ivoa.obscore} table, typically more but occasionally fewer attributes will be required to select an appropriate {\bf response-function} dataset of a given type. The set of attributes required to identify an appropriate {\bf response-function} dataset will typically depend on the type of {\bf response-function}.

@loumir loumir Jul 21, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't follow you on the idea of specializing each table to each kind of response product. My proposal was to allow all responses to be discovered in the table JOIN using either a general term like "response-function" or a specialized one like "psf" or "edisp".
The idea is to define a table template for all responses, and see in practice which columns are suitable to refine the search criteria and add them in the response table. We did not explore yet use cases for fine search in response products.

Comment thread HighEnergyObsCoreExt.tex
The difference between {\em messenger\/} and {\em messenger\_pdgid} is that ({\em e.g.\/}) a data discovery query that includes {\tt messenger = `neutrino'} would enable the user to identify datasets with a neutrino (any kind) messenger particle, whereas a data discovery query that includes {\tt messenger\_pdgid = `+16' OR messenger\_pdgid = `-16'} would restrict the search to include only $\tau$ neutrinos ($\nu_\tau$) or antineutrinos ($\nu_\tau^{-}$).

The downside of {\em messenger\_pdgid} is that PDG ID very unlikely to be recognized outside of the particle astrophysics and high energy particle physics communities. We choose to include both {\em messenger\/} and {\em messenger\_pdgid} because the majority of astrophysicists --- even experienced high-energy astrophysicists --- are unlikely to recognize PDG ID\null. While any astrophysicist is likely to be able to write a query such as {\tt messenger = `photon'}, very few would be able to write that query as {\tt messenger\_pdgid = `+22'} from memory.
The downside of {\em messenger\_pdgid} is that PDG ID very unlikely to be recognized outside of the particle astrophysics and high energy particle physics communities. We choose to include both {\em messenger\/} and {\em messenger\_pdgid} because the majority of astrophysicists --- even experienced high-energy astrophysicists --- are unlikely to recognize PDG ID\null. While any astrophysicist is likely to be able to write a query such as {\tt messenger = `photon'}, very few would be able to cast that query as {\tt messenger\_pdgid = `+22'} from memory.

@loumir loumir Jul 21, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The messenger ( astrophysical thing ) can be described by 'messenger_name' and its type can be encoded as 'messenger_pdgid' when relevant.
I would prefer these column names for more symmetry. But I can live with what is written.

Comment thread HighEnergyObsCoreExt.tex

The optional attribute {\em dataproduct\_subtype} may be used by the data provider to specify more precisely the scientific nature of a data product. Although no vocabulary is defined for {\em dataproduct\_subtype\/}, we recommend that data providers formulate and use a standardized vocabulary for this attribute for data products that are commonly used in \gls{HEA}\null. We have proposed several terms in \S~5 for commonly used \gls{HEA} {\bf response-function} types ({\em e.g.\/}, {\bf aeff}, {\bf edisp}, {\bf psf}), but additional terms could be standardized for other common data products. For example, standardizing using {\bf exposure-map} for an exposure map would enable queries such as ({\em dataproduct\_type\/} = {\bf image}) AND ({\em dataproduct\_subtype\/} = {\bf exposure-map}) to work across multiple facilities. Other possible terms could include (but are not limited to) {\bf significance-map} for a significance map, {\bf probability-map} for a probability map, and {\bf exclusion-map} for an exclusion map ({\em e.g.\/}, as used to adjust TeV background models).

We note in passing that {\bf response-function}s for some facilities and instruments may not be separable into distinct components such as {\bf aeff}, {\bf edisp}, {\bf psf}.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good to mention it . in that case a new term in the response-type vocabulary could be introduced.

@mservillat

mservillat commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

Thanks a lot Ian for this hard work. The text reads well and opens the discussion on important topics. I also noticed that the Radio Interest Group recently wrote its roadmap with this item "Shift focus to advanced data products". It will be interesting to discuss possible commonalities with the HEIG.

My main comment is that most of what is Appendix B in this PR should really be a section 6. As it is now, this appendix arrives after use cases (Appendix A), but presents access concepts that are also used in Appendix A (i.e. direct DataLink). It would thus naturally fit as section 6, except for the proposed table and associated queries, that would follow naturally Appendix A.

The following minor changes in the text would link Appendix B to section 6.3 :

"One possible example of such a table is presented in Table 8."
--> "One possible example of such a table is presented for discussion in Appendix B."

Start of Appendix B:
"We further illustrate in this appendix the possible use of an alternate table as presented in section 6.3. More discussion and use cases will be needed to further explore and generalize this solution."

I then have a comment on the wording in 6.3/B.3 :

"is very likely preferable."
--> "is generally preferable."
It is said just before that "The choice of access method is entirely up to the data provider", it is thus a bit strong and in contradiction to then show that some access method are "very likely" preferable.
--> maybe remove also "entirely" (there is no strict difference between "up to" and "entirely up to"), so the message would be clearer I think.

Typos :

  • ivoa.responsee{_xxx} --> ivoa.response{_xxx}

@mcdittmar

Copy link
Copy Markdown
Collaborator

Mathieu,

My main comment is that most of what is Appendix B in this PR should really be a section 6. As it is now, this appendix arrives after use cases (Appendix A), but presents access concepts that are also used in Appendix A (i.e. direct DataLink). It would thus naturally fit as section 6, except for the proposed table and associated queries, that would follow naturally Appendix A.

I'm liking this content in the Appendix. These seem like options available to the users for implementing a response, and they have varying degrees of exercise in practice. I agree this is good content for the migration process, presumably more on the DAL side (being more response oriented?). Perhaps after consideration and practice, this can change to recommending one approach or another.

@loumir

loumir commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

My main comment is that most of what is Appendix B in this PR should really be a section 6. As it is now, this appendix arrives after use cases (Appendix A), but presents access concepts that are also used in Appendix A (i.e. direct DataLink). It would thus naturally fit as section 6, except for the proposed table and associated queries, that would follow naturally Appendix A.

I also think that the direct access to event-bundle and the datalink strategy should be presented before the Uses cases appendix, in order to understand better use case A1.6 .
The response table can stay as appendix B.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants