Since the apps team will be the main consumer of DyK snippets, it makes sense for us to be the owners of the logic that produces the snippets in the format that is easiest for us to consume.
Therefore we will create a python library that will be a dependency of the DyK Archive service (published as a wheel), and will do the job of taking a full HTML version of a did-you-know archive page, and parsing it into a list of DyK snippet objects, ready for ingestion into the Archive service.
The library will most likely reside on GitLab, and could follow one of these repos as a blueprint:
https://gitlab.wikimedia.org/repos/research/edit-types
https://gitlab.wikimedia.org/repos/data-engineering/workflow_utils/-/blob/main/README.md?ref_type=heads#python-lib-publish-pipeline