Different experiments provide different information about physical constants. Making it easy to submit experimental data so that the uncertainty around fundamental constant’s estimations would allow us to aggregate data of multiple studies, and with it validate different experiments against each other, and better keep track of different constants.

Goal

Build accurate fundamental constant estimates through the aggregation of different experiment, using an MLE methodology

Modelling the likelihood

The main problem when trying to get the likelihood of an “implicit” variable is the unusual transformations through which it might go. For instance, we might try to understand the distribution of a random variable by an experiment which is only affected by . While informative, this measurement will not allow us to learn anything new about that is not periodic. One way of modelling the likelihood function for a given experiment is to apply equation learning where the input is a value of , and the output is the likelihood of observing this value. The exact methodology needs more work, but it will probably apply the results of Discovering Equations using Machine Learning.

Rough idea

  • A dataset exists of collected data for an experiment
  • Models can be added to the dataset. These consist of equations where we map variables to
    • Unobserved variables specific to the experiment
    • Fundamental constants of the universe
    • Measurements based on unobserved variables + noise/errors
  • Based on the models we can use Monte Carlo simulations to get a joint density distribution for the various variables
  • We can compound the distributions of the various experiments in order to get a stronger/more finely tuned measurement of the underlying variables
  • It is unclear what is the best way to describe the probability distributions. A key resource here could be SciKit’s models when displaying continuous marginal distribution histograms
  • The key value add is to see
    • Evolution of the uncertainty around a physical constant over time:
      • WITH references to the key experiments/papers that reduced the uncertainty.
      • We can Score papers based on the information increase ( entropy ratio of the continuous variable before/after the experiment )
    • Joint distributions of the physical constant value. I need to better understand how to use marginal distributions to feed into the joint distribution.

User experience

  • Users can
    • check different constants
    • see the estimated values over time
    • see which experiments provide the most information on each constant.
  • Users can add a constant to the system ( name + definition )
  • Users can submit experiments, containing
    • The formula tested in the experiment
    • The relevant constants used
    • The value of any fixed parameters, including measurement noise terms for any measured variables. The noise can be parameterized through bias, etc, whose variables will not be estimated.

Marketing

Domain: Https://manyfolds.org

ChatGPT guidelines

This platform is a Bayesian aggregation engine designed to estimate uncertain physical constants by combining user-uploaded experimental datasets, user-defined generative models, and a lightweight reputation system. Its objective is to estimate the true value of physical constants. The social layer must not alter the asymptotic posterior relative to a purely statistical system; it is intended only to accelerate convergence through credibility weighting.
 
The system revolves around three core entities: Constants, Datasets, and Models.
 
A Constant is any user-defined bounded real-valued parameter representing a quantity of interest whose value is uncertain. When creating a constant, users must define finite bounds to ensure normalization and numerical stability. Priors can be defined as uniform over the linear scale within bounds or uniform over the log scale within bounds. Although priors may conceptually resemble improper distributions, bounded domains ensure proper normalization in practice. Each constant maintains exactly one global posterior distribution at any given timestamp. Posteriors evolve irreversibly when updated, unless a full recomputation is explicitly triggered. Constants are never jointly modeled; each constant maintains its own marginal posterior. If two credible models use the same constant but imply conflicting likelihood structures, model averaging causes posterior uncertainty to widen. Distinct constants may coexist even if conceptually related. A model may assert equality or functional linkage between constants, and such links are treated like any other model in Bayesian updating. If credible linking models are accepted, posterior collapse occurs irreversibly.
 
Datasets are uploaded by users and contain raw experimental observations in tabular form. All datasets are assumed independent. There is no automatic detection of shared systematics, derived relationships, or lineage. Hierarchical Bayesian uploads are permitted to reduce double-counting effects. The system structurally prefers many separate datasets over a single large dataset. Systematic bias is not automatically detectable; one hundred biased experiments are treated as potentially correct if no contradiction emerges. Retractions are handled via full recomputation centered at a constant or a model.
 
Models are immutable generative functions defining P(data | constants, noise parameters). They must explicitly define likelihood functions and may include arbitrary noise structures, including non-Gaussian noise and nuisance parameters. Constant posteriors are obtained by marginalizing over nuisance parameters. Models are compared strictly via marginal likelihood (Bayes factors). KL divergence is used to penalize asymmetric dataset-model incompatibility. Model averaging is used to compute the final posterior for a constant. When two highly credible models conflict, uncertainty in the constant increases. If credibility differs substantially, the higher-credibility model dominates inference.
 
Bayesian updating produces one global posterior per constant, evolving over real-world timestamps. Posterior mean and variance are stored over time. Contradictory credible data can widen the posterior. Full recomputation is supported and may be performed centered on a constant or a model to correct historical errors. Updates assume independence of datasets. Improper priors are conceptually allowed, but bounded domains ensure practical normalization.
 
The reputation system models credibility as a prior on honesty. Formally, the system models P(data | constants, credibility). Credibility applies to dataset uploaders and model creators, decays over time, and influences marginal likelihood via honesty priors. It is designed not to alter the asymptotic posterior under infinite honest data, but to affect convergence speed. Endorsements increase credibility of users or models, affect prior credibility of future uploads, and modify marginal likelihood weighting in a Bayesian-consistent manner. Penalties are applied when datasets or models are contradicted by later evidence, reducing credibility scores. There is no external fraud arbitration mechanism.
 
Double counting is not explicitly detected. Independence across datasets is assumed. Hierarchical modeling is the only mitigation strategy. Many independent datasets are considered stronger evidence than a single precise dataset.
 
Temporal behavior is based on real-world timestamps. Users can inspect posterior belief states at any past timestamp. Users cannot hypothetically remove a dataset without triggering recomputation. Credibility scores decay over time to reduce overconfidence. Posterior variance may increase when credible contradictory models emerge.
 
For each constant, the platform outputs a single posterior distribution over its bounded domain. This posterior is computed via Bayesian model averaging across credible models. Each model contributes through its marginal likelihood. Dataset likelihoods are credibility-weighted. KL divergence penalizes asymmetric incompatibility. Posterior variance may increase under model conflict.
 
The system assumes constants are never jointly modeled, datasets are independent, models are compared strictly via marginal likelihood, and the social layer must not change the asymptotic posterior given infinite honest data. The system tolerates systematic bias unless explicitly contradicted and does not simulate adversarial attacks prior to launch. The overarching goal is to continuously aggregate experimental evidence to estimate true physical constants using Bayesian updating with reputation-weighted likelihoods and model competition, producing a single evolving posterior per constant.