Search Blogs

Thursday, March 16, 2023

Bayes Rule: Visual Referesher

Bayes rule is a familiar or natural outcome for most familiar with probability theory. In words, it tells us how to update the probability of a random variable(s) given some event(s) has occurred and that we have some prior knowledge or belief about the probability of the random variable(s) from earlier events. The algebra to get to Bayes rule is simple but I found it always best to have a more spatial perspective on what Bayes rule is really stating.

I'll first begin with ta 2D square sample space, $\it{S}$. This space is discrete and we can represent each outcome as a tiny square, $\it{s}$. In this case, we will have a total of 12 tiny squares in $\it{S}$. This means there is a 1/16 chance that any square is randomly selected, hence, $p(s)$.

$$\begin{array}{|c|c|c|c|}\hline  \it{s_1} & \it{s_2}  & \it{s_3}  & \it{s_4}   \\ \hline \it{s_5}  & \it{s_6}  & \it{s_7} & \it{s_8}  \\ \hline \it{s_9}  & \it{s_{10}}  & \it{s_{11}}  &  \it{s_{11}} \\ \hline \it{s_{13}}  & \it{s_{14}} & \it{s_{15}}  & \it{s_{16}} \\ \hline \end{array}$$

$$\mathrm{P}(\it{s}_i)_{\it{S}} = \mathrm{1/16}$$

Now we say have the scenario where we are only interested in two subspaces of $\it{S}$, $\it{S}_A$ and $\it{S}_B$. More specifically we want to know the probabilities of a square randomly occurring in each of these subspaces given they occur in $\it{S}$ and what the probability is of a square occurring in the intersection, or state differently, the probability of a square occurring in both $\it{S}_A$ and $\it{S}_B$.

With this we have the following: $\mathrm{P}(\it{s}_A)$, $\mathrm{P}(\it{s}_B)$, and $\mathrm{P}(\it{s}_A \cap \it{s}_B) = \mathrm{P}(\it{s}_A)$. The updated image of this would look like:

The probability $\mathrm{P}(\it{s}_A)$ in red, $\mathrm{P}(\it{s}_B)$ in blue, and the overlap $P(\it{s}_A \cap \it{s}_B)$. Keep in mind that $\mathrm{P}(\it{s}_A \cap \it{s}_B) = \mathrm{P}(\it{s}_A \cap \it{s}_B) = \mathrm{1/8}$. 

The question we usually want to ask is not what the joint probability, i.e., what's the probability of both $\it{S}_A$ and $\it{S}_B$ squares, but instead is what is the probability of a square in $\it{S}_A$  given that a square in $\it{S}_B$ has been picked/occurred or vice versa.  So what does this mean? We want to compare the relative probabilities of the joint space to that of the given space where the event has occurred:

\begin{equation} \mathrm{P}(\it{s}_A | \it{s}_B) = \frac{\mathrm{P}(\it{s}_A  \cap \it{s}_B)}{\mathrm{P}(\it{s}_B)}\label{eq:bayes1}   \end{equation} 

and

\begin{equation} \mathrm{P}(\it{s}_B| \it{s}_A) = \frac{\mathrm{P}(\it{s}_A  \cap \it{s}_B)}{\mathrm{P}(\it{s}_A)} \label{eq:bayes2}\end{equation}

Notice how these two equations are not the same but we the probability in the joint space, $\mathrm{P}(\it{S}_A \cap \it{S}_B) = \mathrm{P}(\it{S}_B \cap \it{S}_A)$. This had to be the case just by looking at the illustration with the colored cells above. 

The key is that we can now determine the conditional probabilities, that is the probability of a cell in a subspace given a cell in the other subspace has been picked or occurred, by rearrange eq. \ref{eq:bayes1} and eq. \ref{eq:bayes2} for the joint probability and then substituting terms to get:

\begin{equation*} \mathrm{P}(\it{s}_A | \it{s}_B) \mathrm{P}(\it{s}_B) = \mathrm{P}(\it{s}_B | \it{s}_A) \mathrm{P}(\it{s}_A)\end{equation*}

which is rearranged to get the typical Bayes formula:

\begin{equation}\mathrm{P}\left(\it{s}_A | \it{s}_B\right)  = \frac{\mathrm{P}\left(\it{s}_B | \it{s}_A\right) \mathrm{P}\left(\it{s}_A\right)}{\mathrm{P}\left(\it{s}_B\right)} \label{eq:bayesformula}\end{equation}.

At first eq. \ref{eq:bayesformula} might seem expected you could. I mean it is just an outcome of analyzing probabilities of subspaces, but the impact is really how one can this equation to update knowledge. Let us break down the terms in eq. \ref{eq:bayesformula}.

The first term in the numerator is called the likelihood probability. It indicates how probably an event in $\it{S}_B$ is given that an event in $\it{S}_A$ occurs. It can also represent the probability of the observed data given the model and its parameters (i.e. prior over parameters). The second term in the numerator, the prior, informs about previous knowledge of the observations or parameters. Finally, the denominator can be interpreted as the probability of observing a cell in $\it{S}_B$ or you can think about it as the data averaged over all possible values of the model parameters. 

An important aspect of eq. $\ref{eq:bayesformula}$ is that in the case of probability functions, the integration equals one. This just means that over the whole space of probabilities, something must have happened.

In the example given, the probabilities are just uniform discrete values, so we obtain a posterior probability that is just a number that represents our updated knowledge about the probability of a cell in $\it{S}_{A}$ given the cell is in $\it{S}_{B}$. This is a particularly simple and maybe intuitive outcome. What is typically more useful is that we have a probability density function that represents our prior knowledge about an event/outcome and we want to determine the posterior distribution. We then choose a likelihood probability that encodes information about what has been observed given the prior probability and make inferences by sampling the constructed posterior distribution.


Reuse and Attribution

Thursday, March 9, 2023

Book Review: Quantum Entanglement, Jed Brody

I just finished reading through this short monograph on quantum entanglement. The approach taken by the author is to provide what quantum entanglement is through conceptual examples. There are no wave functions or quantum states discussed in this book. At first, the reader is introduced to two very important concepts in the philosophy of physics; realism and locality. In realism, the assumption is that any physical objects have properties regardless of whether another object with agency (i.e., a person) is observing that object. A typical example of this concept is the following questions:

Does a falling tree in the forest make a sound when no one is listening? 

Realism says yes, it does. In the case of the tree, it has a center of mass that gives the tree some gravitational potential energy that upon falling is converted to kinetic energy and then generates sound waves in the air once it hits to ground. The tree had mass, potential energy, and kinetic energy which according to realism exist objectively. The opposing view is that it was the sound wave came into existence because an agent was listening. This seems absurd and it is in classical physics, but not necessarily in quantum physics.

Locality refers to the fact that observing, measuring, or disturbing objects in a region of a space does not affect other objects at arbitrary distances in that space. Here, I'm using space in an abstract sense not necessarily a Euclidean 3D space. I do note that in the book the discussion of locality is with regard to distances in 3D Euclidean geometry but I think I'm correct in that locality applies to non-Euclidean 3D spaces and this would be the more general statement. Locality is a pretty important concept in physics and is one of the reasons we got the famous EPR paper from Einstein. 

After the book presents these two concepts it gradually moves into the concept of hidden variables, that is properties of objects that can change aspects of the object when observed, but yet the variables themselves are never observed. Hidden variables satisfy realism. Much of the subsequent chapters present examples that lead to the famous Bell inequality which arises due to correlations in probabilities. The bell inequality needs to be satisfied for a theory to contain locality, if it is violated the theory is non-local. As it turns out, at least to our ability to experimental test the theory of quantum mechanics, it is a non-local theory without hidden variables because. All experiments that have been conducted to date violate Bell's inequality and suggest that correlations are instantaneous within the quantum mechanics framework. It should be noted that you could have a quantum hidden variable theory (i.e. Bohemian mechanics) that is non-local which would describe experimental results, but I guess the argument against this is why introduce hidden variable theory if a non-hidden variable theory doesn't provide any additional clarity other than satisfying realism.

It is pretty well documented, or at least we are made to think, that Einstein had serious issues with the non-local (dubbed "spooky action at a distance") behavior of quantum theory as well as the mainstream interpretations not satisfying a realism philosophical perspective. More specifically, the Copenhagen interpretation posits that the wavefunction/quantum state is more of a mathematical tool and is not necessarily a physical object since it only provides a way to extract probabilities of observable properties.

Going back to the book, chapters 3 and 4 provide different and simple experimental setups that look at probabilities and their correlations to arrive at Bell's inequality. The author then reminds the reader that quantum mechanics violates this inequality. Chapters 1-4 are written in a direct and comprehendible manner, but the truth is, I find it actually easier to understand the Bell inequality and violations of it by actually following the simple Linear algebra of quantum theory.  Trying to think through all the words describing the setup and outcomes can become burdensome. Given  the current focus on quantum computing, there are a lot of good books that go through the same results using simple linear algebra. I think it would have been easy to introduce most readers interested in this book to the basics of a qubit, Hilbert space, and corresponding operations, which could help readers understand these concepts more easily.

Chapter 5 goes through the potential inconsistencies of quantum mechanics with special relativity. Personally, I found this chapter was not delivered in the most impactful way, but it addresses the original concerns of physicists.  The final concluding sentence that indicates everything is okay in the end is:

"... the linkage between entangled particles conveys neither mass nor messages"

The author ends the book with a chapter regarding realism and its validity.  I think this is the best section of the book. The author gives their thinking on the topic of local realism by stating:

"The only fact that's (almost) certain is local realism cannot account for measured results"

Thus, local realism is a dead concept in the author's eyes As frustrating as it feels, I would agree with the author. The remainder of the chapter deals with interpretations of quantum theory from philosophical perspectives and you get a nice concrete quadrant table to decide what path to take, namely:

Find falsehoods in assumptions Abandon locality & realism
Abandon locality & keep realism Abandon realism & keep locality

I'm not going to go through and explain each of these because I want to leave some excitement, but I think this is the most interesting part of the book. 

I recommend reading this book if you're going to be studying quantum mechanics in any way because it will help with some of the philosophical thinking behind the theory. The reading is extremely accessible to any background and it is very short making for a good weekend read. Here's the book:

MIT Press Store

Reuse and Attribution

Thursday, March 2, 2023

Fine-tuning GPT-3 for a LAMMPS or VASP AI chatbot

The GPT API  enables fine-tuning of the GPT model for your specific application. I'm interested in utilizing this to create a new tool that would allow a user to query a software user manual to generate macros or scripts to perform operations. The idea is to put together a series of prompts and completions that are extracted from user forms like stack overflow/exchange, discourse, etc. as well as from domain users who are willing to contribute. For example, I'm curious about creating a GPT chatbot that can provide users with LAMMPS or VASP scripts based on text prompts about the problem. At the moment, ChatGPT tries to do this but fails to get enough of the specific commands and parameters correct. 

What I'm thinking is if you have a dataset with prompts and completions like:


  {"prompt": "What is the command for\n 
computing thermal conductivity in LAMMPS",\n
"completion": "In order to calculate the\n
thermal conductivty using the Green-Kubo formulas,\n
the heat flux needs to be calculated.\n
The command to do so is:\n
compute ID group-ID heat/flux ke-ID pe-ID stress-ID"}

My hope is that if you fine-tune the GPT model with these examples the user can just ask a AI chatbot more broadly something like:

Please create a LAMMPS input script to calculate the thermal conductivity of graphite at 300K.

Would this approach work for fine-tuning a GPT model? I don't really know, I'm planning on giving it a go. I need to also be cognizant that the number of tokenizations in the dataset for fine-tuning doesn't make it a costly disaster.

I'm wondering is if there is a way to grab the questions and answers in a json format from the LAMMPS discourse community and likewise sources to create the fine-tuning dataset. If not it would be very time-consuming for domain knowledge from individuals. I guess could create some kind of community input form where users provide this. Would do the same for VASP and hopefully most of the other mainstream atomistic packages.  I have a name for a LAMMPS AI chatbot but need to ask the person first if it's okay to eponymize the chatbot after them.


Reuse and Attribution