Jump to content

Main menu Navigation ●Main page ●Contents ●Current events ●Random article ●About Wikipedia ●Contact us ●Donate Contribute ●Help ●Learn to edit ●Community portal ●Recent changes ●Upload file

●Create account ●Log in ●Create account ● Log in Pages for logged out editors learn more ●Contributions ●Talk

(Top) 1 Definition 1.1 Proof 2 Properties 2.1 Expected values 2.2 Transformation 3 Example 4 Maximum likelihood parameter estimation 5 Drawing values from the distribution 6 Relation to other distributions 7 See also 8 References

Matrix normal distribution

●Català ●Türkçe ●中文 Edit links ●Article ●Talk ●Read ●Edit ●View history Tools Actions ●Read ●Edit ●View history General ●What links here ●Related changes ●Upload file ●Special pages ●Permanent link ●Page information ●Cite this page ●Get shortened URL ●Download QR code ●Wikidata item Print/export ●Download as PDF ●Printable version Appearance From Wikipedia, the free encyclopedia

Matrix normal
Notation	${\mathcal {MN}}_{n,p}(\mathbf {M} ,\mathbf {U} ,\mathbf {V} )$
Parameters	$\mathbf {M}$ location (real $n\times p$ matrix) $\mathbf {U}$ scale (positive-definite real $n\times n$ matrix) $\mathbf {V}$ scale (positive-definite real $p\times p$ matrix)
Support	$\mathbf {X} \in \mathbb {R} ^{n\times p}$
PDF	${\frac {\exp \left(-{\frac {1}{2}}\,\mathrm {tr} \left[\mathbf {V} ^{-1}(\mathbf {X} -\mathbf {M} )^{T}\mathbf {U} ^{-1}(\mathbf {X} -\mathbf {M} )\right]\right)}{(2\pi )^{np/2}\|\mathbf {V} \|^{n/2}\|\mathbf {U} \|^{p/2}}}$
Mean	$\mathbf {M}$
Variance	$\mathbf {U}$ (among-row) and $\mathbf {V}$ (among-column)

Instatistics, the matrix normal distributionormatrix Gaussian distribution is a probability distribution that is a generalization of the multivariate normal distribution to matrix-valued random variables.

Definition[edit]

The probability density function for the random matrix X (n × p) that follows the matrix normal distribution ${\mathcal {MN}}_{n,p}(\mathbf {M} ,\mathbf {U} ,\mathbf {V} )$ has the form:

p(\mathbf {X} \mid \mathbf {M} ,\mathbf {U} ,\mathbf {V} )={\frac {\exp \left(-{\frac {1}{2}}\,\mathrm {tr} \left[\mathbf {V} ^{-1}(\mathbf {X} -\mathbf {M} )^{T}\mathbf {U} ^{-1}(\mathbf {X} -\mathbf {M} )\right]\right)}{(2\pi )^{np/2}|\mathbf {V} |^{n/2}|\mathbf {U} |^{p/2}}}

where $\mathrm {tr}$ denotes trace and Misn × p, Uisn × n and Visp × p, and the density is understood as the probability density function with respect to the standard Lebesgue measure in $\mathbb {R} ^{n\times p}$ , i.e.: the measure corresponding to integration with respect to $dx_{11}dx_{21}\dots dx_{n1}dx_{12}\dots dx_{n2}\dots dx_{np}$ .

The matrix normal is related to the multivariate normal distribution in the following way:

\mathbf {X} \sim {\mathcal {MN}}_{n\times p}(\mathbf {M} ,\mathbf {U} ,\mathbf {V} ),

if and only if

\mathrm {vec} (\mathbf {X} )\sim {\mathcal {N}}_{np}(\mathrm {vec} (\mathbf {M} ),\mathbf {V} \otimes \mathbf {U} )

where $\otimes$ denotes the Kronecker product and $\mathrm {vec} (\mathbf {M} )$ denotes the vectorizationof $\mathbf {M}$ .

Proof[edit]

The equivalence between the above matrix normal and multivariate normal density functions can be shown using several properties of the trace and Kronecker product, as follows. We start with the argument of the exponent of the matrix normal PDF:

{\begin{aligned}&\;\;\;\;-{\frac {1}{2}}{\text{tr}}\left[\mathbf {V} ^{-1}(\mathbf {X} -\mathbf {M} )^{T}\mathbf {U} ^{-1}(\mathbf {X} -\mathbf {M} )\right]\\&=-{\frac {1}{2}}{\text{vec}}\left(\mathbf {X} -\mathbf {M} \right)^{T}{\text{vec}}\left(\mathbf {U} ^{-1}(\mathbf {X} -\mathbf {M} )\mathbf {V} ^{-1}\right)\\&=-{\frac {1}{2}}{\text{vec}}\left(\mathbf {X} -\mathbf {M} \right)^{T}\left(\mathbf {V} ^{-1}\otimes \mathbf {U} ^{-1}\right){\text{vec}}\left(\mathbf {X} -\mathbf {M} \right)\\&=-{\frac {1}{2}}\left[{\text{vec}}(\mathbf {X} )-{\text{vec}}(\mathbf {M} )\right]^{T}\left(\mathbf {V} \otimes \mathbf {U} \right)^{-1}\left[{\text{vec}}(\mathbf {X} )-{\text{vec}}(\mathbf {M} )\right]\end{aligned}}

which is the argument of the exponent of the multivariate normal PDF with respect to Lebesgue measure in $\mathbb {R} ^{np}$ . The proof is completed by using the determinant property: $|\mathbf {V} \otimes \mathbf {U} |=|\mathbf {V} |^{n}|\mathbf {U} |^{p}.$

Properties[edit]

If $\mathbf {X} \sim {\mathcal {MN}}_{n\times p}(\mathbf {M} ,\mathbf {U} ,\mathbf {V} )$ , then we have the following properties:^[1]^[2]

Expected values[edit]

The mean, or expected value is:

E[\mathbf {X} ]=\mathbf {M}

and we have the following second-order expectations:

E[(\mathbf {X} -\mathbf {M} )(\mathbf {X} -\mathbf {M} )^{T}]=\mathbf {U} \operatorname {tr} (\mathbf {V} )

E[(\mathbf {X} -\mathbf {M} )^{T}(\mathbf {X} -\mathbf {M} )]=\mathbf {V} \operatorname {tr} (\mathbf {U} )

where $\operatorname {tr}$ denotes trace.

More generally, for appropriately dimensioned matrices A,B,C:

{\begin{aligned}E[\mathbf {X} \mathbf {A} \mathbf {X} ^{T}]&=\mathbf {U} \operatorname {tr} (\mathbf {A} ^{T}\mathbf {V} )+\mathbf {MAM} ^{T}\\E[\mathbf {X} ^{T}\mathbf {B} \mathbf {X} ]&=\mathbf {V} \operatorname {tr} (\mathbf {U} \mathbf {B} ^{T})+\mathbf {M} ^{T}\mathbf {BM} \\E[\mathbf {X} \mathbf {C} \mathbf {X} ]&=\mathbf {V} \mathbf {C} ^{T}\mathbf {U} +\mathbf {MCM} \end{aligned}}

Transformation[edit]

Transpose transform:

\mathbf {X} ^{T}\sim {\mathcal {MN}}_{p\times n}(\mathbf {M} ^{T},\mathbf {V} ,\mathbf {U} )

Linear transform: let D (r-by-n), be of full rank r ≤ n and C (p-by-s), be of full rank s ≤ p, then:

\mathbf {DXC} \sim {\mathcal {MN}}_{r\times s}(\mathbf {DMC} ,\mathbf {DUD} ^{T},\mathbf {C} ^{T}\mathbf {VC} )

Example[edit]

Let's imagine a sample of n independent p-dimensional random variables identically distributed according to a multivariate normal distribution:

\mathbf {Y} _{i}\sim {\mathcal {N}}_{p}({\boldsymbol {\mu }},{\boldsymbol {\Sigma }}){\text{ with }}i\in \{1,\ldots ,n\}

When defining the n × p matrix $\mathbf {X}$ for which the ith row is $\mathbf {Y} _{i}$ , we obtain:

\mathbf {X} \sim {\mathcal {MN}}_{n\times p}(\mathbf {M} ,\mathbf {U} ,\mathbf {V} )

where each row of $\mathbf {M}$ is equal to ${\boldsymbol {\mu }}$ , that is $\mathbf {M} =\mathbf {1} _{n}\times {\boldsymbol {\mu }}^{T}$ , $\mathbf {U}$ is the n × n identity matrix, that is the rows are independent, and $\mathbf {V} ={\boldsymbol {\Sigma }}$ .

Maximum likelihood parameter estimation[edit]

Given k matrices, each of size n × p, denoted $\mathbf {X} _{1},\mathbf {X} _{2},\ldots ,\mathbf {X} _{k}$ , which we assume have been sampled i.i.d. from a matrix normal distribution, the maximum likelihood estimate of the parameters can be obtained by maximizing:

\prod _{i=1}^{k}{\mathcal {MN}}_{n\times p}(\mathbf {X} _{i}\mid \mathbf {M} ,\mathbf {U} ,\mathbf {V} ).

The solution for the mean has a closed form, namely

\mathbf {M} ={\frac {1}{k}}\sum _{i=1}^{k}\mathbf {X} _{i}

but the covariance parameters do not. However, these parameters can be iteratively maximized by zero-ing their gradients at:

\mathbf {U} ={\frac {1}{kp}}\sum _{i=1}^{k}(\mathbf {X} _{i}-\mathbf {M} )\mathbf {V} ^{-1}(\mathbf {X} _{i}-\mathbf {M} )^{T}

and

\mathbf {V} ={\frac {1}{kn}}\sum _{i=1}^{k}(\mathbf {X} _{i}-\mathbf {M} )^{T}\mathbf {U} ^{-1}(\mathbf {X} _{i}-\mathbf {M} ),

See for example ^[3] and references therein. The covariance parameters are non-identifiable in the sense that for any scale factor, s>0, we have:

{\mathcal {MN}}_{n\times p}(\mathbf {X} \mid \mathbf {M} ,\mathbf {U} ,\mathbf {V} )={\mathcal {MN}}_{n\times p}(\mathbf {X} \mid \mathbf {M} ,s\mathbf {U} ,{\tfrac {1}{s}}\mathbf {V} ).

Drawing values from the distribution[edit]

Sampling from the matrix normal distribution is a special case of the sampling procedure for the multivariate normal distribution. Let $\mathbf {X}$ be an nbyp matrix of np independent samples from the standard normal distribution, so that

\mathbf {X} \sim {\mathcal {MN}}_{n\times p}(\mathbf {0} ,\mathbf {I} ,\mathbf {I} ).

Then let

\mathbf {Y} =\mathbf {M} +\mathbf {A} \mathbf {X} \mathbf {B} ,

so that

\mathbf {Y} \sim {\mathcal {MN}}_{n\times p}(\mathbf {M} ,\mathbf {AA} ^{T},\mathbf {B} ^{T}\mathbf {B} ),

where A and B can be chosen by Cholesky decomposition or a similar matrix square root operation.

Relation to other distributions[edit]

Dawid (1981) provides a discussion of the relation of the matrix-valued normal distribution to other distributions, including the Wishart distribution, inverse-Wishart distribution and matrix t-distribution, but uses different notation from that employed here.

References[edit]

^ A K Gupta; D K Nagar (22 October 1999). "Chapter 2: MATRIX VARIATE NORMAL DISTRIBUTION". Matrix Variate Distributions. CRC Press. ISBN 978-1-58488-046-2. Retrieved 23 May 2014.

^ Ding, Shanshan; R. Dennis Cook (2014). "DIMENSION FOLDING PCA AND PFC FOR MATRIX- VALUED PREDICTORS". Statistica Sinica. 24 (1): 463–492.

^ Glanz, Hunter; Carvalho, Luis (2013). "An Expectation-Maximization Algorithm for the Matrix Normal Distribution". arXiv:1309.6609 [stat.ME].

Dawid, A.P. (1981). "Some matrix-variate distribution theory: Notational considerations and a Bayesian application". Biometrika. 68 (1): 265–274. doi:10.1093/biomet/68.1.265. JSTOR 2335827. MR 0614963.
Dutilleul, P (1999). "The MLE algorithm for the matrix normal distribution". Journal of Statistical Computation and Simulation. 64 (2): 105–123. doi:10.1080/00949659908811970.
Arnold, S.F. (1981), The theory of linear models and multivariate analysis, New York: John Wiley & Sons, ISBN 0471050652

Probability distributions (list)

Discrete
univariate

with finite
support

with infinite
support

Continuous
univariate

supported on a bounded interval	arcsine ARGUS Balding–Nichols Bates beta beta rectangular continuous Bernoulli Irwin–Hall Kumaraswamy logit-normal noncentral beta PERT raised cosine reciprocal triangular U-quadratic uniform Wigner semicircle
supported on a semi-infinite interval	Benini Benktander 1st kind Benktander 2nd kind beta prime Burr chi chi-squared noncentral inverse scaled Dagum Davis Erlang hyper exponential hyperexponential hypoexponential logarithmic F noncentral folded normal Fréchet gamma generalized inverse gamma/Gompertz Gompertz shifted half-logistic half-normal Hotelling's T-squared inverse Gaussian generalized Kolmogorov Lévy log-Cauchy log-Laplace log-logistic log-normal log-t Lomax matrix-exponential Maxwell–Boltzmann Maxwell–Jüttner Mittag-Leffler Nakagami Pareto phase-type Poly-Weibull Rayleigh relativistic Breit–Wigner Rice truncated normal type-2 Gumbel Weibull discrete Wilks's lambda
supported on the whole real line	Cauchy exponential power Fisher's z Kaniadakis κ-Gaussian Gaussian q generalized normal generalized hyperbolic geometric stable Gumbel Holtsmark hyperbolic secant Johnson's S_U Landau Laplace asymmetric logistic noncentral t normal (Gaussian) normal-inverse Gaussian skew normal slash stable Student's t Tracy–Widom variance-gamma Voigt
with support whose type varies	generalized chi-squared generalized extreme value generalized Pareto Marchenko–Pastur Kaniadakis κ-exponential Kaniadakis κ-Gamma Kaniadakis κ-Weibull Kaniadakis κ-Logistic Kaniadakis κ-Erlang q-exponential q-Gaussian q-Weibull shifted log-logistic Tukey lambda

Mixed
univariate

continuous-
discrete

Rectified Gaussian

Multivariate
(joint)

Discrete:
Ewens
multinomial
- Dirichlet
- negative
Continuous:
Dirichlet
- generalized
multivariate Laplace
multivariate normal
multivariate stable
multivariate t
normal-gamma
- inverse
Matrix-valued:
LKJ
matrix normal
matrix t
matrix gamma
- inverse
Wishart
- normal
- inverse
- normal-inverse
- complex

Directional

Univariate (circular) directional: Circular uniform; univariate von Mises; wrapped normal; wrapped Cauchy; wrapped exponential; wrapped asymmetric Laplace; wrapped Lévy
Bivariate (spherical): Kent
Bivariate (toroidal): bivariate von Mises
Multivariate: von Mises–Fisher; Bingham

Degenerate
and singular

Degenerate: Dirac delta function
Singular: Cantor

Families

Retrieved from "https://en.wikipedia.org/w/index.php?title=Matrix_normal_distribution&oldid=1119144125" Categories: ●Random matrices ●Continuous distributions ●Multivariate continuous distributions Hidden categories: ●Articles with short description ●Short description matches Wikidata ●This page was last edited on 30 October 2022, at 23:34 (UTC). ●Text is available under the Creative Commons Attribution-ShareAlike License 4.0; additional terms may apply. By using this site, you agree to the Terms of Use and Privacy Policy. Wikipedia® is a registered trademark of the Wikimedia Foundation, Inc., a non-profit organization. ●Privacy policy ●About Wikipedia ●Disclaimers ●Contact Wikipedia ●Code of Conduct ●Developers ●Statistics ●Cookie statement ●Mobile view