Full Transcript

·YouTLDR

Random Sampling in Statistics: Expected Value and Variance of the Sample Mean

16:08EnglishTranscribed Jul 20, 2026
0:00

welcome back okay so we're talking about

0:03

the theory of random sampling to say

0:06

something about a large but unknown

0:08

population in terms of a random smaller

0:11

sample of that population this is useful

0:14

all over in statistics and this is kind

0:16

of an entry point to more advanced uh

0:19

topics so we showed last time that you

0:22

can have this population um which is

0:24

kind of a large population its PDF may

0:27

or may not even be known but it a mean

0:30

and and a variance and then um the

0:34

sample statistics if I take a subsample

0:36

a little n subsample of that big n

0:40

population those samples become random

0:43

variables and the average of those

0:46

random variables

0:47

xbar um is hopefully an estimate of the

0:52

population mean mu so the kind of uh

0:55

sample mean should be a good estimate of

0:58

the population mean under some circum

1:00

ances and we can also compute things

1:02

like the sample variance and so on and

1:04

so forth so this opens up a ton of

1:06

questions um all kinds of questions come

1:08

up

1:09

so question one um does the expectation

1:14

of xar does

1:17

xar equal mu we we've kind of hinted

1:21

that this sample mean should converge uh

1:25

to me in the as as little n gets bigger

1:27

and bigger as my sample size gets bigger

1:29

but can I actually show that the

1:31

expectation of this random variable is

1:33

in fact the population mean mu that

1:35

would be very useful we're going to do

1:36

that today another question um what is

1:40

the variance of xbar meaning we we have

1:45

a pretty good gut feeling that this

1:48

sample mean should be gausian

1:51

distributed for reasonably large n we

1:54

know that from the central limit theorem

1:55

this should be kind of normally

1:57

distributed hopefully with a center

1:59

around the true mean mu but what's the

2:02

variance of that distribution is it a

2:04

fat distribution is it skinny we want

2:06

the variance of xar to be really really

2:08

small because that means xar is a really

2:11

really tight estimate of mu so this has

2:14

implications um about the convergence of

2:17

these values with n how fast is xar

2:20

converged to me how efficient is it

2:23

things like that there's questions you

2:24

know are is there a bias is it is the

2:27

expectation of xar mu plus some constant

2:30

Offset you know is there any bias in my

2:35

estimates those are all powerful

2:38

statistics questions we're going to ask

2:41

and answer those for this very very

2:43

simple case of simple random sampling to

2:46

iner things about a population but these

2:49

questions hold much more generally in

2:51

statistics in data analysis and even in

2:54

machine learning okay does my you know

2:56

if I sample data do I converge to a good

3:00

model of a much bigger complex process

3:03

okay good so we're going to jump in uh

3:06

and we're going to start with the

3:07

expectation value because that's always

3:08

easier to work within the variance

3:10

because the formula is simpler than the

3:12

formula for variance so we want to show

3:16

that um the expectation value of

3:19

xar equals mu meaning that xar is an

3:24

unbiased estimate of mu so we say that

3:28

xbar is an

3:34

unbiased

3:36

estimate

3:38

of mu this is statistics language for

3:41

the expectation value of xar equals mu

3:44

and there's no constant error so the x

3:47

xar is unbiased meaning it it converges

3:50

its expectation value is exactly mu okay

3:53

so we're going to prove this now this is

3:55

pretty easy to prove uh maybe I will do

3:57

this in green so the expectation of xar

4:03

um is literally I'm just going to plug

4:05

this in to the expectation this is equal

4:08

to the

4:09

expectation of 1 / n sum I = 1 to n of

4:14

each of my random variables

4:17

XI now we know that we can pop this

4:19

constant out and this the expectation of

4:21

a sum is the sum of expectations so this

4:24

equals 1 / n sum I = 1 to

4:30

n um

4:33

expectation of each of these

4:37

XIs here's a really important fact that

4:41

I need you to to believe and to know and

4:44

I'm just going to write it down here the

4:47

E for any uh for any

4:51

individual X for any

4:54

individual uh

4:56

indiv individual

5:01

sample

5:03

XI the expected value of

5:07

XI is equal to

5:09

Mu uh and the

5:13

variance of

5:15

x i is equal to Sigma squ where mu is

5:20

the true population mean and sigma squar

5:22

is the true population variance this is

5:25

super important any one of these

5:27

individual samples it's expected value

5:30

is mean mu and its expected variance is

5:33

Sigma squar where those are the

5:35

population values you can actually

5:38

convince yourself of this pretty easily

5:40

you can write down this expectation

5:42

value um this I'll just do it for for

5:44

the expected value and you can convince

5:46

yourself also for the variance um this

5:50

is the sum over every single possible n

5:53

over all of the big n uh J equal 1 of

5:57

all of the little values X J times the

6:01

probability that my random variable x i

6:03

equals little

6:05

XJ that's just the definition of

6:07

expected value of this random variable

6:08

it's the sum over all the possible

6:10

things it could be times the probability

6:11

that it is actually that

6:13

thing and there are each of these um the

6:18

chance that I drew any one of these for

6:20

x i is just 1 over n that's the

6:23

probability so this equals the sum over

6:26

uh big n of little X J * a probability

6:30

of 1 over big n this is the definition

6:34

of

6:35

my population mean it's 1 / n times the

6:40

sum of all of those little XIs so you

6:42

can convince yourself anyway that each

6:44

of these random variables each of these

6:46

XIs their expected value is Mu and their

6:49

expected variance their variance is

6:52

Sigma squar you can think about it

6:54

because each of these X's is pulled from

6:56

this population so you can kind of say

6:59

that uh x i is

7:03

distributed according to whatever the

7:06

distribution of my

7:09

population was okay whatever my

7:11

population distribution is each of these

7:13

XIs is randomly sampled from that

7:15

population distribution so anyway this

7:18

let's go back to to what we're trying to

7:20

show we're trying to show that the

7:21

expectation of xar equals mu so we take

7:24

our sample mean xar we plug it into this

7:27

expectation and it's the sum of all of

7:30

these little the the these random

7:31

variables x i * 1 over little n the

7:35

constant pops out the sum of an expect

7:38

the expectation of a sum is the sum of

7:39

the expectations and now each of these

7:42

expectation

7:43

values is Mu so I have um essentially

7:49

this

7:50

equals uh 1/ n time the sum of IAL 1 to

7:55

little n of mu each of these is equal to

7:58

Mu this is n * mu * 1/ n this whole

8:02

thing just equals mu the expected value

8:06

of xar is equal to Mu very very cool

8:10

this means that xar the sample mean is

8:13

an unbiased estimate uh of the

8:16

population mean mu and hopefully as n

8:20

gets bigger and bigger this expected

8:23

value um sorry the this distribution of

8:26

xar gets Tighter and Tighter and Tighter

8:28

around this expected value you good um

8:31

maybe I'll just draw a little picture so

8:34

um probably I have some population

8:38

distribution and I'm actually going to

8:39

draw it to be kind of gnarly um but

8:42

let's say it has some mean value some

8:46

mu the sample mean

8:50

xar by the central limit theorem we'll

8:52

prove this later but by the central

8:54

limit theorem xar is going to be a

8:57

normally distributed variable about its

8:59

expected value of

9:02

mu so

9:05

xar should be normally distributed with

9:08

its expectation value centered around mu

9:13

and we want xar to get Tighter and

9:16

Tighter and Tighter we want the spread

9:19

of possible X bars to be really really

9:21

small around this value of mu as n gets

9:24

larger that spread of course is related

9:26

to the variance of this uh of this xar

9:30

quantity so now let's talk about what's

9:33

the

9:35

variance of xar okay variance of xar

9:38

tells me how good this estimate is for

9:42

um increasing sample size little n okay

9:46

good um this result makes intuitive

9:48

sense now let's talk about the variance

9:51

uh of

9:52

xar so I'm going to actually

9:55

prove a slight approximation what I'm

9:58

going to write down is not the exact

10:01

variance of xar it's an approximation to

10:03

the variance of xar making an assumption

10:07

that each of these X's is

10:10

independent now remember we sampled

10:13

without replacement so every time I drew

10:16

a sample my population got a little

10:18

smaller that technically builds in a

10:21

small amount of dependence between these

10:22

variables but for really really big and

10:26

for really really big populations you

10:28

can kind of assume that these are are

10:29

independent and that's what I'm going to

10:30

write down here and then I'm going to

10:32

write down the correction for finite n

10:35

for finite population size so this is an

10:39

approximation uh this equals again I'm

10:42

going to plug in this expression into

10:44

xar this equals the variance of the sum

10:47

VAR of 1 / n * X1 plus dot dot dot plus

10:53

X little

10:55

n and I'm just going to again remind you

10:58

this is uh

10:59

um this is if if these are independent

11:05

samples then I can say this this is

11:10

um actually sorry if they're independent

11:13

samples then I can split these into the

11:14

sum of a bunch of variances so I'll wait

11:16

I'll I'll I'll wait to write down my

11:18

Independence assumption in a minute um

11:21

so my 1/n pops out as a 1 over n^ 2

11:25

that's how variance of a constant times

11:27

a variable you can pop that constant

11:29

squared out so this equals 1 over n^

11:34

squar times the variance of this sum and

11:38

that is the sum of the individual

11:39

variances that is VAR

11:43

X1 plus dot dot dot plus VAR

11:49

xn now I've used this assumption this is

11:53

true if my X eyes are independent

12:00

and that's true for very very large

12:04

population size and much much greater

12:06

than one like n a million or 100,000 or

12:10

10,000 this is going to be a very good

12:12

approximation technically there is joint

12:16

co-variance between these variables and

12:18

so this step is actually not exactly

12:20

true it's really kind of this is

12:21

approximately equal to this for very

12:24

large population size so be on the watch

12:27

for me making those kinds of approx IM

12:30

again we're trying to compute the

12:31

variance of our sample mean we want that

12:34

variance to be small it's equal now

12:36

approximately to 1 n^ 2times the sum of

12:39

the variances of all of those individual

12:40

elements and the sum of those variances

12:43

each of those variances are the

12:45

population variance Sigma squar so I can

12:48

write this

12:49

now as you know each of these this is

12:52

just um let's

12:55

say

12:56

uh this is n * Sigma

13:00

s and so this whole thing is

13:03

approximately equal to n / n^ 2 * Sigma

13:06

squ that's Sigma squar over n and

13:11

actually this is the result from the

13:13

central limit theorem so I want you to

13:14

go back and and check out that Central

13:17

limit theorem uh video this is the

13:21

result from the central limit

13:23

theorem um that that if you have the sum

13:27

of a bunch of independent random

13:29

variables each with their own variance

13:31

Sigma squar then the sum of those

13:33

variables would have um this uh variance

13:37

okay so this is actually all coming from

13:40

the central limit theorem this is um I

13:42

guess law of large numbers this is

13:44

Central limit theorem

13:48

good now I'll show this in the next

13:50

video I'll actually go through the Gory

13:51

details of deriving this in the next

13:53

video but remember this is only true for

13:56

very very large n very large population

14:00

so for

14:03

finite uh population

14:05

size Big N technically this VAR X bar

14:12

there is a correction and again I'm

14:13

going to derive this in the next lecture

14:15

there's a correction it's Sigma 2ar over

14:20

n * 1 - little n minus1 over big n minus

14:27

one okay and again

14:29

and this is approximately equal to Sigma

14:33

2 over little

14:34

n when uh little n is much less than big

14:39

n when I have a really big

14:41

population um and my sample size is

14:44

small compared to that really big

14:45

population then I recover this this very

14:48

very good approximation to the variance

14:50

of xar so for small populations and

14:53

small samples you need this this finite

14:55

size correction most of the time we're

14:57

going to end up using this result from

14:58

the Central limit theorem we're going to

15:00

assume that our sample mean xar is a

15:04

normally distributed random variable

15:06

with mean mu and variance Sigma squ Over

15:10

N where Sigma squar and mu are the

15:13

variance and mean of our overall

15:15

population so this xbar tells us a lot

15:18

about this unknown

15:20

population um so measuring this xar

15:23

measuring all of this sample taking this

15:25

random sample and Computing this xar the

15:27

sample mean tells me a ton about the

15:30

population and as n gets bigger and

15:32

bigger and bigger this variance gets

15:35

smaller and smaller and smaller meaning

15:36

we

15:38

converge uh to the true population mean

15:42

with a relatively small sample n okay

15:46

super cool stuff in the next lecture

15:48

this is going to be a technical lecture

15:49

I'm actually going to derive this finite

15:51

n correction uh to the variance of xar

15:54

it's pretty technical you can probably

15:55

skip it if you like but if you want to

15:57

know where it comes from um all write

15:59

this out in terms of the co-variances um

16:01

for the shrinking without replacement

16:03

population okay thank you

More transcripts

Explore other videos transcribed with YouTLDR.

Get the TLDR of any YouTube video

Transcribe, summarize, and repurpose videos in 125+ languages — free, no signup required.

Try YouTLDR Free