Related resources
Transcript
0:00
Hello everyone and welcome to today's session. What is user focused
0:05
observability? My name is Colin Contrary. Uh I'm on the marketing team at Embrace. I'd like to thank everyone
0:12
for being here today. And before we get started, uh I'd like to set the stage a bit. So now that Embrace has joined
0:19
forces with Chronosphere. We're in the unique position of combining our expertise in the observability
0:26
challenges that affect front-end teams with the powerhouse that is Chronosphere who works on back-end infrastructure
0:32
challenges that affect some of the world's biggest companies. And therefore, we thought it would be fun to cover this concept of user focused
0:39
observability by having a Q&A between experts in the front-end and backend
0:44
observability domains. So, you're going to hear many questions and answers as our two speakers have this session. Um,
0:51
obviously, if you think of questions as this session is going on, please use the Q&A functionality in Zoom and we'll all
0:58
lot time at the end to answer any questions as they come up. And with all of that said, I'd like to give a little
1:05
setup for our two speakers today. So, we'll start with Aloque. uh Aloc is a technology leader at Chronosphere where
1:12
he helps organizations navigate the complexities of modern cloudnative observability. Uh prior to joining
1:18
Chronosphere, Aloc spent several years at Splunk driving enterprise data and monitoring solutions and with deep
1:25
expertise in telemetry, system reliability and data scale. He is passionate about empowering engineering
1:32
teams to optimize performance and accelerate innovation. uh two guesses if Aloque is going to be our back-end
1:38
observability expert in this session. Uh and our second speaker is going to be Christine. Uh Christine leads solution
1:45
engineering and embrace where she turns data into delightful user experiences. So with a knack for bridging sales,
1:52
product and customer success, she's been championing front-end teams since 2017.
1:57
Her analytical mind meets creative problem solving to make technology work better for everyone. So Christine,
2:03
obviously you'll be repping the front end in this discussion. And so with those two uh intros out of the way, we
2:08
have our speakers ready to go. I'm now going to hand it over to Aloque to kick off the discussion. Colin, thanks so much uh for such a kind
2:15
and very nice introduction to myself and Christine. Um and thank you everyone for
2:20
joining today. I know you've got many other things to do during the day. So really appreciate you taking out the time uh to join us in this in this
2:27
fireside chat. Uh before I begin, I also want to welcome the entire Embrace team
2:33
along with Christine, Colin and Eduardo on the call today. Uh they are now part of the Chronosphere family which only
2:39
recently became part of the PaloAlto Networks family. So a couple of acquisitions here, but welcome Embrace
2:44
to the team. Uh as you all know, Embrace is a leader in the in the front end observability space, front end and
2:50
mobile observability space. And I myself I'm from the Chronosphere team. Uh we
2:56
are a leader in the observability space as it is defined or backend observability as as we can call it
3:02
today. Um our claim to fame Chronosphere's claim to fame is our
3:08
ability to uh work with large cloudnative workloads and enable customers to manage their costs very
3:14
effectively. So combined, Chronosphere and Embrace are really a complimentary uh capability allows you to go all the
3:21
way from front-end experiences down to the health of your back end all within one one offering. So with that, I'm
3:28
going to jump into my first question which is Christine. Uh welcome to the show. Um start off by could you start
3:35
off by telling us what front-end observability is and how that is different from backend observability?
3:41
Yeah. So, front-end observability measures the app experience as the user
3:47
experiences it on their actual device, their browser, their mobile phone and
3:53
the connectivity. All of that affects front-end observability.
3:58
Whereas backend observability measures the services and the infrastructure that
4:03
they run on. So all of this ties to the type of data that we collect and how you
4:08
want to slice and dice it. So for the front end, you don't control the
4:14
environment at all. So you don't control the device. You know, we have a lot of
4:21
folks that have a number of Android devices that users are using. Some of them are old, some of them are new. You
4:28
don't get to control the OS version that they're on. You don't get to control the connectivity. Someone is walking through
4:35
a tunnel, going on an elevator. All of those things are environmental and
4:40
scenarios that you're not able to control versus the back end where you are able to control the machines and the
4:46
software that's behind it and what version you're on. The second part um uh
4:52
to continue the theme of not having control is it's not just your code. So,
5:00
it is third-party SDKs that you're implementing in your app to
5:07
look at your analytics to help with push notifications, CDNs.
5:13
There's a lot of code that has to interact with your mobile or web app that doesn't belong to you. And so,
5:20
those are the sources of some of the problems. I'll talk with front-end
5:25
developers and the first thing that they want to know is is it my code or is it someone else's code and that is like a
5:32
good fork in the road for them to decide okay is it my problem or do I need to complain about this to someone else
5:41
and the last piece is more mobile specific but you actually can't control
5:47
when your fix is adopted so mobile does releases maybe every one to two weeks.
5:53
And there are times when people will release a version
5:58
that has a ton of errors. Hopefully not, but it can go out in the wild. That
6:05
version stays out in the wild for potentially years. I've seen
6:10
instances where some people never update. And so you get haunted with some
6:16
of these mistakes that um are made. So mobile developers really try and aim to
6:22
code for perfection because they don't know when when their code is actually going to the
6:29
new version will get updated. So again like I think the big difference between front end and backend observability is
6:36
control and the interactions that you are you're going through.
6:43
That's a really good explanation and as you were talking about that it gave me
6:48
uh shivers at the time I used to be a front-end developer actually a mobile developer way back in time and I know
6:55
exactly how that feels. Uh now I work in the backend observability space and yes it's you're absolutely right it's it's
7:01
something it's fully control it's not fully controlled environment but it's something we fully understand as opposed
7:06
to what's happening out in customer devices customer front-end experiences and so on. So thanks for that excellent
7:12
explanation. Uh let's jump into a few more details. Can you tell me what types
7:18
of met metrics front-end teams use uh and and to to get the outcomes they
7:23
need? Yeah. So I break it out typically into
7:28
two camps. One is reliability, stability
7:33
and the second is performance. So for reliability, the tried andrue for mobile
7:39
is your crash free rate. When I look at most mobile apps, I want to say it's
7:45
99.7% crashree sessions. Um, in the past 5
7:50
years, I cannot remember a time a single time where my app has crashed. I think
7:55
we've moved far beyond that. And most apps have a really good crash free rate and they're quite stable in that way. on
8:03
web you have JS errors, unhandled exceptions. I think the challenge with
8:09
those is um you're not really able to see the impact. So for those you might
8:15
see a high number of events, a high number of users affected, but you actually don't get a feel for what did
8:22
the user do when they experienced that error. Did they leave because the page was unusable or nothing happened and it
8:31
was perfectly fine. So even if it's the number one issue on your list,
8:36
maybe it's not even worth your time. So that's the reliability camp. And then
8:42
the second part of it is performance. How fast is it? So web has this
8:48
beautiful thing called core web vitals. You have largest contentful paint, your
8:54
LCP, INP, um interaction to next paint which is measuring responsiveness and
9:01
cls cumulative layout shift measuring the smoothness of it. Um it's amazing
9:09
because one it's easy to capture you don't have to do any extra implementation. The browser already has
9:16
the representation of the page in the DOM and as a result the standard browser
9:21
APIs can measure these metrics without requiring any application specific implementation.
9:28
The other part of it is that the web PF community has really invested in this
9:33
and so they have looked at it and seen okay what actually
9:40
affects conversion. So they've gone through iterations of what core web vitals looks like and they really try
9:47
and tie it to the data to say what is the most impactful thing towards conversion. How can we make this easy
9:54
for everyone to have better web performance. So it really is like a gold
10:00
standard in terms of measuring performance. Mobile, we're about 15 years behind on
10:10
that. Um, so there aren't a lot of out ofthebox
10:16
ways of measuring performance and even with the out ofthe-box solutions,
10:21
everyone approaches it in a slightly different way and there's no consistency behind it and that's a challenge. I
10:28
think it in the end it makes it hard for developers to aim for better performance
10:33
if there's no consistent definition of what good performance should be. And so
10:39
I think that has has made it challenging for mobile
10:44
development and people that are focused on mobile performance to move in the same direction as web. The things that
10:51
I've seen, they're not all great. They're but they're a good starting point. Um I see
10:59
with a lot of traditional enterprise appdex score, which is measuring
11:06
response times. um that is the start of something. I
11:12
think what I like where I tend to guide people is that is
11:22
it a lot of the good uh networking calls that you have will drown out the really
11:28
important bad ones. So for example in your app, your e-commerce app, you're
11:34
looking at products and you get a ton of 200s. Everything is very fast. Your checkout's slow. The number of
11:40
networking calls that you experience before checkout and the number of networking calls you have during
11:46
checkout, it's a high ratio. So in in the situation of appdex just by pure
11:53
volume, the very impactful errors are actually going to get drowned out. So, I
11:59
like the attempt at measuring it without implementation, but I don't think it gets people what they really want, which
12:05
is how does this tie to revenue? Like, is it going to point me in the direction of something that is going to have a
12:11
meaningful change in my app and am I going to get the ROI that I expect if I
12:17
try and improve this score? [clears throat] The other things that I've seen are um
12:24
time to initial display um and then for an attempt at
12:30
responsiveness on Android you have your ANR rate on iOS you have your iOS hang
12:36
rate and slow and frozen frames. So my whole thing with mobile is like
12:44
and the embrace team is is pioneering this and we're trying to establish core
12:50
mobile vitals and not for the sake of just saying we capture these metrics out
12:56
of the box and it is very closely tied to core web vitals. I think there is the
13:02
spirit behind core core web vitals which is there's constant research behind it
13:08
and whether or not it actually impacts revenue. Does it actually move the
13:14
needle when you make an improvement to a screen load? Is that actually going to
13:19
improve conversion? Are you going to get more people that are adding to cart, more people that are booking reservations with you? And so I think
13:27
that aspect of it, our team has put a lot of time and effort to one try and
13:35
auto instrument it so that you don't have to put in a lot of effort and you can have a starting point to improve the
13:43
mobile experience. That's pretty cool to hear actually and
13:48
given I'm you know closely working with you guys on this I also know about that but still very very exciting that that's
13:54
going on. Uh what this completely reminds me of is in the backend space there are all these um methods and best
14:02
practices that organizations follow. So for example, if you're monitoring Kubernetes or virtualized environment,
14:08
you have these best practices provided by the originators of those products could be an open source, you know, if
14:14
Kubernetes, it's open source. So it's coming from that from that group of people. In more vendor defined products,
14:19
it's coming from them saying this is how to monitor my database or this is how to monitor my virtualized environment. uh
14:25
or and take that forward. It sounds a lot like the RED method as well. Reed for those who don't know uh RED metrics
14:31
sometimes they're called which is rate errors and duration. So me measuring sort of requests per second or fail
14:39
requests per second in the case of errors and duration things like latency P50s, P90s, P99s. So it creates this
14:46
from from what I understand what I'm understanding is like these are the standards and methods to follow for as best practices. So with that with my
14:55
little siloquy over there going back to going back to mobile and web metrics
15:01
what's uh what's missing from these metrics that people need? Um I think there is first uh for the
15:10
metrics that I've already talked through it's the amount of data. Uh what I
15:15
typically see is that people sample front-end data like they sample backend
15:20
data. And to go back to what we talked about in the beginning, it's not the same. So if you have a very standardized
15:29
environment, if you take 1% of uh 1% of data, it's still quite
15:36
representative of everything else that's happening on the front end. You have
15:42
low-end Android devices that are running on this operating system. You have high-end Android devices with this
15:48
screen size running on this system. So you have these smaller populations and
15:53
so when I see people that are heavily sampling their front end um their
16:00
front-end data because of cost reasons because it's cost prohibitive to to look
16:06
at more data. That tends to be like the first starting point is do you actually have enough data to understand how your
16:14
users are experiencing this? Because with 1% sampling, you might not capture
16:22
the low-end Android devices, but that may also make up 1% of your revenue. All
16:28
of these things are are highly impactful to the bottom line. And so I think the
16:34
first thing is for the metrics that you collect, are you collecting enough data
16:39
to actually make the right decisions? And the second part of the metrics that
16:46
people are already collecting is do you actually have the context and the
16:52
information to solve it if it's not a great number. So for example, ANRS and
17:00
iOS hangs, Google and iOS provide those numbers um
17:06
for developers to look at. If you have a new version that comes out and then
17:12
there's a delta or a change, it's easy to correlate that with a new version.
17:18
But say that it's not correlated and a change is not correlated with a
17:23
particularly new version. There actually isn't a lot of information for developers to actually
17:30
solve that. So what I've heard for from developers on
17:35
the Android side is the stack trace has to be the right stack trace.
17:41
They may take that stack trace at a point in time that doesn't really help them fix the problem. And on iOS, it's
17:49
an aggregated number with no details on the con the context of what is
17:54
happening. So we have some good data that is provided
18:02
but not enough data to actually make it actionable. And so for the the
18:08
metrics that I've already listed, I think the biggest thing is having enough data um one and then two is just um
18:18
having enough context to fix it. The thing that I stress and I emphasize
18:25
or people to measure that they're not currently measuring is actually the user experience. So once the payload hits,
18:33
what happens? That is the actual bridge between your backend and your front-end teams.
18:40
That payload lands and how long does it take for the user to do what they had intended. Um my the
18:48
example that I I really like to talk I typically talk through is
18:54
LCP on a product details page. You have your image that loads for a page and you
19:02
say, "Okay, I have the image of the product LCP is fast, but what if the add
19:12
to cart button never shows up or it takes 1 to two seconds?" I know I want
19:18
it and my LCP may be very fast because I'm optimizing for that because we're
19:25
measuring that. But if I'm not measuring the time to something useful and as a
19:30
user that is when the add to cart button is is available and present for me to
19:37
click on then there's a disconnect that is actually the key part of the conversion
19:43
and so there may be a lot of networking calls coming up and it's creating a lag
19:49
and so the add to cart button never shows up. So what you want to do is actually measure the user experience and
19:55
what you intend the user to do and a lot of that requires manual instrumentation
20:01
uh user timing on web and performance traces on mobile and I think this is a
20:08
newer concept for front-end teams. I think back-end teams are very very
20:14
comfortable with tracing um much more comfortable with front-end teams and so
20:20
it really has been an exercise I think of education and and allowing people to
20:29
to to really explore and go okay how is my code being experienced by the user
20:37
and I think those are the metrics that I see most teams missing but it does
20:42
require a performance-based culture and understanding how are my metrics
20:47
actually uh affecting my end user. Interesting you mentioned backend teams
20:53
are more comfortable. We see a mixed results out there when it comes to backend teams and using tracing as well.
20:58
Really depends on how much they are how their organization is uh geared towards performance monitoring essentially how
21:05
how how is the product being perceived. I had two questions. Yeah, go ahead. Sorry. I'm curious to look like uh
21:11
back-end teams what is the percentage of teams you would say are like comfortable with tracing versus not
21:18
it is it varies so much I think it comes down to are organizations leaning
21:24
towards tracing or not and so you what we see is when an organization is all in
21:30
on tracing and I say in quotes because they'll have tons of old legacy applications if they don't want to go instrument they want to modify the code
21:36
to spit out you know put use the hotel SDKs to spit out traces, but for the new applications, they may be all in on
21:43
tracing. Even then, they're heavily sampling that data because it's just too much data that it produces. So, it and
21:50
then you have organizations who really haven't bought into the tracing world and they are dabbling in it, but they're
21:57
still predominantly metrics and logs. And this is so it really is a mixed bag
22:03
out there. But the ones who are all in on tracing really get a lot more value out get a lot of value out of that kind
22:10
of performance monitoring you could say because they have a really really deep understanding of what's going on in their product and how their product in
22:16
this case the backend experience the backend product is being performing. Um
22:22
which you can't just get from metrics and logs. Not that those data pieces go away, they're still relevant. But traces
22:28
just give you a ton of rich data that you help you solve problems faster. Yeah, I think that is like what I've
22:34
seen is organizations that adopt tracing are opening themselves up to
22:41
understanding the unknown. So before it was errors. I know what I am looking for
22:48
and I want to see if it spikes. I want to see if I introduce these errors in a
22:53
new version. They're very defined problems and they're very
22:58
like very direct in terms of solving it. People that are adopting tracing, these problems are becoming more amorphous and
23:06
it's it's how long does it take? Um is it erroring out? It these are they're
23:12
it's introducing more complex problems for developers to solve, but they're
23:18
inviting it. And so they actually get more visibility and coupling it with logs and metrics, they have a much
23:25
richer understanding of what's happening in their system. And I think those teams are the ones that are winning because
23:31
they care about it. They're measuring it. They know it's uncomfortable, but they want to know what's actually going
23:37
on. Yeah, absolutely. I I'm I'm with you on that. I have two questions. Uh first one
23:43
is for the uninitiated like myself, what does ANR mean? What does LCP mean? And then I have an actual question after
23:48
that. Yeah. Uh ANR is your application is not responsive. Um for Google it is actually
23:57
a very a stricter definition of okay your application is not responsive for 5
24:04
seconds and you get prompted by the operating system that says do you want to quit? That
24:12
is a good measurement. But I don't know many people that will tolerate 5 seconds
24:17
of just nothing happening and then say yes I want to quit. Most people will leave
24:24
before that. So even that measurement when I talk to developers especially on
24:29
consumer apps your customer is already gone. Like the patience level that people have on mobile is very very low.
24:38
And so I typically guide people away from just looking at that metric. You want to see, okay, the main thread was
24:45
blocked. They weren't able to do something. Even something as long as 500 milliseconds will be impactful in in
24:52
their flow. And so the threshold for that is actually much lower and your users will tell you when they abandon
24:58
your flow. And then your second question was LCP. Correct. Yes.
25:04
Uh LCP is largest contentful paint. And so that is a core web vital. So that's
25:10
when the largest piece on your page has loaded. Typically it is like an image or
25:17
just like hero text that's on that page. Well, thank you for educating me. Um
25:23
before we're chatting before we chatted about tracing, you were talking about user experience. Can you elaborate on
25:28
what you mean by user experience exactly? Yeah. So, normally observability teams will look at an
25:36
endto-end trace and an API call and it measures the request out um and then the
25:43
response back and it stops the moment the payload gets there. User experience
25:49
I like to say is an extension of that. You want to see okay how long did it
25:55
take um for the user to do something actionable on the screen. It's not just
26:01
when the payload lands. The payload has the device has to do some work on the payload. So, for example, I've seen
26:08
instances where an API takes 100 milliseconds. All of your dashboards are
26:14
green and your customers are complaining that this is a slow app. I can't search.
26:22
I can't purchase. The button disappears. all of these things that [clears throat]
26:27
you're getting in the actual user complaints, but all your dashboards look green. And
26:33
so, because it's not measuring, okay, I'm actually waiting for the device to
26:39
do things. So, what I've seen is the back end gives you a payload and it gives it to you very quickly, but is
26:45
that the right payload? Um, is it structured in a way that is fast for the device? So, I've seen instances where
26:53
they give you a payload and you're supposed to parse the data, you're supposed to sort the data, and you're
26:58
working on a really old iPad or you're working on a low-end Android device.
27:04
That's going to take time. It's going to take time to show those search results.
27:09
And so, the user as it actually experiences that payload is much longer than okay, the 100 milliseconds it took
27:16
to get to the device. And so really um that is like what I emphasize is it
27:24
doesn't stop when the payload gets to the device. It's what the device has to do with that payload and measuring out
27:31
okay the time to something useful. I think that is like the thing that we want to put in our developers like way
27:39
of thinking is if the user can get to something useful
27:44
that is what you actually want to measure and then the optimization happens between the two teams and it's the bridge between the two two teams
27:52
then you can say hey yes I got the API quickly that's I I got the call and the
27:58
payload very quickly but it's not in the format that is actually usable for our
28:04
user base and then you can have a productive conversation and that's where
28:09
um I think it bridges the gap between back-end teams and front-end teams and
28:15
creates like a more harmonious experience for the user.
28:21
That's actually quite interesting. Uh it takes me back to a few roles ago before Chronosphere where I was was responsible
28:29
for teams uh for a consumer product essentially and we ran into these exact conversation which is well things are
28:36
being served up everything looks fine. Um but the images that we were serving up to images slash the the user profiles
28:43
we were serving serving up to people. Yeah. The blob was I mean the content was just too big. So ultimately the end user
28:50
experience was things were slow and whereas the back end was serving up exactly at the pace it ought to have.
28:56
Yes. So very interesting stuff. Yeah. I've seen instances where um even
29:02
images aren't optimized for the device that they're on. Right. And so they are
29:07
really high resolution images that you don't really need. Um, and so these are
29:14
things that that slip by because we can't be perfect. We can't we can't
29:21
fully release code that's absolutely perfect all the time. But if we have good measurements and
29:29
h it's almost like a backs stop like okay if we're measuring it and it's slow we at least know that there's a problem
29:35
and we can go and find that problem. So, I like to say you want to measure the user experience because that tells you
29:43
when something is wrong. There are so many different types of errors that no
29:49
one had can ever predict and especially with mobile. I've seen instances where
29:55
your network calls are happening out of order and it produces a blank screen. How are you supposed to know that? Like
30:02
there is no way that you are going to create a log that says I want to make sure networking call A is going to
30:09
happen before networking call B. If you do that then you would have coded for it to begin with. And so like let's relieve
30:17
developers of that burden of having to develop so perfectly. And if we have
30:23
measurements of oh they're a lot of users are dropping off on this
30:28
particular page go back and say okay why are they dropping off
30:35
and seeing that experience it helps them actually solve the problem. So again,
30:40
there is no way that we could even predict the number and types of errors
30:46
that um users will actually experience because people will surprise you in
30:51
terms of their behavior. And I have seen that multiple times. uh offline, online,
30:58
airplane mode, low power mode, um all of these things that like you could never
31:05
have thought of, but if your user base is doing it quite frequently, you kind of have to roll with the punches and
31:11
code for it. Yeah, it's very it's a very rubber meets the road kind of problem. You forgot one permutation, which is giving it to a
31:18
5-year-old child and seeing what happens when they when they start clicking because that is Yes. And then they change all of your
31:23
settings. [laughter] They change all of your settings and then you're like, I don't even know how
31:30
I got here. Yeah. No, it's it's very true. I sometimes I look at my son using some apps and I feel like, how could anyone
31:36
have ever thought of testing for a child who was going to do this combination of things?
31:41
Yes. So, uh, all right. So, that was super interesting as always as through this whole session. Uh, are there other
31:47
examples you could provide? Um, yeah. Uh, this is interesting
31:53
because I've seen mobile teams like in my time with working with mobile teams
32:00
grow a ton. So, I've seen when I was first starting out, I would meet mobile
32:05
teams where it was like one to three developers managing multi-million dollar
32:10
apps. Um, and they were building features, maintaining the the app, but
32:16
because they're only like three people, they always knew what was going on. things would break, but you only had two
32:22
other people to look at and YouTube like that team would figure it out. Now you
32:27
have a lot more feature teams. You have a lot more teams working on a particular
32:32
screen or screen on an app or a screen on a web page. And what ends up
32:37
happening is like a lot of times and specifically I see this more in
32:43
traditional enterprise, they don't talk to each other at all. Every team works on their one thing and
32:52
they have no idea what all of the other teams are doing. And so what you end up
32:57
happen what ends up happening is you have one team that will make a call to get the information.
33:04
The second team doesn't know that that information already exists will also make another call. The third team will
33:11
also make a call. And so then you have five duplicative network calls. all
33:17
grabbing the same information. Now, the like mildest I would say the
33:24
more mild result of that is slowness. You have five networking calls where you should only have one. They're all
33:32
relatively fast, maybe 100 milliseconds each, but you're creating some congestion. You're creating a queue, and
33:39
there's some slowness to that. I would say that's like a mild uh result of
33:45
that. The more extreme case is you're pulling that data five times and that
33:51
data may not be consistent each and every time. So then you get into these
33:56
weird states where a devel a developer will be like I have no idea how that
34:03
happened. Um why? Because the first time you called it maybe it's looking at inventory there were three left. The
34:11
second time you pulled it there was still three left. by the time you pulled it the fifth time, there's nothing left.
34:17
And so data changes every time you pull it. Um, and I I've seen developers look
34:25
at this and go, I've gotten into this weird state. I don't know h possibly how I could do it. And it's because no one
34:31
has the full visibility of what it looks like
34:37
like when it's all working together. everyone has a specific view of their
34:42
piece of the code and not a holistic view of everything happening together. And so again, when you're measuring that
34:50
user experience, you're just saying, I care from when that page starts loading to when something meaningful is
34:56
happening on that page. I don't know what errors could come up in that time
35:01
period. I'm open to whatever it could be, but I know if it's slow, I got to fix it. Um,
35:09
but yes, I see I see it a lot of times and people look at it and they're completely surprised and it's a deeper
35:15
architecture problem. Like it's how these teams are interacting with one another, how they're talking and how
35:21
they're making these decisions. Um, and I think that has also been a challenge
35:26
in terms of how do I know what's happening in my app
35:32
without having all of the pieces and the information there. Well, that's super super interesting.
35:38
Christine, I think we're now at the end of our session. I wanted to just thank you for everything you share. Did I miss
35:45
anything by the way? Anything you wanted to share? That was that was it. Yeah. All right. Uh this was super exciting. Um I used to
35:52
be a front-end developer actually to be more honest more correct a mobile developer and we dealt with this quite I
35:58
dealt with this directly quite regularly and so it's always enlightening to understand this in more detail and as
36:04
time has gone on how how these uh how there's now better solutions to the problems I used to face back then albeit
36:10
I'm no longer a developer. Do you have that visceral pain though when I talk about it? I did. I do. Which
36:16
is why I I didn't mention the specific apps I worked on cuz that, you know, we're on a call like this. [laughter]
36:22
There are very specific apps I dealt with. It was a specific app where the the the essentially the image is being
36:29
rendered by the back end were far too high res.
36:36
And I was wondering why it took so long that my screen and my cells would load, you know, the table the in the iPhone
36:42
iOS, I forget what it's called now, but the the table view would load pretty quickly. The content, the text would
36:48
load super quickly, but the images would just render like truly crazy slow. And ultimately, it took time, but I figured
36:54
out what it was. But I if you know I it so it's these kinds of things that it really, you know, struck me as uh yeah,
37:01
now there's better ways to monitor this. And thank God for those who are building apps today. But it's still a problem today. Like you
37:07
have a problem today. That's right. Because I when I talked about it, I could see your face. I was like, "Ah, I
37:12
think he's experience." Exact thing. Uh there's several other examples. When I was running a team, I
37:19
wasn't doing the apps myself, but I was responsible for the mobile app and back end. Anyway, um Christine, this has been
37:24
super helpful, uh enlightening. I'm going to hand it back to Colin. Uh and I
37:29
think we're open to Q&A at this point. Yeah. So, uh, thanks Aloc and Christine
37:35
for this awesome discussion. Uh, we've had a few questions come in, so I do want to get to them. And, uh, while
37:40
we're doing this, if uh, you digest some of this information that they've shared and you have questions, please feel free
37:46
to continue adding them. We'd like to answer, uh, as many as we have time for. Um, but I do want to start with one
37:52
that, uh, we we covered a bit, but always great to kind of like dive in a bit more, add a bit more color. Maybe
37:58
this is one where we can do a handoff between the two of you. Um because this is the the million-dollar question. Uh we had a question come in which is um
38:06
how do you truly observe full stack systems from the front end to the back
38:12
end? Uh I know we we get this question a lot. I feel like maybe Christine we can start um with you. Maybe you can share a
38:19
little bit about uh some of the questions you've gotten maybe from uh people more familiar with backend about
38:24
what they envision this looks like and kind of h how we've helped them in terms of understanding how you can take you
38:30
know user focus observability integrate it with what they're used to like traditional observability maybe some of
38:36
the gaps there etc and then maybe Aloc you can chime in with what Chronosphere customers has asked for what they
38:41
envision this looking like I think this is a good question to to dive in so I'll hand it to you Christine
38:46
yeah [snorts] so I I know I've been talking about user the user experience and measuring that. But um the way that
38:54
I like to talk to teams is first you want to measure that user experience and
39:00
get the aggregated information there like how long does it actually take? Does it take 1 second? Does it take 5
39:06
seconds? And then from there you start to narrow it down for those slow instances. Is it
39:13
related to an API call? Is it related to a back-end service? And so I think that
39:19
is like the first thing is you have to measure the user experience and have those aggregated that aggregated
39:26
information and truly connecting it to the back end is okay I have this API
39:33
call that is airing out or is particularly slow and then starting to trace it through the back end to see
39:39
what services it has touched and what are the potential issues that
39:46
it has experienced. Yeah. And sorry Christina, did I cut you
39:52
off there? No. No. Okay. Um, one addition in terms of like stitching it
39:57
together for a full end to end from from a full end to end visibility perspective, there are a couple ways you
40:04
can do this. One is very obvious. Any data that the front end uh telemetry is
40:10
producing that Christine has talked about throughout this whole session. How do you get those metrics or traces? How
40:16
do you get them into the same platform as your backend tool? That could be a chronosphere, it could be something else you use. How do you get them in one
40:22
place? Uh why that's important is because the most basic version of doing end to end getting achieving this end
40:29
toend visibility is having a tool that allows you to actually stitch this together. So if they're in one place, you can then build dashboards etc.
40:35
That's just the basic version of this, right? Where still the onus is on you as the owners of these services and
40:40
applications to be responsible for building these experiences for yourself. But at least now they're in one place.
40:46
You're not trying to fight three different tools to say what does this mean? Now let's go to the two other ways
40:52
to achieve well the next way next approach to do this which is without it being in a single tool. There truly is
40:58
AI today these days. You can actually have MCPs talking to several different tools and you can ask questions of these
41:04
MCP servers but again you need to have them somewhere and you can ask a ton of these questions and you're going to make
41:10
sure the MCP tools are in a good place as well. You can't you can't just say someone has an MCP doesn't necessarily
41:15
mean you have everything you need. Again, the onus is still upon yourself to go build what it means to stitch this
41:21
together, but now AI could help you stitch it together with varying results obviously and you can build scales and
41:28
then continue on from there. The version that I love is a combination of what I mentioned with AI but also a unified
41:35
tool that not just is taking all the data in but also is generating a knowledge graph of what does that mean and how and stitching it all together
41:41
for you under the hood. So you are doing less work uh to stitch it together. So
41:47
the MCP example or the dashboard example I gave still requires someone to know what they're looking for. The final
41:53
example is a product that actually strings things together for you under the hood, perhaps surfaces them, but at
41:58
least gives you access to that knowledge graph, you could say, and allows you to stitch it together. So, it's it's it's
42:06
not an the answer wasn't particularly uh shocking when I said put it in one tool, but it's the knowledge graph port, which
42:12
is how to stitch it all together. Tools that do that will take you a long way uh in getting those answers.
42:20
Back to you, Colin. Awesome. Thank you so much. Uh so we have another question. Uh I'm going to send this one to uh Christine first. Uh
42:27
but this is a question that we've gotten a lot. Um which is how do you convince stakeholders that core web vitals are
42:34
only the first step and that they should work on speed more? Maybe this is someone who's a web performance
42:40
engineer, but uh definitely like uh Christine, you touched on this a little bit about Corora vitals being this the
42:46
these Google metrics that they collected a lot of data to to come up with these numbers that are broadly representative,
42:53
right, of websites in general, but obviously every website uh is different. Every user's tolerance of that tool is
42:59
different. So can you talk a little bit about how core vitals are just the first step? There's a lot more that comes into
43:05
incorporating that into a user focused observability practice. Yeah, core web vitals are great and I
43:12
say that it's like a great standard to start with and it's a good building block.
43:18
Again, it's the measurement and the intent of what what it actually measures like product details. That tends to be
43:24
the thing that I always go back to. It's only measuring the [clears throat] the
43:30
largest contentful paint when you're looking at that product details page and it's not measuring something necessarily
43:37
useful. I think for core web vitals there is the plateau of you can improve
43:43
them to a certain point and then there's not additional benefit. You have to say
43:48
that core web vitals is the first stepping stone where you're taking that first piece of it. And then the actual
43:56
measurement that you are going for is okay time to something actually useful. And we see this with more mature
44:02
performance cultures is that they're changing the measurement. They're augmenting core web vitals and expanding
44:09
it to say time to something useful. And that's what they're optimizing for. And what you'll see is a lot of these teams
44:16
are doing trade-offs with you're doing experiments and there's high personalization like I want to
44:24
suggest to eloque these five items as you're checking out because I know that
44:29
you buy this frequently that has a cost um and it's not just
44:34
your largest contemp it's a part of it's beyond that it is I've already made my
44:41
purchase and it is a carousel at the very bottom like that is where you're actually tying the additional um average
44:48
cart value, the number of items you're putting in. Those are kind of the personalizations
44:54
um and ways that you're getting a little bit more. And you see this with a lot of
45:00
apps. They're investing in it because they know that they can get a little bit more at the very end with those impulse
45:06
buys. So the company itself has already invested in it and it's made a decision like this is very important. We just
45:14
have to start saying okay well now how do we measure it in a way that is that
45:21
is in line with the intention and core web vitals like I will always say is a
45:27
great great great starting point but teams grow out of it and so what is like
45:33
the next the next thing that your team needs to to measure and I know it's very
45:38
it's very tough it's a very uphill battle um to educate But we are here to
45:44
support you on that. Nice. And Christine, this is probably
45:50
similar to what you just covered, but uh in case this is an opportunity to to go a little bit more in depth, we did have
45:56
a question which was uh how do you connect your performance data to real business impact?
46:03
I like to think of this as like how is performance affect your funnel? And this
46:09
is a you know I'm going from my add to cart to
46:16
my my view cart like if it's slow is
46:22
that going to affect my conversion to the next stage. So always tying the
46:28
performance with the actual next stage that you're attempting to get to. So if
46:33
you're able to make that conversion correlation and you see okay as I improve speed the conversion to the next
46:40
step is increasing that is like the most direct way of correlating your
46:46
performance data with your real business impact and you have to do it piece by piece it's not going to be as direct as
46:53
like okay I have uh increased this like
46:58
ACV by $5 um but it's okay If I have a
47:04
higher add to cart conversion, then that means with the average number of costs
47:09
for the items in my basket, it will increase by this amount. So that's
47:14
that's how I'd connect performance data with the real business impact.
47:21
Awesome. Thanks, Christine. Um, we did have a question uh just asking about what are Embrace's open-source
47:27
offerings? Uh, can you dive a little bit into into that? Yeah. So our um mobile
47:34
SDKs and our web SDKs are open source. So you are free and open to use them. I
47:40
think the thing that is challenging with front-end data is like I talked about
47:46
things like going offline um not having connectivity. Those are some of the
47:53
challenges with sending the data and storing the data that Embrace manages
47:58
with its paid offering versus the open-source where you're managing that
48:04
yourself. So that's like the first fork in the road. The second is the actual visualization of everything. I think you
48:12
will get your logs, your traces, your aggre, you can do your aggregated
48:18
metrics from our open-source solution. But what really matters when you're trying to resolve these really
48:24
challenging problems like I talked about with the multiple networking calls, the
48:29
duplicative network calls or um seeing, you know, like very slow
48:37
instances, you actually wanted to see it visualized in a combined way. I see a
48:44
lot of teams that will use our SDK first to just to collect a aggregated data,
48:49
but when they really want to start to dig into it and resolve the problem, they realize that they need the context
48:55
and they need more detail behind it, they need the visualization of what is happening in the session and that
49:01
combined view to actually solve the problem. So I think our open source solution is great. I think it's a really
49:08
great starting point to start getting the base level metrics when the team
49:14
starts to realize okay I have this now I want to improve it and you you attack
49:20
the lowhanging fruit first and you have more challenging things to look at that tends to be when you want to come to the
49:26
embrace platform the other part of it is our user flows which is the product
49:32
funnel where you go okay I'm going from step one to step two and you're discovering new issues Again, that is
49:38
very specific to our platform and it's a back-end way of analyzing the data. Um,
49:45
that tends to be where we see the differentiation. But I do recommend most people like if you're really starting
49:52
with metrics and you want to get a feel for it, our open source solution is
49:57
really good. Awesome. Thanks, Christine. And uh I
50:04
think we have I think one more question uh so far. So I think I'm going to send this uh your way aloque. And the the
50:11
question was uh I'm looking for full OTEL traceability from front end to backend with spans with a single trace
50:18
ID and not to get robbed with ingestion costs. Uh the word costs comes up. I'm sure Alo, you're excited to to chime in
50:25
here a little bit about Chronosphere, but uh would you like to answer this one? Yeah, I'll take a quick stab. I mean
50:30
again keep in mind there's nuance in all situations. So it you know the mileage will vary but one of the things I
50:36
mentioned about Chronosphere at the beginning of this call was that one of our pieces of claim to fame is that we
50:43
help customers try to ma not try to strongly manage their cost profile
50:49
relative to the value they're getting. So we won't get into the metrics and logging portions of this because that's not what's relevant today. But our
50:55
tracing product has a very rich feature set that does some things that are
51:01
common in the industry or common in tracing platforms but goes beyond that as well. So obviously we have head sampling, we have tail sampling but we
51:07
also have these thing additional things called user behavior sampling uh meaning sorry behavior sampling uh automated
51:14
behavior sampling. So what you can do is create all these rules in your system that and now with AI you can more easily
51:20
create those rules you know talk to use AI to generate rules more uh rapidly and
51:25
and more automatically if you'd like. Uh but basically there'll be conditional rules. So let's say you're facing an
51:31
outage. You can keep your sampling rate really low or high depending on how you know describe it. But basically collect
51:36
very few traces up until the point there's an actual incident and then you can or an incident is beginning and you
51:42
can use what we call behaviors in the chronosphere platform to say now I want more of those traces. This way allows
51:48
you to keep your overall cost pretty low but when you need the data you can start pulling that data in. Conditional
51:53
behavior is the key here which is when this particular service or this issue spikes which is going back to like if
51:59
you had some of these metrics in one tool you can use the metrics to say I'm seeing a spike in the data now collect
52:04
some more traces or go back to baseline and that overall helps you constrain costs quite a bit.
52:13
Awesome. Thank you so much Aloque. Well I I think we've gotten through all the questions uh that we have time for and
52:19
we're getting up to the top of the hour. So I want to thank you so much Aloque and Christine for this awesome
52:24
discussion and uh answering all these wonderful questions. Um I'd like to thank everyone for being here in
52:30
attendance today. Uh I hope this session helped you understand a bit more about the challenges of measuring app
52:36
performance, reliability and enduser experience in web and mobile apps. I know Christine you gave a lot of great
52:42
uh examples of some of the challenges that both these front-end teams face. Um, hopefully now you have a good
52:47
starting point for evaluating your own user focused observability practices. Uh, and if you have any questions or
52:53
want some guidance on how to improve them, we are happy to help. Uh, so I want to give one more thanks again. Alok
53:00
and Christine, thanks everyone for being here. Uh, have a fantastic day.
53:05
Thanks everyone. Thank you everyone. Thanks everyone. Bye.
Highlights
3:48
Front-end observability is defined by what you don't control.Not the device, the OS version, or the network condition — and on mobile, not even when a fix reaches users, since a broken release can sit in the wild for years.
7:31
Mobile performance monitoring is still lagging behind the web's.Core Web Vitals work because they're free to capture and tied to real conversion data — native mobile has no equivalent yet, so teams fall back to blunt instruments like Apdex scores.
15:12
1% sampling can quietly erase your lowest-end users.Front-end traffic is far more fragmented than back-end infrastructure, so a sampling rate that works for server metrics can drop entire device and OS cohorts from view.
25:39
Green dashboards, angry users.An API can return in 100ms and still feel broken — the payload still has to be parsed, sorted, and rendered on-device before it's actually usable, and that gap is invisible to backend-only monitoring.
43:17
Core Web Vitals are a first step, not a finish line.Mature performance teams graduate to measuring "time to something useful" tied to a specific product moment, not just a generic paint metric.